Reading lines beyond SUB in Python

Question

Newbie question. In Python 2.7.2., I have a problem reading text files which accidentally seem to contain some control characters. Specifically, the loop

for line in f

will cease without any warning or error as soon as it comes across a line containing the SUB character (ascii hex code 1a). When using f.readlines() the result is the same. Essentially, as far as Python is concerned, the file is finished as soon as the first SUB character is encountered, and the last value assigned line is the line up to that character.

Is there a way to read beyond such a character and/or to issue a warning when encountering one?

score 8 · Accepted Answer · answered Mar 01 '12 at 17:16

8

On Windows systems 0x1a is the End-of-File character. You'll need to open the file in binary mode in order to get past it:

f = open(filename, 'rb')

The downside is you will lose the line-oriented nature and have to split the lines yourself:

lines = f.read().split('\r\n')  # assuming Windows line endings

answered Mar 01 '12 at 17:16

Ethan Furman

63,992
20
159
237

1

for linux line endings, use `lines = f.read().split('\n')` – Nik A. Jul 01 '16 at 06:36

score 6 · Answer 2 · answered Mar 01 '12 at 17:06

6

Try opening the file in binary mode:

f = open(filename, 'rb')

answered Mar 01 '12 at 17:06

NPE

486,780
108
951
1,012

3

thank you a thousand times. – Emre Aydin Aug 29 '15 at 10:01

Reading lines beyond SUB in Python

2 Answers2

Linked