Skip to content

UTF-8 Text File Size: Bytes, BOM and Line Endings

To measure a UTF-8 text file’s size, count the original file bytes. A decoded character count answers a different question. Some Unicode characters use several UTF-8 bytes, a leading byte order mark adds three bytes, and CRLF line endings use two bytes where LF uses one. Keep those choices separate before comparing a file with a byte limit.

A file containing é😀 followed by CRLF has eight UTF-8 content bytes. Add a UTF-8 BOM at the start and the original file has eleven bytes. After decoding with leading-BOM handling, the text has four Unicode code points: the accented letter, the emoji, carriage return and line feed.

Measure the original file before decoding

Python’s Path.read_bytes returns a file’s binary contents. Taking the length of that bytes object measures the original bytes, including a leading signature and original line endings. It loads the entire file, so the example below is for a small file you can read into memory.

A Python str is a sequence of Unicode code points. Its length describes decoded text. Code points are also distinct from user-perceived characters: a letter followed by a combining accent can contain more than one point. Use the String Length guide when the separate question is which text-counting unit your rule requires.

Compare six exact UTF-8 files

The table names the text after a possible leading BOM is skipped. Quoted text uses JSON notation: \n represents one actual LF, and \r\n represents CR followed by LF. The quotes and backslash spellings are labels, not extra file content. Text bytes come from re-encoding that decoded text as UTF-8; file bytes come from the original binary input.

Decoded textLeading BOMCode pointsText UTF-8 bytesOriginal file bytes
"A\nB"
No333
"A\r\nB"
No444
"é😀\n"
No377
"é😀\n"
Yes3710
"é😀\r\n"
Yes4811
""
Yes003

The empty-text row still contains three file bytes because the original file consists only of the BOM. The two emoji rows with LF have identical decoded text and content-byte totals; their original lengths differ by the signature. Replacing LF with CRLF adds one byte and one code point in these examples.

Read a known UTF-8 file with Python

For a small sample.txt containing a leading UTF-8 BOM, é😀 and CRLF, run this read-only snippet. It prints the original file length and two measurements of the decoded text.

from pathlib import Path

data = Path("sample.txt").read_bytes()
text = data.decode("utf-8-sig")  # Known UTF-8; skip a leading BOM.
print("File bytes:", len(data))
print("Text code points:", len(text))
print("Text UTF-8 bytes:", len(text.encode("utf-8")))

The output is File bytes: 11, Text code points: 4 and Text UTF-8 bytes: 8. The utf-8-sig codec skips the three-byte signature only at the beginning when decoding. Encoding the resulting text with utf-8 does not add that leading signature back.

This assumes the input is known UTF-8. It does not guess an unknown encoding. Strict decoding raises a decoding error for invalid UTF-8; replacing or ignoring invalid bytes changes the text you would count. A U+FEFF inside the text remains content and should not be removed as though it were the leading signature.

Preserve the line endings and name the rule

Default text-mode open uses universal-newline reading: CR, LF and CRLF become LF in the returned string. For the BOM-and-CRLF example, reading with encoding="utf-8-sig" and default newline handling produces three code points instead of four. Counting those translated characters or re-encoding them cannot establish the original file’s eleven-byte length.

Binary reading keeps the original bytes. If you need text-mode reading without newline translation, newline="" returns the original line-ending characters. Line counting is a further task: the Line Counter guide explains how to define a line before counting one.

The Unicode BOM guidance gives the UTF-8 signature as EF BB BF. UTF-8 has no big-endian or little-endian distinction; whether a BOM belongs in a file depends on its format or protocol. Follow the receiving system’s stated rule rather than adding or stripping it as a universal fix.

When a rule names complete file bytes, use the original binary length. When it names decoded text, state the encoding, leading-BOM handling, line-ending treatment and counting unit. Neither figure alone establishes disk space allocation, transfer overhead or acceptance by an upload service.