CSV Field Length: Quoted Source vs. Parsed Value
Published · Updated
Count a CSV field after parsing when the limit applies to the imported value. Surrounding quotes and extra quotes used for escaping belong to the file spelling. For example, "a,b" has five code points in the CSV text and becomes the three-code-point value a,b.
If the limit applies to an upload file or encoded bytes, measure that exact representation instead. A field value, a quoted field, a record and a complete file have different boundaries.
What happens to the quotes
Under the common comma and double-quote rules in RFC 4180, quotes around a field mark its boundary. A quote inside the value is written twice. The file spelling "He said ""go""" therefore becomes He said "go": sixteen code points become twelve.
For an always-quoted field under this convention, the raw code-point length is n + q + 2, where n is the value’s code-point length and q is its number of literal double quotes. Each literal quote adds one extra escape quote, and the surrounding pair adds two. An exporter that leaves an ordinary field unquoted has different overhead.
Five field lengths before and after parsing
These examples quote every field, preserve its text and use UTF-8. Each raw-field count excludes the delimiter and record terminator. In the displayed spellings, \r\n stands for an actual carriage return and line feed: two code points, rather than four literal backslash-letter characters. The empty value is labelled explicitly.
| Parsed value | Raw CSV field | Parsed code points | Parsed UTF-8 bytes | Raw code points | Raw UTF-8 bytes |
|---|---|---|---|---|---|
a,b | "a,b" | 3 | 3 | 5 | 5 |
He said "go" | "He said ""go""" | 12 | 12 | 16 | 16 |
é😀 | "é😀" | 2 | 6 | 4 | 8 |
line1\r\nline2 | "line1\r\nline2" | 12 | 12 | 14 | 14 |
(empty) | "" | 0 | 0 | 2 | 2 |
Parse the field before counting it
The Python csv module reads fields using the chosen dialect. This example uses a StringIO stream with newline preservation, a comma delimiter, doubled quote escaping and strict parsing. It prints each parsed value, its code-point count and its UTF-8 byte count.
import csv, io
raw = '"a,b","He said ""go""","é😀","line1\r\nline2",""\r\n'
reader = csv.reader(io.StringIO(raw, newline=""),
delimiter=",", quotechar='"',
doublequote=True, strict=True)
values = next(reader)
for value in values:
print(repr(value), len(value), len(value.encode("utf-8")))
The output counts are 3 / 3, 12 / 12, 2 / 6, 12 / 12 and 0 / 0, in that order. Python strings contain Unicode code points. The pair é😀 has two code points and six UTF-8 bytes; JavaScript String.length counts three UTF-16 units for the same text. Visible-character counts can use another rule; choose the unit named by the receiving system.
For a file, open the text stream with its known encoding and newline="", then pass that stream to the reader. Configure the delimiter and quoting to match the exporter. A semicolon delimiter or backslash escape convention needs different settings. Handle parse errors before treating a length as valid.
Keep commas, spaces and embedded line breaks
Splitting CSV text at every comma would break the quoted value a,b. Splitting at every newline would also break the multiline value. Its embedded CRLF remains two code points in this example; the record-ending CRLF is outside the value. Spaces in the value remain part of its count. Trimming or changing newline spelling changes the measured input.
The W3C tabular-data guidance recommends UTF-8 and describes dialect and line-ending choices. The non-ASCII examples here use that declared encoding. They do not imply that every CSV exporter or importer uses UTF-8 or preserves text identically.
Match the count to the documented limit
Use the parsed value for a field-text rule, its encoded bytes for a field-byte rule, or the original file bytes for a whole-upload rule. A whole file can include headers, separators, record endings and an encoding marker. Re-exporting parsed values may change the original spelling, so a newly generated CSV is not a measurement of the original file.
A parser’s configured field-size limit and the destination’s field limit are separate checks. This guide supplies no universal maximum or acceptance guarantee. Use the String Length guide for counting units, and Line Counter for the separate question of lines versus records.