Skip to content

JSON String Length: Escaped Source vs. Decoded Text

A JSON string has a written representation and a decoded value. Count the representation when you need the length of the JSON text; parse it first when you need the length of the text it contains. Then choose the unit: UTF-16 code units, Unicode code points, grapheme clusters or UTF-8 bytes.

For example, "\u00E9" contains 8 UTF-16 code units as written, including the quotation marks. Its decoded value is é: 1 UTF-16 code unit, 1 code point and 2 UTF-8 bytes. The difference comes from measuring different inputs.

What to include in the count

A complete JSON string literal includes its opening and closing quotation marks. Those delimiters disappear from the decoded value. A backslash escape also becomes the value it describes: \n represents a line feed, while \\ represents a literal backslash. The escaping rules are defined in RFC 8259, section 7.

If you are debugging a field limit, check whether the receiving system measures the JSON source or the parsed field. A payload-byte limit and a decoded-text limit require different measurements. There is no single JSON character limit that applies to every API.

Worked counts for exact JSON literals

The source column below includes the delimiting quotes. Decoded counts exclude those quotes and measure the value returned by JSON.parse. The emoji example uses a valid surrogate pair.

JSON sourceSource UTF-16 unitsDecoded UTF-16 unitsDecoded code pointsDecoded UTF-8 bytes
"abc"5333
"\n"4111
"\u00E9"8112
"\uD83D\uDE00"14214
"\\n"5222

The last row decodes to a backslash followed by the letter n, not a line feed. That is why its decoded length is 2. Count the exact representation you received, rather than replacing every backslash by hand.

Measure the decoded string in JavaScript

This example starts with one complete JSON string literal. String.raw keeps the backslash spelling in the JavaScript source so that JSON.parse performs the JSON decoding step. A normal JavaScript string literal adds another escaping layer.

const source = String.raw`"\uD83D\uDE00"`;
const value = JSON.parse(source);
if (typeof value !== "string") {
  throw new TypeError("Expected a JSON string literal");
}
console.log(source.length);                         // 14
console.log(value.length);                          // 2
console.log([...value].length);                     // 1
console.log(new TextEncoder().encode(value).length); // 4
const segments = new Intl.Segmenter("en", { granularity: "grapheme" });
console.log([...segments.segment(value)].length);   // 1

JavaScript string length counts UTF-16 code units. String iteration counts code points; a code point count can still differ from what looks like one character. Unicode grapheme clusters approximate user-perceived characters. Check that Intl.Segmenter is available in your target runtime before using the final line.

JSON.parse throws for invalid JSON, and valid JSON can represent an object, an array or another value instead of a string. Keep the type check when this calculation is specifically for a string. For a whole document, parse it and select the intended string field before measuring that field.

Measure a payload byte limit separately

new TextEncoder().encode(source).length counts the UTF-8 bytes of the supplied JSON text. Use the complete serialized payload as the input when the limit applies to the payload. Measuring one decoded field leaves out keys, punctuation and the other values. This calculation excludes compression and protocol overhead.

Using Count Character with JSON text

The Count Character counter uses JavaScript string length for its with-spaces character total. Pasting escaped JSON measures the pasted representation; the counter does not decode JSON first. Paste the decoded text when that is the text you want to measure, and check the destination’s own counting rule before relying on the result.

For the difference between code units, code points and bytes outside JSON, see the String Length guide.