URL Length Before and After Percent-Encoding
Published · Updated
Measure the URL text after the serializer has encoded it if you need its encoded length. Measure the original text separately if the limit applies to the value before encoding. A component, a query string and a complete URL contain different text, so name the input as well as the unit.
For example, é becomes %C3%A9 under encodeURIComponent: one UTF-16 code unit becomes six ASCII characters. The emoji 😀 becomes %F0%9F%98%80, which contains twelve ASCII characters.
Why percent-encoding changes length
A percent triplet consists of % and two hexadecimal digits. It represents one byte, as defined in RFC 3986, section 2.1. JavaScript’s encodeURIComponent first uses UTF-8 for characters it escapes. An escaped byte therefore occupies three characters in the result.
Do not multiply the original character count by three. Some characters stay unescaped; others need several UTF-8 bytes. The exact output depends on the serializer and the input.
Worked component lengths
These rows use exactly the text shown, without surrounding quotes. Original length means JavaScript UTF-16 code units; original bytes means UTF-8. The encoded output is ASCII, so its character count also equals its UTF-8 byte count.
| Original text | Original UTF-16 units | Original UTF-8 bytes | encodeURIComponent output | Encoded length |
|---|---|---|---|---|
a b | 3 | 3 | a%20b | 5 |
é | 1 | 2 | %C3%A9 | 6 |
😀 | 2 | 4 | %F0%9F%98%80 | 12 |
a+b | 3 | 3 | a%2Bb | 5 |
a/b | 3 | 3 | a%2Fb | 5 |
Measure the component, query or full URL
This runnable example records the original text, encoded component and query. String.length counts UTF-16 units; TextEncoder.encode supplies UTF-8 bytes. The query includes its key and equals sign. A complete serialized URL also includes the scheme, host, path and query marker.
const text = "é 😀";
const component = encodeURIComponent(text);
console.log(text.length); // 4 UTF-16 units
console.log(new TextEncoder().encode(text).length); // 7 UTF-8 bytes
console.log(component.length); // 21 ASCII characters
const params = new URLSearchParams({ term: text });
console.log(params.toString());
// term=%C3%A9+%F0%9F%98%80
console.log(params.toString().length); // 24 query characters
Use the input that matches the documented limit. For a byte limit, encode that exact serialized string with TextEncoder and measure the returned array. A component count leaves out the surrounding URL. These calculations do not establish a server’s limit or include HTTP headers and other protocol overhead.
A space is not always %20
URLSearchParams uses form-style query serialization: a space becomes +, while a literal plus becomes %2B. For a b, encodeURIComponent returns a%20b, but new URLSearchParams({ q: "a b" }).toString() returns q=a+b. For a+b, that query is q=a%2Bb.
Pass the raw value to URLSearchParams. Passing encodeURIComponent("é") as the value instead produces q=%25C3%25A9: the percent signs are encoded again. Measure the output from the same code path that will send the request.
Check input validity and count units
encodeURIComponent throws URIError for a lone surrogate. Handle that error or validate the input; silently replacing characters would change the text being measured.
The String Length guide explains why UTF-16 units, code points and visible characters can differ. For JSON escapes, use the separate JSON source and decoded-text guide; JSON escaping and URL encoding require different transformations.