HTML Character Count: Source, textContent and innerText
Published · Updated
An HTML character count needs two choices: the text to extract and the unit to count. Measure the original source when tags and character-reference spelling belong to the limit. Read textContent for descendant DOM text, or innerText for text shaped by rendering. Then count that selected string in code points, UTF-16 units or bytes.
For example, <p>A&B</p> contains fourteen source code points. Its paragraph text is A&B, with three. A source-code counter and a text counter are measuring different inputs.
Choose the text before measuring its length
HTML character references such as & and 😀 represent decoded text. Tags describe structure. Their spelling is part of the source string, while the resulting text belongs to DOM text nodes.
textContent concatenates descendant text without using its rendered appearance. It includes hidden descendant text and text inside script or style elements, but excludes comments. Select the content element you intend to measure rather than taking every descendant of the whole page.
innerText uses rendered appearance. Hidden descendants, line breaks and whitespace can change its result. For a detached or unrendered element, it falls back to textContent. Reading it can also trigger layout, so repeatedly measuring it inside a large loop has a different cost from reading textContent.
Six fragments with three different text scopes
The table uses attached paragraph elements in Chrome, white-space: normal, no text transformation and hidden descendants set to display: none. All counts are Unicode code points. Quoted text uses JSON notation: \n means one actual line feed, and two literal spaces remain two. These results describe the declared context, not arbitrary page CSS.
| HTML source | Source points | textContent | DOM points | innerText | Rendered points |
|---|---|---|---|---|---|
| 14 | | 3 | | 3 |
| 16 | | 1 | | 1 |
| 15 | | 2 | | 2 |
| 30 | | 3 | | 2 |
| 13 | | 2 | | 3 |
| 11 | | 4 | | 3 |
The hidden span keeps B in textContent but omits it from innerText. A <br> contributes a rendered line break here without adding a text node between A and B. The two source spaces survive in textContent; normal whitespace collapses them in innerText. Changing white-space to pre changes that rendered-space example.
Count an existing element with JavaScript
For a paragraph already present as <p id="sample">😀</p>, this snippet reads its DOM text and measures three units. It only reads the element; it does not insert or execute pasted HTML.
const element = document.querySelector("#sample");
if (!element) throw new Error("Choose an existing text element.");
const text = element.textContent;
console.log({
codePoints: [...text].length,
utf16Units: text.length,
utf8Bytes: new TextEncoder().encode(text).length
});
The result is { codePoints: 1, utf16Units: 2, utf8Bytes: 4 }. JavaScript String.length counts UTF-16 units. String iteration yields code points, while TextEncoder.encode yields UTF-8 bytes. Code points still differ from user-perceived characters: the table’s combining-accent value has two code points.
If rendered text is your chosen scope, read element.innerText instead and measure that returned string. For the line-break row it contains A\nB, so its three points include the line feed. State whether whitespace and line breaks belong to your rule before removing anything.
Keep original source and current DOM separate
innerHTML serializes current descendants as markup. It is not a recovery of the exact original file or response spelling, especially after parsing or page changes. Keep the original input when the limit applies to that source. Encoding extracted text as UTF-8 measures that text’s bytes, not the complete HTML response.
A form’s current field value and a content element’s descendant text are different scopes too. This guide measures the existing content element and supplies no universal CMS or platform limit. Use the String Length guide for the separate choice of counting units, or Keyboard Symbols when you need to enter a character rather than extract webpage text.