OneLess field notes Text essentials · 03
Why Do Character Counters Give Different Results?
You paste the same text into two counters. The numbers disagree. Before cutting a sentence, check what each counter is actually counting.
Open Character CounterThe rule explains the number.
The quick answer
“Character” can mean more than one thing.
First, check whether both counters include spaces and line breaks. Then check their counting unit: a reader-perceived character, a Unicode code point, or a UTF-16 code unit. A combined emoji or accented letter can produce different totals under those rules.
A difference alone does not prove a bug. Compare the exact input and the stated counting rule before deciding which result you need.
01 / Start with the simple difference
Do spaces and line breaks count?
In OneLess, Characters includes spaces and line breaks. Without spaces excludes whitespace, including tabs and newlines. These two totals answer different questions.
A B
3 characters · 2 without spaces.
A B
3 characters · 2 without spaces, using one LF newline.
A line that wraps because the text box is narrow is different from an inserted newline. Resizing the box does not add characters. For files, newline encodings can also differ: CRLF contains two code points, while LF contains one. The examples here use LF.
02 / Look beneath the label
Three rules behind the number.
For ordinary English letters, these counts often match. Emoji and combining marks make the distinction easier to see.
Grapheme clusters
Groups that approximate the characters a reader perceives. A letter with a combining accent can form one group.
e + ◌́ → one graphemeUnicode code points
The individual Unicode values in the text. The letter and combining accent in that example are two code points.
U+0065 + U+0301 → twoUTF-16 code units
The units counted by JavaScript’s string length property. Some code points, including 😀, need two of these units.
"😀".length → 2Unicode describes grapheme boundaries in Unicode Text Segmentation. MDN explains JavaScript’s UTF-16 counting in its String.length reference. Graphemes approximate perceived characters; they are not a count of every shape drawn by a font.
03 / Same text, different totals
A few characters tell the story.
All three columns below include whitespace. The accented examples deliberately use different underlying sequences, even though they can look identical.
| Text | Graphemes | Code points | UTF-16 units |
|---|---|---|---|
| A BOne space | 3 | 3 | 3 |
| 😀One smile | 1 | 1 | 2 |
| éU+00E9 · precomposed | 1 | 1 | 1 |
| ée + combining accent | 1 | 2 | 2 |
| 👨👩👧👦Family sequence | 1 | 7 | 11 |
The family sequence contains four emoji code points joined by three zero-width joiners. Its appearance depends on your font and device. The totals above describe the underlying text, even if your device displays the sequence as separate symbols.
Try it yourself
Four lines. One reproducible check.
Copy this block into Character Counter. It has one ordinary space and three LF newlines, with no newline after the family emoji.
A B 😀 é 👨👩👧👦
The third line uses an e followed by a combining accent.
The same block has 16 code points and 21 UTF-16 code units.
OneLess uses Intl.Segmenter to count grapheme clusters when available. In browsers without it, the fallback counts code points, so combined characters can produce larger totals. Word Counter uses the same character-counting rule.
04 / Match the count to the task
Which number should you use?
Read the destination’s rule
For a submission form, caption, or application, check whether its limit includes spaces and how it treats special characters. Use the receiving field’s own validation as the final check. A general counter cannot guarantee that every platform will accept the same text.
Compare the exact same text
Include the same title, punctuation, blank lines, and trailing spaces in both tools. An extra newline at the end is easy to miss. If totals differ only after pasting, check whether the destination changed the text.
Find the smallest example that disagrees
Try a plain word first, then a space, an emoji, and an accented letter. The comparison table helps identify a counting-rule difference. If identical text still disagrees under the same rule, a bug or a segmentation-version difference may need investigation.
A little more clarity
Questions behind a surprising total.
Can invisible characters still count?
Yes. A zero-width space can be present without looking like an ordinary blank. OneLess’s Without spaces total does not remove every invisible Unicode character. Do not delete all invisible characters blindly: joiners can be part of an emoji sequence.
Are bytes the same as characters?
No. Bytes measure encoded data, and their total depends on the encoding. A field asking for a byte limit needs a check for that encoding; a character total alone is insufficient.
Does Word Counter count characters differently?
OneLess’s Word Counter and Character Counter share the same character-counting approach. Their main emphasis differs: one puts words first, the other puts characters first. Compare the character totals using the same pasted text.
Should I remove spaces to make the text fit?
Only if removing them makes sense for the content. For accidental spacing, Remove Extra Spaces can help. For awkward copied line breaks, see our PDF text cleanup guide. Read the result before submitting; a smaller count is not useful if the meaning changes.