Subtoken Blindness

Intermediate

Why models can sound brilliant and still miss tiny letter, digit, and counting tasks that humans solve instantly.

Last updated: Sep 13, 2026

The strawberry problem is a symptom

The famous question "how many r letters are in strawberry?" is not interesting because the answer is hard. It is interesting because the error feels alien. A model can summarize a paper, write code, and reason through a plan, then stumble on a small character-level task.

A better name: subtoken blindness

Subtoken blindness is the tendency of token-based language models to lose reliability when the task depends on structure inside tokens: individual letters, repeated characters, exact digit positions, separators, or column-by-column arithmetic.

Not just tokenization

Tokenization is the entry point, but not the whole explanation. Research on letter counting finds that models can often recognize letters yet fail to count them consistently, especially when counts exceed two or when the task requires stable step-by-step state.

Tasks worth testing

Strawberry and cranberry

A correct answer for strawberry alone does not establish reliable letter counting. Test held-out words, counts, prompt variants and a named model version. Different results do not by themselves reveal whether a model memorized, counted or used another strategy.

Long digit addition

Multi-digit addition looks simple to humans because we use a strict algorithm. A base LLM sees a string of tokens and predicts text. Without tool use or a learned reliable procedure, it may produce plausible-looking arithmetic.

IDs, filenames, and weird strings

The same pattern appears in hashes, SKUs, coupon codes, exact casing, invisible whitespace, and off-by-one character positions. The content has little semantic meaning, so fluency is the wrong capability.

Try the tokenizer view

Interactive failure mode

Token view vs. character task

The tokenizer sees reusable chunks. The human task often asks for letters, digits, or exact positions inside those chunks.

strawberry
Tokens
3
Characters
10
Letter r
3

Count the letter r

Expected deterministic answer
3

This counts characters in code. Token boundaries alone do not determine whether a language model will answer correctly.

Reliable workaround: externalize the hidden structure. Spell characters one by one, or route arithmetic to code/a calculator.

Why this matters

These failures are useful because they reveal the boundary between language fluency and exact symbolic manipulation. The model is not a database, not a calculator, and not a character array unless the system gives it tools or forces the structure into the context.

What works in practice

  • Ask the model to spell out the string before counting.
  • Use code, a calculator, or validation logic for arithmetic and exact string work.
  • Treat viral benchmark fixes cautiously; nearby variants often expose the same weakness.
  • Design agent tools so exact operations happen outside the language model.

Sources and further reading