The strawberry problem is a symptom
The famous question "how many r letters are in strawberry?" is not interesting because the answer is hard. It is interesting because the error feels alien. A model can summarize a paper, write code, and reason through a plan, then stumble on a small character-level task.
A better name: subtoken blindness
Subtoken blindness is the tendency of token-based language models to lose reliability when the task depends on structure inside tokens: individual letters, repeated characters, exact digit positions, separators, or column-by-column arithmetic.
Not just tokenization
Tokenization is the entry point, but not the whole explanation. Research on letter counting finds that models can often recognize letters yet fail to count them consistently, especially when counts exceed two or when the task requires stable step-by-step state.
Tasks worth testing
Strawberry and cranberry
A correct answer for strawberry alone does not establish reliable letter counting. Test held-out words, counts, prompt variants and a named model version. Different results do not by themselves reveal whether a model memorized, counted or used another strategy.
Long digit addition
Multi-digit addition looks simple to humans because we use a strict algorithm. A base LLM sees a string of tokens and predicts text. Without tool use or a learned reliable procedure, it may produce plausible-looking arithmetic.
IDs, filenames, and weird strings
The same pattern appears in hashes, SKUs, coupon codes, exact casing, invisible whitespace, and off-by-one character positions. The content has little semantic meaning, so fluency is the wrong capability.
Try the tokenizer view
Token view vs. character task
The tokenizer sees reusable chunks. The human task often asks for letters, digits, or exact positions inside those chunks.
Count the letter r
This counts characters in code. Token boundaries alone do not determine whether a language model will answer correctly.
Why this matters
These failures are useful because they reveal the boundary between language fluency and exact symbolic manipulation. The model is not a database, not a calculator, and not a character array unless the system gives it tools or forces the structure into the context.
What works in practice
- Ask the model to spell out the string before counting.
- Use code, a calculator, or validation logic for arithmetic and exact string work.
- Treat viral benchmark fixes cautiously; nearby variants often expose the same weakness.
- Design agent tools so exact operations happen outside the language model.