AI capability is not a smooth curve
A stronger model is not simply a model that is a little better at everything. The frontier is jagged: huge peaks appear in domains with clear feedback, abundant practice data, and verifiable rewards, while nearby-looking tasks can remain surprisingly brittle.
Success on a benchmark does not establish success on every similar-looking task. Letter-counting errors such as the historical "strawberry" example illustrate this point; they are not a current universal failure of language models. Evaluate the exact model, prompt, tools and task you intend to use.
Similar tasks, different success criteria
An evaluation plan, not a measured capability map. These pairs show why success on one task should not be assumed on its neighbor. Which model fails where must be measured.
Task A
Write a function that passes the visible examples.
Task B
Run that function correctly on empty inputs, Unicode and invalid data.
Check: Test ordinary and edge cases separately. Passing examples says nothing about untested inputs.
Controlled comparison: New test inputs with the same model and prompt.
Verifiable rewards
Math answers, unit tests, compiler errors, game scores, and benchmark checks create objective signals. If the model can try many attempts and receive reliable feedback, training can push hard on the behavior.
Synthetic practice loops
Once a task has a verifier, systems can generate new problems, sample many solutions, score them, and train on the winners. That makes coding and math unusually scalable compared with vague knowledge work.
Tool-shaped environments
Coding agents, browser agents, and office agents can be placed in realistic sandboxes. The model acts, the environment changes, and a verifier decides whether the goal was actually reached.
Why the valleys remain
Hidden exactness
Some tasks look semantic but depend on exact characters, positions, files, cells, or UI state. Token-based models can miss that structure unless tools expose it explicitly.
Weak verifiers
Many useful office, research, and planning tasks do not have a clean yes/no checker. Without a reliable target metric, reinforcement learning can optimize the wrong proxy.
Distribution shift
A model may be excellent on a benchmark-shaped version of a problem and fragile when the same skill appears in a messy real workflow with missing context and ambiguous goals.
The contrast that confuses people
Olympiad-style math
The final answer can often be checked. Training can reward correct derivations, reject wrong attempts, and build strong reasoning traces around the task format.
Coding with tests
A coding agent can edit files, run tests, inspect failures, and improve. The feedback loop is expensive, but the success signal can be very concrete.
Exact character operations
Character-level operations have a different success criterion from fluent language. Check exact output and tool use; do not assume that a given current model will fail this task.
Office and browser work
The model may understand the goal, but realistic environments need accounts, documents, state reset, permissions, UI recovery, and verifiers that know what success means.
Why this hurts adoption
Jagged capability makes AI feel unreliable in a very specific way. People see a system solve a hard problem, infer broad competence, then lose trust when it fails a simple-looking task. The visible failure is not just a bug; it breaks the user's mental model of what the system is.
How to work with the jagged frontier
- Ask whether the task has an objective success signal, not whether it feels easy to a human.
- Prefer workflows with tests, validators, scripts, checklists, or reviewable artifacts.
- Use tools for exact operations: counting, arithmetic, file inspection, spreadsheet logic, and browser state.
- Benchmark your actual workflow instead of trusting demos from a nearby but cleaner task.
- Treat task format, training, available tools and evaluation criteria as separate contributors.