Jagged Frontier

Intermediate

Why frontier models can solve olympiad-level problems and still fail simple-looking tasks at the edge of their capability.

Last updated: Sep 13, 2026

AI capability is not a smooth curve

A stronger model is not simply a model that is a little better at everything. The frontier is jagged: huge peaks appear in domains with clear feedback, abundant practice data, and verifiable rewards, while nearby-looking tasks can remain surprisingly brittle.

Success on a benchmark does not establish success on every similar-looking task. Letter-counting errors such as the historical "strawberry" example illustrate this point; they are not a current universal failure of language models. Evaluate the exact model, prompt, tools and task you intend to use.

Similar tasks, different success criteria

An evaluation plan, not a measured capability map. These pairs show why success on one task should not be assumed on its neighbor. Which model fails where must be measured.

Task A

Write a function that passes the visible examples.

Task B

Run that function correctly on empty inputs, Unicode and invalid data.

Check: Test ordinary and edge cases separately. Passing examples says nothing about untested inputs.

Controlled comparison: New test inputs with the same model and prompt.

Verifiable rewards

Math answers, unit tests, compiler errors, game scores, and benchmark checks create objective signals. If the model can try many attempts and receive reliable feedback, training can push hard on the behavior.

Synthetic practice loops

Once a task has a verifier, systems can generate new problems, sample many solutions, score them, and train on the winners. That makes coding and math unusually scalable compared with vague knowledge work.

Tool-shaped environments

Coding agents, browser agents, and office agents can be placed in realistic sandboxes. The model acts, the environment changes, and a verifier decides whether the goal was actually reached.

Why the valleys remain

Hidden exactness

Some tasks look semantic but depend on exact characters, positions, files, cells, or UI state. Token-based models can miss that structure unless tools expose it explicitly.

Weak verifiers

Many useful office, research, and planning tasks do not have a clean yes/no checker. Without a reliable target metric, reinforcement learning can optimize the wrong proxy.

Distribution shift

A model may be excellent on a benchmark-shaped version of a problem and fragile when the same skill appears in a messy real workflow with missing context and ambiguous goals.

The contrast that confuses people

Possible strength

Olympiad-style math

The final answer can often be checked. Training can reward correct derivations, reject wrong attempts, and build strong reasoning traces around the task format.

Possible strength

Coding with tests

A coding agent can edit files, run tests, inspect failures, and improve. The feedback loop is expensive, but the success signal can be very concrete.

Separate check

Exact character operations

Character-level operations have a different success criterion from fluent language. Check exact output and tool use; do not assume that a given current model will fail this task.

Messy frontier

Office and browser work

The model may understand the goal, but realistic environments need accounts, documents, state reset, permissions, UI recovery, and verifiers that know what success means.

Why this hurts adoption

Jagged capability makes AI feel unreliable in a very specific way. People see a system solve a hard problem, infer broad competence, then lose trust when it fails a simple-looking task. The visible failure is not just a bug; it breaks the user's mental model of what the system is.

How to work with the jagged frontier

  • Ask whether the task has an objective success signal, not whether it feels easy to a human.
  • Prefer workflows with tests, validators, scripts, checklists, or reviewable artifacts.
  • Use tools for exact operations: counting, arithmetic, file inspection, spreadsheet logic, and browser state.
  • Benchmark your actual workflow instead of trusting demos from a nearby but cleaner task.
  • Treat task format, training, available tools and evaluation criteria as separate contributors.

Related concepts