Verifiable Rewards

Expert

How LLMs and agents can be trained on tasks with objective outcome checks, from coding sandboxes to Office-style environments.

Last updated: Sep 13, 2026

Training on outcomes, not opinions

Verifiable-reward training works when a model or agent can attempt a real task and an external checker can score the result. This is why math, code, browser tasks, and tool-using agents are so interesting: the reward can come from the world, not only from human preference labels.

The reward loop

  1. 1. Define the target

    Write the intended outcome and how to verify it. Passing a test must correspond to the user’s actual task.

  2. 2. Run a candidate

    Provide an isolated, resettable environment with appropriate data, tools and permissions.

  3. 3. Check the result

    Compare the resulting artifact or state with reference values, tests or constraints. Inspect the checker’s blind spots.

  4. 4. Use the feedback

    Rank candidates or collect training data. Updating model weights requires an actual training procedure; running this demo does not train a model.

Test the checker before trusting its reward

All functions and tests run locally. The task is to sum an array of integers. The public example alone cannot distinguish a general solution from an overfit answer. These pass counts describe only the shown test suite.

values => 3
TestInputExpectedActualResult
Public example[1,2]33Passed

Tests passed: 1 / 1

Compare the weak checker with the exact checker. More test passes under a weak reward do not establish task success. Held-out examples reduce this particular loophole; they are not a proof of correctness.

Realistic environments are the product

Code

Repository, visible tests, held-out regressions and runtime behavior.

Documents

Source data, formulas, expected totals and required structure.

Browser workflow

Seeded application state and checks of the final state, not just a click trace.

Why this is hard

  • A checker can reward a shortcut that violates the intended task.
  • Keep evaluation cases separate from optimization cases and inspect suspicious successes.
  • Test cases are finite. Generalization depends on task coverage and deployment conditions.