Training on outcomes, not opinions
Verifiable-reward training works when a model or agent can attempt a real task and an external checker can score the result. This is why math, code, browser tasks, and tool-using agents are so interesting: the reward can come from the world, not only from human preference labels.
The reward loop
1. Define the target
Write the intended outcome and how to verify it. Passing a test must correspond to the user’s actual task.
2. Run a candidate
Provide an isolated, resettable environment with appropriate data, tools and permissions.
3. Check the result
Compare the resulting artifact or state with reference values, tests or constraints. Inspect the checker’s blind spots.
4. Use the feedback
Rank candidates or collect training data. Updating model weights requires an actual training procedure; running this demo does not train a model.
Test the checker before trusting its reward
All functions and tests run locally. The task is to sum an array of integers. The public example alone cannot distinguish a general solution from an overfit answer. These pass counts describe only the shown test suite.
values => 3
| Test | Input | Expected | Actual | Result |
|---|---|---|---|---|
| Public example | [1,2] | 3 | 3 | Passed |
Tests passed: 1 / 1
Compare the weak checker with the exact checker. More test passes under a weak reward do not establish task success. Held-out examples reduce this particular loophole; they are not a proof of correctness.
Realistic environments are the product
Code
Repository, visible tests, held-out regressions and runtime behavior.
Documents
Source data, formulas, expected totals and required structure.
Browser workflow
Seeded application state and checks of the final state, not just a click trace.
Why this is hard
- A checker can reward a shortcut that violates the intended task.
- Keep evaluation cases separate from optimization cases and inspect suspicious successes.
- Test cases are finite. Generalization depends on task coverage and deployment conditions.