Why Evaluate Agents?
Agent evaluation is critical for understanding performance, catching regressions, and improving reliability. Without measurement, you're flying blind.
Key Metrics
Important metrics to track for agent systems.
Task Success Rate
Percentage of tasks completed correctly.
Efficiency
Steps taken, tokens used, time elapsed per task.
Accuracy
Correctness of agent outputs and decisions.
Reliability
Consistency across repeated runs of the same task.
Evaluation Approaches
Different ways to evaluate agent performance.
Unit Tests
Test individual tools and components in isolation.
Integration Tests
Test the full agent loop with mock environments.
Benchmarks
Standard task suites for comparing agents.
Human Evaluation
Expert review for nuanced quality assessment.
Mini evaluation: three attempts with the same rubric
Run the task below with a model three times yourself, paste the answers and assess each criterion. This worksheet does not call a model or provide model answers.
Use the same model version, prompt and settings, starting a fresh conversation without earlier answers each time. Record system instructions, available tools, temperature and a seed if applicable. Do not select only the best answers. Start a separate series when comparing another model.
Mark unknown settings as unknown. These three repetitions only show variation on this one task; they do not measure general model quality.
Identical prompt for each attempt
Fictional source text for this exercise: [Q1] Workshop Orion starts on Tuesday at 09:30, lasts 90 minutes and has six participants. [Q2] Room B12 has eight seats and a projector. Room C4 has four seats and no projector. Choose a room with a projector that can accommodate everyone. Respond with exactly two bullet points: (1) the room and justification with source labels; (2) the weekday, start and calculated end time with source labels. Use only the source text.
Rubric defined before the test
- Room: B12 is selected, correctly justified by eight seats for six participants and the projector.
- Time: Tuesday, a 09:30 start and the calculated 11:00 end are correct.
- Sources and format: exactly two bullet points, appropriate [Q1] and [Q2] references alongside the claims, and no unsupported additions.
An attempt succeeds when all three criteria are met. You provide the judgments; equivalent wording is allowed. Only an answer with all three criteria rated enters the success rate.
- Fully evaluated
- 0 / 3
- Successful attempts
- 0
- Success rate
- Not yet available
Successful / fully evaluated attempts: 0 / 0. Open attempts: 3. Open attempts do not count as failures.
This worksheet is kept only in the current view. Save your answers and notes before leaving or reloading the page.
Common LLM Benchmarks
Standard benchmarks used to evaluate and compare language model capabilities across different tasks.
MMLU
Massive Multitask Language Understanding - 57 subjects from STEM to humanities. Tests broad knowledge.
HellaSwag
Commonsense reasoning about everyday situations. Tests understanding of physical world.
HumanEval
Code generation benchmark with 164 programming problems. Tests coding ability.
GSM8K
Grade school math word problems. Tests multi-step mathematical reasoning.
ARC
AI2 Reasoning Challenge - science questions requiring reasoning beyond pattern matching.
MATH
Competition-level mathematics problems. Tests advanced mathematical reasoning.
Benchmark Caveats
- ⚠Benchmarks can be gamed - models may be trained on test data
- ⚠High scores don't guarantee real-world performance
- ⚠Many benchmarks are saturated - top models score similarly
- ⚠Benchmarks often miss important capabilities like following instructions
LLM-as-a-Judge
Using language models to evaluate other model outputs - a scalable but imperfect approach.
How It Works
A capable LLM (the "judge") is prompted to evaluate outputs from another model. The judge scores responses on criteria like helpfulness, accuracy, and safety.
Advantages
Scalable
Can evaluate thousands of outputs quickly without human annotators.
Consistent
Same criteria applied uniformly (unlike human fatigue/variation).
Cost-effective
Much cheaper than hiring human evaluators at scale.
Flexible
Easy to adjust evaluation criteria by changing the prompt.
Problems & Biases
Self-preference Bias
Models tend to prefer outputs similar to what they would generate.
Position Bias
Judges may favor the first or last option regardless of quality.
Verbosity Bias
Longer responses often rated higher even when less accurate.
Style Over Substance
Well-formatted wrong answers may beat poorly-formatted correct ones.
Capability Ceiling
Judge can't reliably evaluate outputs beyond its own capability level.
Best Practices for LLM Judges
- →Use the most capable model available as the judge
- →Randomize option order to mitigate position bias
- →Request reasoning before scores (chain-of-thought)
- →Validate against human judgments on a subset
- →Use multiple judges and aggregate scores
CLASSIC Framework
The ICLR 2025 CLASSIC framework evaluates five dimensions: Cost, Latency, Accuracy, Stability and Security. Use the paper’s definitions and report your task, model and evaluation setup.
C - Cost
Total cost of ownership including API calls, compute, infrastructure, and maintenance. Track cost per task and cost per successful outcome.
L - Latency
Time to first token, end-to-end response time, and task completion time. Critical for user experience and real-time applications.
A - Accuracy
Correctness of outputs measured against ground truth. Includes factual accuracy, logical consistency, and task-specific precision.
S - Stability
Consistency of outputs across identical inputs. Low variance indicates reliable behavior; high variance suggests unpredictable performance.
S - Security
Resistance to prompt injection, jailbreaks, and data leakage. Includes input validation, output filtering, and access control.
Keep the five original dimensions distinct. Add interpretability or compliance as explicit project criteria where relevant, with their own definitions and checks.
Agent-Specific Benchmarks
Modern benchmarks designed specifically to evaluate AI agents on complex, multi-step tasks in realistic environments.
AgentBench
Evaluates LLMs as agents across 8 environments: OS, database, knowledge graph, web browsing, and more. Tests real-world tool use.
GAIA
General AI Assistants benchmark with 466 questions requiring multi-step reasoning, web browsing, and tool use. Human-verified answers.
Berkeley Function-Calling Leaderboard
Tests function calling accuracy across simple, parallel, and nested calls. Includes real-world API scenarios and edge cases.
SWE-bench
Real GitHub issues from popular Python repos. Agents must understand context, write code, and pass existing tests.
WebArena
Tests agents on realistic web tasks across e-commerce, forums, and content management sites with complex multi-page workflows.
TAU-bench
Tool-Agent-User benchmark testing agents on real customer service scenarios with tools, policies, and user interactions.
Interactive Evaluation
Dynamic evaluation approaches that test agent behavior in changing environments and adversarial conditions.
Beyond Static Benchmarks
Static benchmarks have fixed questions and answers. Interactive evaluation tests how agents adapt to dynamic environments, handle unexpected situations, and maintain performance under changing conditions.
Environment Perturbation
Change the environment during task execution—modify files, alter API responses, introduce errors—to test agent robustness and recovery.
Adversarial User Simulation
Simulate users who give ambiguous instructions, change their minds, or try to manipulate the agent. Tests real-world resilience.
Multi-Turn Consistency
Evaluate coherence across long conversations with context shifts. Check if the agent maintains accurate state and follows instructions over time.
Curriculum Difficulty
Start with easy tasks and progressively increase complexity. Identifies capability boundaries and graceful degradation patterns.
Interactive evaluation better predicts real-world performance than static benchmarks alone.
Best Practices
Guidelines for effective agent evaluation.
- ✓Test edge cases and failure modes, not just happy paths.
- ✓Track costs alongside quality metrics.
- ✓Use versioned evaluations to catch regressions.
- ✓Include adversarial tests for security.
Key Takeaways
- 1Evaluation is essential—unmeasured systems can't be improved
- 2Combine automated tests with human evaluation
- 3Track multiple metrics: success, efficiency, cost
- 4Build evaluation into your development workflow
- 5LLM-as-a-judge is useful but has significant biases to account for
- 6Use the CLASSIC framework for comprehensive enterprise evaluation
- 7Agent-specific benchmarks like AgentBench and GAIA test real-world capabilities
Additional project criteria: interpretability and compliance may be useful, but are not two extra CLASSIC dimensions.