Common Agent Failure Modes
Understanding typical agent failures helps in building more robust systems and setting appropriate expectations.
Tool Misuse
Agents may call tools incorrectly, with wrong parameters, or at inappropriate times.
Infinite Loops
Agents can get stuck repeating the same actions without making progress.
Goal Drift
Agents may gradually shift focus away from the original task objective.
Over-confidence
Agents may proceed with actions despite uncertainty or incomplete information.
Tool Hallucination
Agents sometimes "invent" tool parameters or even entire tools that don't exist. This usually happens when the tool definition is ambiguous or when the model tries to force a solution.
Looping Issues
Agents can get trapped in repetitive cycles where they perform the same action, receive the same error, and try again without changing strategy.
Cost & Latency
Agent workflows can contain several model calls and tool steps. Cost and latency depend on their size, prices and dependencies.
A worked cost and latency example
Chosen teaching inputs, not a provider quote or benchmark: input costs $1 per million tokens and output $4 per million. No cached tokens or tool fees. All token counts and model wait times below are assumed. The agent makes three dependent calls and waits another 0.4 s and 0.6 s for tools.
| Call | Input tokens | Output tokens | Model wait (s) | Cost |
|---|---|---|---|---|
| Single chat call | 1,000 | 200 | 2 | $0.0018 |
| Agent call 1 | 1,000 | 100 | 2 | $0.0014 |
| Agent call 2 | 1,800 | 200 | 3 | $0.0026 |
| Agent call 3 | 1,200 | 300 | 1 | $0.0024 |
Cost = (input tokens × $1 + output tokens × $4) / 1,000,000. The agent totals 4,000 input tokens and 600 output tokens: ($4,000 + $2,400) / 1,000,000 = $0.0064.
Single-call total: $0.0018 · 2 s
Three-call total: $0.0064 · 7 s
Sequential agent latency = 2 + 3 + 1 + 0.4 + 0.6 = 7 seconds. It is not inferred from the call count. Independent branches could overlap; their latency would follow the critical path instead of this sum.
Key Takeaways
- 1Implement safeguards like iteration limits and cost controls
- 2Add human-in-the-loop checkpoints for critical actions
- 3Monitor agent behavior and log all actions for debugging
- 4Design clear success and failure criteria
For each model call, count uncached input tokens, cached input tokens and output tokens at the applicable rates, then add tool costs. Sequential latency includes each model and tool wait; parallel branches contribute their critical path. Five steps do not imply five times the cost or latency.