Understanding Agent Context
Agent context includes the system prompt, conversation history, tool definitions, and retrieved information. Managing this context efficiently is crucial for agent performance.
Think of context as the agent's working memory—everything it needs to understand the task and respond appropriately.
Explore Context Layers
Click each layer to see example content
Context Anatomy
Illustrative 8,192-token budget and preset counts, not a tokenizer measurement. The application must choose what to send. Overflow may be rejected or truncated according to the API and request options; the model does not automatically archive omitted content.
7,750 / 8,192 tokens
Defines the agent's role, capabilities, and behavioral guidelines.
Use only the supplied announcement. Do not infer weekend opening. Say “not specified” if asked for information absent from it.
What Actually Gets Sent to the Model
messages = [
// 1. System prompt (highest priority)
{ "role": "system", "content": "You are a helpful coding assistant..." },
// 2. Conversation history
{ "role": "user", "content": "Help me fix this bug" },
{ "role": "assistant", "content": "I'll read the file first" },
// 3. Tool calls and results
{ "role": "assistant", "tool_calls": [{"name": "read_file"}] },
{ "role": "tool", "content": "def buggy_function():..." },
// 4. Latest user message
{ "role": "user", "content": "Thanks, what was the issue?" },
]
tools = [
{ "name": "read_file", "description": "Read file contents", ... },
{ "name": "write_file", "description": "Write to a file", ... },
]Context Components
System Prompt
Defines the agent's role, capabilities, and behavioral guidelines.
Tool Definitions
Descriptions of available tools and how to use them.
Conversation History
Previous messages, tool calls, and their results.
Retrieved Information
External knowledge fetched during the conversation.
Context Management Strategies
Sliding Window
Keep the most recent N messages. Simple but may lose important early context.
Summarization
Periodically compress older messages into summaries. Preserves key information while reducing tokens.
context = [system, summary] + recent_messages
Priority-based Truncation
Assign priority scores to messages. System prompts and recent turns get highest priority.
Common Pitfalls
Context Overflow
An oversized request may be rejected or truncated according to the API and options. For example, the Responses API defaults to rejecting overflow when truncation is disabled. The application must manage its budget.
Tool Definition Bloat
Too many tools or overly verbose descriptions eat into context space. Keep tool definitions concise and only include tools relevant to the task.
Lost in the Middle
Some long-context retrieval experiments find lower accuracy for evidence in the middle. This is task- and model-dependent; test evidence placement and retrieval on your workload.
Stale Retrieved Data
RAG results from earlier in conversation may become outdated as discussion evolves. Refresh retrieved data when the topic shifts.
Key Takeaways
- 1Context management is key to agent reliability
- 2Prioritize recent and relevant information
- 3Tool definitions should be clear and unambiguous
- 4Summarization helps maintain context over long sessions