Context Anatomy

Intermediate

Breaking down the structure of context windows and how agents manage information.

Last updated: Sep 13, 2026

Understanding Agent Context

Agent context includes the system prompt, conversation history, tool definitions, and retrieved information. Managing this context efficiently is crucial for agent performance.

Think of context as the agent's working memory—everything it needs to understand the task and respond appropriately.

📊

Explore Context Layers

Click each layer to see example content

Context Anatomy

Illustrative 8,192-token budget and preset counts, not a tokenizer measurement. The application must choose what to send. Overflow may be rejected or truncated according to the API and request options; the model does not automatically archive omitted content.

7,750 / 8,192 tokens

Defines the agent's role, capabilities, and behavioral guidelines.

Use only the supplied announcement. Do not infer weekend opening. Say “not specified” if asked for information absent from it.

What Actually Gets Sent to the Model

messages = [
  // 1. System prompt (highest priority)
  { "role": "system", "content": "You are a helpful coding assistant..." },

  // 2. Conversation history
  { "role": "user", "content": "Help me fix this bug" },
  { "role": "assistant", "content": "I'll read the file first" },

  // 3. Tool calls and results
  { "role": "assistant", "tool_calls": [{"name": "read_file"}] },
  { "role": "tool", "content": "def buggy_function():..." },

  // 4. Latest user message
  { "role": "user", "content": "Thanks, what was the issue?" },
]

tools = [
  { "name": "read_file", "description": "Read file contents", ... },
  { "name": "write_file", "description": "Write to a file", ... },
]

Context Components

📋

System Prompt

Defines the agent's role, capabilities, and behavioral guidelines.

🔧

Tool Definitions

Descriptions of available tools and how to use them.

💬

Conversation History

Previous messages, tool calls, and their results.

📚

Retrieved Information

External knowledge fetched during the conversation.

Context Management Strategies

Sliding Window

Keep the most recent N messages. Simple but may lose important early context.

messages = messages[-MAX_MESSAGES:]

Summarization

Periodically compress older messages into summaries. Preserves key information while reducing tokens.

summary = llm("Summarize this conversation: " + old_messages)
context = [system, summary] + recent_messages

Priority-based Truncation

Assign priority scores to messages. System prompts and recent turns get highest priority.

priority: system > tools > recent_user > recent_assistant > old_history

Common Pitfalls

Context Overflow

An oversized request may be rejected or truncated according to the API and options. For example, the Responses API defaults to rejecting overflow when truncation is disabled. The application must manage its budget.

Tool Definition Bloat

Too many tools or overly verbose descriptions eat into context space. Keep tool definitions concise and only include tools relevant to the task.

Lost in the Middle

Some long-context retrieval experiments find lower accuracy for evidence in the middle. This is task- and model-dependent; test evidence placement and retrieval on your workload.

Stale Retrieved Data

RAG results from earlier in conversation may become outdated as discussion evolves. Refresh retrieved data when the topic shifts.

Key Takeaways

  • 1Context management is key to agent reliability
  • 2Prioritize recent and relevant information
  • 3Tool definitions should be clear and unambiguous
  • 4Summarization helps maintain context over long sessions

Primary sources