What is chat compaction?
Chat compaction is the deliberate compression of long agent conversations into smaller, structured representations. The goal is not just to fit more text into a context window. The real goal is to keep old work recoverable, so an agent can continue a task later without re-reading every raw turn.
Compaction is navigation, not truth.
Why agents need it
Active sessions
Long-running agents accumulate user turns, tool calls, results, retries, and intermediate reasoning. If nothing is compacted, the context window fills up and the agent starts losing useful recent state or paying heavily to keep it all alive.
Inactive sessions
Old sessions often fall out of hot cache entirely. Compaction gives you a resumable version of cold history, so the agent can restart from a compacted summary plus targeted recall instead of rehydrating the whole thread.
The four building blocks
Compress
Convert older turns into summaries once they no longer need to live in the hot context window.
Structure
Store compacted history in linked chunks or hierarchies so old work remains navigable instead of becoming a blob.
Retrieve
Pull only the relevant compacted fragments back when the current task depends on earlier decisions or facts.
Expand
Recover exact wording, values, and causal chains on demand instead of guessing from summaries alone.
The core pattern
fresh_context = recent_messages + tool_results
compacted_history = summaries + recall_index
if context_too_large:
compacted_history.add(compact(older_turns))
matches = retrieve(compacted_history, current_task)
if exact_details_needed:
restored = expand(matches)
context = fresh_context + restored
else:
context = fresh_context + matchesInteractive compaction explorer
Compare overloaded live sessions with old inactive sessions that need to be resumed later.
The original transcript remains in the conversation archive. The summary is an index for retrieving evidence, not an exact replacement for original statements.
Content sent in the next model request
- Decision, 10 September: do not publish until the integration test passes.
- Evidence: the integration test failed in run test-42.
- Current task: prepare a release note draft.
Original conversation archive (retained throughout)
- Decision, 10 September: do not publish until the integration test passes.
- Evidence: the integration test failed in run test-42.
- Current task: prepare a release note draft.
For the exact test-run identifier, retrieve the original evidence instead of reconstructing it from a summary.
What naive compaction breaks
Naive truncation drops old context entirely. Naive summarization keeps only a plausible paraphrase. Both approaches fail in the same way: they make old work cheaper to carry, but harder to trust.
- Important decisions lose their rationale.
- Open loops disappear, so the agent forgets what was still unresolved.
- Exact commands, paths, timestamps, and values get blurred into summary language.
- Old summaries can contradict newer evidence if they are treated like ground truth.
Good compaction in practice
Design rules
- Keep a fresh tail for active work.
- Compact old turns into structured, queryable units.
- Retrieve summaries for relevance, expand for precision.
- Let newer evidence override stale summaries.
Separation of concerns
- Fresh context is for immediate reasoning.
- Compacted history is for navigation and resumability.
- Expansion is for exact facts and causal detail.
- Long-term memory stores distilled facts, not every turn.
Key takeaways
- 1Compaction is not just about saving tokens. It is about preserving recoverability.
- 2Good compaction helps both overloaded live sessions and older inactive sessions that are no longer cached.
- 3Summaries are recall cues, not proof. Exact claims should come from targeted expansion.
- 4The best systems separate fresh context, compacted history, retrieval, and long-term memory.