Coding / AI Engineering / Context Engineering / billion-context
CHAPTER 04 / 11 · CONTEXT ENGINEERING
Basic approaches
The common context-management toolkit
Most agent systems combine several methods: keep a recent window, compact older turns, prune noisy outputs, retrieve relevant sources, persist useful state, and isolate focused tasks. Each changes a different part of what reaches the model.
THE NOTEBOOK
Concise learning notes · source-grounded mechanismsThe practical starting point
Ask two separate questions: what should be visible on this request? and what should remain stored for later? A database can retain the entire transcript while a request contains only a small working set.
The methods below are complementary. The same application can use a recent window, a summary, retrieved evidence and persistent notes together.
1. Sliding window / history trimming
Keep the last few turns or a token-limited suffix, alongside essential instructions. Treat a tool call and its result as a valid unit; cutting through the middle of that exchange can produce an invalid message sequence.
Useful for: short support exchanges and local conversational continuity. Tradeoff: an older constraint may leave the working view even when it still matters. Store or pin critical state separately.
Examples: LangChain short-term memory ↗ · OpenAI Agents SDK session-memory cookbook ↗
2. Summarization / compaction
Replace older turns with a shorter account of goals, decisions, unresolved problems and progress; retain recent turns verbatim when useful. Compaction can be triggered by a token threshold, requested explicitly, or run in the background. Its scope depends on the implementation.
Useful for: maintaining continuity through a long task. Tradeoff: a summary can omit exact details, and repeated rewriting can introduce drift. Saving the original transcript provides a separate recovery path; the summary itself is not an inverse encoding.
Example: Claude compaction overview ↗
3. Selective pruning / tool-result clearing
Remove or shorten specific low-value content: stale file dumps, duplicate reads and already-consumed logs. Keep recent or protected results. This targets noisy items rather than necessarily summarizing the whole conversation.
Useful for: tool-heavy coding and research. Tradeoff: a result that seems stale can become relevant again. Preserve artifacts or allow re-reading.
Example: Claude context editing ↗
4. Retrieval / RAG / just-in-time loading
Keep documents or historical records outside the prompt. At query time, select relevant passages through keyword search, semantic search, database queries or file tools and insert them into the working context. RAG combines retrieval with generation; a vector database is one implementation option.
Useful for: large document collections and occasional exact lookups. Tradeoff: an answer depends on finding the right evidence. Chunking, ranking and query quality affect what is found; a search miss is not proof that the source lacks the information.
Example: LangChain retrieval ↗
5. Structured memory / persistent notes
Write durable facts and project state into explicit records: user preferences, constraints, decisions, TODOs and artifact paths. Load selected records when needed. Unlike a chronological transcript, these records organize what the task should remember.
Useful for: returning to a project or carrying state across sessions. Tradeoff: extraction can miss details; outdated notes need updating. Keeping a fact does not automatically keep the evidence or conversation that produced it.
Example: LangChain long-term memory ↗
6. Context isolation / subagents
Give a focused task its own context and return a concise finding plus source references to the coordinating agent. The exploratory trace stays outside the coordinator’s working view.
Useful for: separable research or code investigations. Tradeoff: coordination and extra model calls add work, and a handoff may omit assumptions. Isolation can reduce the main agent’s context without reducing total system token usage.
Example: Anthropic context-engineering guide ↗
Compare the mechanisms
| Method | What changes in the working view? | What needs separate care? |
|---|---|---|
| Sliding window | Only a recent suffix stays visible | Older constraints |
| Compaction | Older turns become a summary | Exact details and summary drift |
| Selective pruning | Chosen outputs disappear or shrink | Later re-reading |
| Retrieval | Relevant stored passages enter on demand | Search coverage and source retention |
| Structured memory | Selected durable state is loaded | Freshness and provenance |
| Context isolation | Detailed exploration moves to another agent | Handoff quality and coordination cost |
A combined design
Instructions + durable state + summary + recent turns + retrieved evidence can all contribute to one request. Their combined size still has to fit the model’s input budget.
A larger context window provides capacity. Capacity alone does not choose which information deserves the space. The next chapter examines billion-context’s particular combination of these familiar ideas.
CHECK YOUR INTUITION
A summary omits an exact test value. What determines whether the agent can recover it later?
Choose an answer to reveal the reasoning.
FOLLOW THE SOURCE
Common patterns checked against official provider and framework documentation. Examples illustrate mechanisms rather than market-share rankings.
LangChain · short-term memory ↗Claude · compaction ↗Claude · context editing ↗LangChain · retrieval ↗LangChain · long-term memory ↗Anthropic · context engineering ↗OpenAI · session memory ↗