Context engineering: CCDV-F study guide
CCDV-F · Prompt and Context Engineering, topic weight 3.8% of the exam
Context Engineering sits in Prompt and Context Engineering, which is 11.0% of the CCDV-F exam, and carries 3.8% on its own. It tests how you keep the context window small and relevant as a conversation or agent task grows: what to prune, when to compact, and when to give work to a subagent with its own context.
What the official guide covers
The Claude Certified Developer Foundations exam guide (version 1.0, effective July 2026) describes this topic as context and memory management techniques for Claude applications.
| What the guide lists | What it means in practice |
|---|---|
| Context window management | Know what fills the window on every turn and measure it |
| Prevention of context bloat: tool output pruning | Return only what Claude needs from a tool, and clear old tool results |
| Prevention of context bloat: compaction | Replace older turns with a summary when a long task must continue |
| Prevention of context drift | Keep the facts and instructions the task depends on in view as history grows or is summarised |
| Context isolation through subagents or multi-step workflows | Run noisy work in a separate context and pass back only the result |
| Memory management | Keep notes outside the window and load them when they are needed |
What goes wrong as context grows
Everything in a request counts toward the context window, including the output. In an agent loop, tool results pile up fastest: search results, file contents and logs that Claude used once.
Two problems follow:
- Bloat. Each turn resends the whole history, so cost and latency rise turn by turn. Anthropic's docs also describe context rot: as the token count grows, accuracy and recall degrade.
- Drift. The facts that matter get buried under newer material or lost when history is summarised. Claude starts working from recent noise instead of the original instructions and confirmed details, such as an order number the customer gave in the first turn.
Anthropic's engineering guidance frames the goal as the smallest set of high-signal tokens that gets the outcome you want. A larger window fixes neither problem.
Techniques, from lightest to heaviest
| Technique | How it works | Cost to you |
|---|---|---|
| Prune at the source | Your tool code returns selected fields, a page of results or a length-capped excerpt | Nothing is lost that Claude needed |
| Clear old tool results | Context editing removes older tool results server-side once input passes a threshold | Claude can call the tool again if it needs the data |
| Just-in-time retrieval | Keep IDs, paths or queries in context and load the content with a tool when needed | Extra tool calls |
| Memory | Claude writes notes to storage outside the window and reads them back later | Notes must be kept current |
| Compaction | Older turns are replaced by a summary | Detail in the summarised turns can be lost |
| Subagent isolation | A subagent explores in its own context window and returns only a summary | The main agent sees conclusions, not raw data |
Tool result clearing is available on the Messages API as context editing, behind the context-management-2025-06-27 beta header. The clear_tool_uses_20250919 strategy clears older tool results once input passes your trigger, keeps the most recent results (keep), and never touches tools you list in exclude_tools. Clearing happens on the server before the prompt reaches Claude. Your application keeps the full, unedited history. A second strategy, clear_thinking_20251015, manages old thinking blocks. Clearing tool results invalidates cached prompt prefixes from the point it clears, so there is a trade-off with prompt caching.
Compaction summarises older turns. On the Messages API, server-side compaction (in beta) writes the summary for you and returns it as a block that takes the place of the turns it summarises. The summary keeps only what the summarising prompt asks for. A generic "summarise the conversation" tends to drop task state such as which files changed, which option was chosen at a decision point, and which error was hit and how it was fixed. The API lets you replace the default summarising prompt with your own instructions inside the compaction object; name every field and decision that must survive, plus the user's latest open request. The same applies if you write your own summariser. In Claude Code, /compact summarises the conversation and accepts focus instructions, such as /compact keep the failing test names. /clear starts a new conversation with empty context, which is the better choice when you switch to unrelated work.
Subagents run in their own context window. In Claude Code, a subagent starts with its own system prompt and the delegated task, not your conversation history. It does its exploration there and returns only a summary, so file contents and logs never reach the main conversation.
Clearing old tool results (Python)
import anthropic
client = anthropic.Anthropic()
def prune_search(raw: dict) -> dict:
"""Tool output pruning: send Claude only what it needs."""
return {
"total": raw["total"],
"results": [
{"id": r["id"], "title": r["title"], "snippet": r["snippet"][:300]}
for r in raw["results"][:10]
],
}
response = client.beta.messages.create(
model=MODEL,
max_tokens=4096,
tools=TOOLS,
messages=messages, # your full history; the API edits its own copy
betas=["context-management-2025-06-27"],
context_management={
"edits": [
{
"type": "clear_tool_uses_20250919",
"trigger": {"type": "input_tokens", "value": 60000},
"keep": {"type": "tool_uses", "value": 3},
"exclude_tools": ["get_case_facts"], # never clear the confirmed facts
}
]
},
)
print(response.context_management.applied_edits) # what was cleared, and how many tokens
prune_search runs in your tool handler before the result enters messages. Context editing then clears older results once the prompt passes 60,000 input tokens, keeping the last three, and never clears get_case_facts.
Stopping drift
Pruning and compaction reduce bloat, but both can remove a detail that later turns depend on. Protect those details explicitly:
- Keep confirmed facts (IDs, amounts, decisions, constraints) in one structured block that you send on every turn outside the history being summarised, or return them from a tool that is excluded from clearing.
- Give compaction instructions that name what must survive.
- Put standing instructions in the system prompt, not in an early user turn that may be summarised away.
The CCAR-F guide covers the same problem from the architect side in 5.1 Context across long conversations.
Measure with production-sized tool output
Test fixtures are usually small. A tool that returns a short record in testing may return a record plus attachments and history in production, several times longer. A session that ran twenty turns cleanly in development can then fill its budget by turn eight.
- Measure each tool's result on the largest real input you can find, not on fixtures.
- Log input tokens per turn, and gate requests with the token counting endpoint before they would exceed your budget.
- Read the symptom carefully. When tool selection gets worse after a fixed number of turns, the system prompt and early instructions are often being crowded out by old tool results. Check how full the window is before you rewrite tool descriptions. If the same tool was chosen correctly in earlier turns, the description is not the cause. Raising
max_tokenswill not help either: it limits the reply, not what Claude reads.
Choosing what survives between sessions
| Memory scope | What carries over | Fits | What you give up |
|---|---|---|---|
| In-context only | Nothing after the session ends | Short sessions that fit in the window | All state at session end; cost grows every turn |
| External storage | State saved to a database or file and loaded at the start | The same user or task across days, or several agent instances | Read and write code, plus a lookup on each start |
| Summarised memory | A summary injected into the next session | Long conversations that would outgrow the window | Any detail the summary did not keep |
| Stateless | Nothing, by design | Jobs that finish and close, such as formatting one file | Any follow-up that depends on an earlier job |
Choose the scope at design time. Replaying every past session in context works in a demo and then fails a few sessions into real use, when the injected history alone fills most of the budget.
Which technique fits
| Symptom | Choose | Why |
|---|---|---|
| Tool results take up most of the window | Prune in the tool code, then clear old tool results | Removes the bulk without losing the conversation |
| A long task must continue and old turns matter only in outline | Compaction | A summary keeps the thread in far fewer tokens |
| Exact values from early turns are wrong after compaction | A persistent facts block outside the summarised history | Summaries lose detail; the block is resent unchanged |
| Exploring a codebase floods the main context | A subagent | Exploration stays in its own window; only the summary returns |
| Work spans several sessions | Memory notes or a progress file | The next session reads the notes instead of the old history |
| A Claude Code session is full of unrelated earlier work | /clear | Start fresh rather than summarise noise |
| Tool choice degrades after the same number of turns each session | Measure context per turn, then prune or compact earlier | Old results are crowding out the instructions |
Rules that decide exam answers
- A bigger context window is not the fix. Context rot and cost still grow with every token.
- Prune before you summarise. Clearing old tool results is the lightest change; compaction loses more.
- Summaries lose detail. Facts that must stay exact belong in a block that is never summarised or cleared.
- Subagents return summaries, not transcripts. That is what isolates the main context.
- Context editing is server-side and breaks the cache where it clears. Your client keeps the full history; weigh clearing against the savings from prompt caching.
- Never strip thinking blocks yourself to save space. In a tool loop they must go back unchanged; use the API's thinking-clearing strategy instead.
Where it appears in the exam
Context Engineering is 3.8% of the exam, inside Prompt and Context Engineering (11.0%). The guide's description points to questions about long-running agents, long conversations and coding sessions where context grows: rising cost per turn, a window that fills up, or an agent that forgets early instructions or facts.
Two sample questions
These are original Timo practice questions. They are not official exam questions.
Build exercise
- Build an agent loop with a search tool that returns large results. Log
usage.input_tokenson every turn. - Add
prune_searchto the tool handler and compare the per-turn input. - Add context editing with a low trigger, run a long session and print
applied_editsto see what was cleared. - In Claude Code, run
/context, then move a "map this repository" task into a subagent and run/contextagain to compare.
Practise this topic
- Claude Certified Developer practice exam: free, 20 questions, no sign-up
- CCDV-F study guide: all topics
- Same topic in another exam: CCAR-F 5.1 Context across long conversations
- Previous topic: Cost and Token Management
- Next topic: Prompt Engineering
Sources
- Claude Certified Developer Foundations Exam Guide, version 1.0, effective July 2026 (Anthropic), topic: Context Engineering (Domain 6, Prompt and Context Engineering)
- Anthropic documentation: Context editing
- Anthropic documentation: Compaction overview
- Claude Code documentation: Create custom subagents
- Anthropic: Effective context engineering for AI agents
By Amotion AI