TimoBy Amotion AI

Context engineering: CCDV-F study guide

CCDV-F · Prompt and Context Engineering, topic weight 3.8% of the exam

Context Engineering sits in Prompt and Context Engineering, which is 11.0% of the CCDV-F exam, and carries 3.8% on its own. It tests how you keep the context window small and relevant as a conversation or agent task grows: what to prune, when to compact, and when to give work to a subagent with its own context.

What the official guide covers

The Claude Certified Developer Foundations exam guide (version 1.0, effective July 2026) describes this topic as context and memory management techniques for Claude applications.

What the guide listsWhat it means in practice
Context window managementKnow what fills the window on every turn and measure it
Prevention of context bloat: tool output pruningReturn only what Claude needs from a tool, and clear old tool results
Prevention of context bloat: compactionReplace older turns with a summary when a long task must continue
Prevention of context driftKeep the facts and instructions the task depends on in view as history grows or is summarised
Context isolation through subagents or multi-step workflowsRun noisy work in a separate context and pass back only the result
Memory managementKeep notes outside the window and load them when they are needed

What goes wrong as context grows

Everything in a request counts toward the context window, including the output. In an agent loop, tool results pile up fastest: search results, file contents and logs that Claude used once.

Two problems follow:

  • Bloat. Each turn resends the whole history, so cost and latency rise turn by turn. Anthropic's docs also describe context rot: as the token count grows, accuracy and recall degrade.
  • Drift. The facts that matter get buried under newer material or lost when history is summarised. Claude starts working from recent noise instead of the original instructions and confirmed details, such as an order number the customer gave in the first turn.

Anthropic's engineering guidance frames the goal as the smallest set of high-signal tokens that gets the outcome you want. A larger window fixes neither problem.

Techniques, from lightest to heaviest

TechniqueHow it worksCost to you
Prune at the sourceYour tool code returns selected fields, a page of results or a length-capped excerptNothing is lost that Claude needed
Clear old tool resultsContext editing removes older tool results server-side once input passes a thresholdClaude can call the tool again if it needs the data
Just-in-time retrievalKeep IDs, paths or queries in context and load the content with a tool when neededExtra tool calls
MemoryClaude writes notes to storage outside the window and reads them back laterNotes must be kept current
CompactionOlder turns are replaced by a summaryDetail in the summarised turns can be lost
Subagent isolationA subagent explores in its own context window and returns only a summaryThe main agent sees conclusions, not raw data

Tool result clearing is available on the Messages API as context editing, behind the context-management-2025-06-27 beta header. The clear_tool_uses_20250919 strategy clears older tool results once input passes your trigger, keeps the most recent results (keep), and never touches tools you list in exclude_tools. Clearing happens on the server before the prompt reaches Claude. Your application keeps the full, unedited history. A second strategy, clear_thinking_20251015, manages old thinking blocks. Clearing tool results invalidates cached prompt prefixes from the point it clears, so there is a trade-off with prompt caching.

Compaction summarises older turns. On the Messages API, server-side compaction (in beta) writes the summary for you and returns it as a block that takes the place of the turns it summarises. The summary keeps only what the summarising prompt asks for. A generic "summarise the conversation" tends to drop task state such as which files changed, which option was chosen at a decision point, and which error was hit and how it was fixed. The API lets you replace the default summarising prompt with your own instructions inside the compaction object; name every field and decision that must survive, plus the user's latest open request. The same applies if you write your own summariser. In Claude Code, /compact summarises the conversation and accepts focus instructions, such as /compact keep the failing test names. /clear starts a new conversation with empty context, which is the better choice when you switch to unrelated work.

Subagents run in their own context window. In Claude Code, a subagent starts with its own system prompt and the delegated task, not your conversation history. It does its exploration there and returns only a summary, so file contents and logs never reach the main conversation.

Clearing old tool results (Python)

import anthropic

client = anthropic.Anthropic()

def prune_search(raw: dict) -> dict:
    """Tool output pruning: send Claude only what it needs."""
    return {
        "total": raw["total"],
        "results": [
            {"id": r["id"], "title": r["title"], "snippet": r["snippet"][:300]}
            for r in raw["results"][:10]
        ],
    }

response = client.beta.messages.create(
    model=MODEL,
    max_tokens=4096,
    tools=TOOLS,
    messages=messages,                      # your full history; the API edits its own copy
    betas=["context-management-2025-06-27"],
    context_management={
        "edits": [
            {
                "type": "clear_tool_uses_20250919",
                "trigger": {"type": "input_tokens", "value": 60000},
                "keep": {"type": "tool_uses", "value": 3},
                "exclude_tools": ["get_case_facts"],   # never clear the confirmed facts
            }
        ]
    },
)
print(response.context_management.applied_edits)   # what was cleared, and how many tokens

prune_search runs in your tool handler before the result enters messages. Context editing then clears older results once the prompt passes 60,000 input tokens, keeping the last three, and never clears get_case_facts.

Stopping drift

Pruning and compaction reduce bloat, but both can remove a detail that later turns depend on. Protect those details explicitly:

  • Keep confirmed facts (IDs, amounts, decisions, constraints) in one structured block that you send on every turn outside the history being summarised, or return them from a tool that is excluded from clearing.
  • Give compaction instructions that name what must survive.
  • Put standing instructions in the system prompt, not in an early user turn that may be summarised away.

The CCAR-F guide covers the same problem from the architect side in 5.1 Context across long conversations.

Measure with production-sized tool output

Test fixtures are usually small. A tool that returns a short record in testing may return a record plus attachments and history in production, several times longer. A session that ran twenty turns cleanly in development can then fill its budget by turn eight.

  • Measure each tool's result on the largest real input you can find, not on fixtures.
  • Log input tokens per turn, and gate requests with the token counting endpoint before they would exceed your budget.
  • Read the symptom carefully. When tool selection gets worse after a fixed number of turns, the system prompt and early instructions are often being crowded out by old tool results. Check how full the window is before you rewrite tool descriptions. If the same tool was chosen correctly in earlier turns, the description is not the cause. Raising max_tokens will not help either: it limits the reply, not what Claude reads.

Choosing what survives between sessions

Memory scopeWhat carries overFitsWhat you give up
In-context onlyNothing after the session endsShort sessions that fit in the windowAll state at session end; cost grows every turn
External storageState saved to a database or file and loaded at the startThe same user or task across days, or several agent instancesRead and write code, plus a lookup on each start
Summarised memoryA summary injected into the next sessionLong conversations that would outgrow the windowAny detail the summary did not keep
StatelessNothing, by designJobs that finish and close, such as formatting one fileAny follow-up that depends on an earlier job

Choose the scope at design time. Replaying every past session in context works in a demo and then fails a few sessions into real use, when the injected history alone fills most of the budget.

Which technique fits

SymptomChooseWhy
Tool results take up most of the windowPrune in the tool code, then clear old tool resultsRemoves the bulk without losing the conversation
A long task must continue and old turns matter only in outlineCompactionA summary keeps the thread in far fewer tokens
Exact values from early turns are wrong after compactionA persistent facts block outside the summarised historySummaries lose detail; the block is resent unchanged
Exploring a codebase floods the main contextA subagentExploration stays in its own window; only the summary returns
Work spans several sessionsMemory notes or a progress fileThe next session reads the notes instead of the old history
A Claude Code session is full of unrelated earlier work/clearStart fresh rather than summarise noise
Tool choice degrades after the same number of turns each sessionMeasure context per turn, then prune or compact earlierOld results are crowding out the instructions

Rules that decide exam answers

  • A bigger context window is not the fix. Context rot and cost still grow with every token.
  • Prune before you summarise. Clearing old tool results is the lightest change; compaction loses more.
  • Summaries lose detail. Facts that must stay exact belong in a block that is never summarised or cleared.
  • Subagents return summaries, not transcripts. That is what isolates the main context.
  • Context editing is server-side and breaks the cache where it clears. Your client keeps the full history; weigh clearing against the savings from prompt caching.
  • Never strip thinking blocks yourself to save space. In a tool loop they must go back unchanged; use the API's thinking-clearing strategy instead.

Where it appears in the exam

Context Engineering is 3.8% of the exam, inside Prompt and Context Engineering (11.0%). The guide's description points to questions about long-running agents, long conversations and coding sessions where context grows: rising cost per turn, a window that fills up, or an agent that forgets early instructions or facts.

Two sample questions

These are original Timo practice questions. They are not official exam questions.

Question 1

A research agent calls a web search tool about 40 times per session, and each result is around 3,000 tokens. The agent only needs the latest few results to decide its next step. Cost per turn keeps rising, and answers late in the session get worse. Which change fixes the cause with the least loss of useful context?

Answer: B. Old tool results are the bulk, and clearing them keeps the conversation while removing what Claude no longer uses. A leaves the bloat and the context rot in place, C summarises far more than needed and loses detail, and D still sends every result on every turn.

Question 2

A support agent compacts its conversation when it gets long. After compaction it sometimes quotes the wrong refund amount or order number, both of which the customer gave in the first turn. What should the developer do?

Answer: D. Exact values must not depend on a summary, so they live in a block that is resent unchanged. A trades drift for a hard failure at the limit, B still relies on the summary keeping exact values, and C adds a full-transcript read on every turn, which is the bloat compaction was meant to remove.

Build exercise

  1. Build an agent loop with a search tool that returns large results. Log usage.input_tokens on every turn.
  2. Add prune_search to the tool handler and compare the per-turn input.
  3. Add context editing with a low trigger, run a long session and print applied_edits to see what was cleared.
  4. In Claude Code, run /context, then move a "map this repository" task into a subagent and run /context again to compare.

Practise this topic

Sources