Context across long conversations: CCAR-F task statement 5.1
CCAR-F · Context Management & Reliability (15% of the exam)
Task statement 5.1 sits in Context Management & Reliability, 15% of the CCAR-F exam. It tests one skill: keeping the facts a long conversation depends on (amounts, dates, order numbers, what the customer was told) exact while the history grows, gets trimmed or gets summarised.
What the official guide covers
The Claude Certified Architect Foundations exam guide (version 1.0, effective July 2026) lists this under task statement 5.1, "Manage conversation context to preserve critical information across long interactions":
| Knowledge of | Skills in |
|---|---|
| Progressive summarisation turns numbers, percentages, dates and customer expectations into vague summaries | Extracting transactional facts (amounts, dates, order numbers, statuses) into a persistent "case facts" block sent with every prompt, outside the summarised history |
| Each API request must carry the complete conversation history to keep the conversation coherent | Keeping structured issue data (order IDs, amounts, statuses) in a separate context layer when one session covers several issues |
| Tool results pile up and use tokens out of proportion to their value, such as 40+ fields per order lookup when 5 matter | Trimming verbose tool output to the relevant fields before it accumulates |
| The "lost in the middle" effect: models read the start and end of long inputs reliably but may miss findings in the middle | Putting a key findings summary at the start of aggregated input and using explicit section headers |
| Making subagents return structured data with metadata (dates, sources, method) instead of verbose content and reasoning when the next agent has a small context budget |
How context grows
The Messages API does not remember earlier turns. Every request carries the whole conversation, and everything in the request counts toward the context window: the system prompt, every message (including tool results and documents) and the tool definitions. If you drop turns to save space, Claude loses them. If the input alone is larger than the context window, the API rejects the request with a "prompt is too long" error.
Three things cause trouble in a long support conversation:
- Tool results. An order lookup that returns 45 fields adds all 45 to the history, and they are re-sent with every later request.
- Several issues in one chat. Facts from three orders mix in one stream of text.
- Length itself. In a long input, details in the middle are the most likely to be missed.
Ways to shorten history, and what each one loses
| Method | What it does | What can be lost |
|---|---|---|
| Your own summary of older turns | Replaces old turns with a short summary | "£84.20 refund for order 7731, delivered 12 March" becomes "customer wants a refund for a recent order" |
| Server-side compaction (Claude API) | The API replaces older turns with a summary it writes; you can supply your own summarisation instructions | The same, unless your instructions name what must survive |
| Context editing, tool result clearing | Once input passes a threshold, older tool results are replaced with placeholder text (clear_tool_uses_20250919) | Any value you did not extract before the result was cleared |
/compact in Claude Code | Summarises the session; /compact Focus on the API changes steers what is kept | The same; a "Compact instructions" section in CLAUDE.md sets a default |
Every method shortens history by rewriting or removing it. The fix is the same for all of them: copy the facts that matter into a place no summary touches, before the summary runs.
If you summarise, name what must survive
A summariser told only "summarise the conversation so far" writes a general recap. It keeps the topic and drops the details a later turn depends on. Whichever method you use (your own summary call, compaction with your own instructions, or /compact with a focus), write the instruction as a list of what must come through unchanged:
Summarise the conversation above for the agent that continues it.
Keep exactly as written: every order ID, amount, currency, date and status;
what the customer asked for and what we promised them;
what has already been tried and the result of each attempt;
the question that is still open.
Drop greetings, repeated apologies and tool output already reflected in case_facts.
This makes the summary better, but it is still a model rewriting text. Treat it as a backstop. The case facts block below stays the source of truth for numbers.
The case facts block
Keep a small structured record of facts and rebuild it into the prompt on every request. Summaries can then compress the chat freely, because the numbers live somewhere else.
import json
import os
import anthropic
client = anthropic.Anthropic()
MODEL = os.environ["CLAUDE_MODEL"] # a current model ID
RETURN_FIELDS = ["order_id", "status", "total", "delivered_on", "return_window_ends"]
case_facts = {"customer_id": None, "issues": {}} # one entry per order in this session
def trim_order(order: dict) -> dict:
"""Keep only the fields the returns flow needs."""
return {k: order.get(k) for k in RETURN_FIELDS}
def record_order(order: dict, customer_expects: str = "") -> dict:
slim = trim_order(order)
case_facts["issues"][slim["order_id"]] = {**slim, "customer_expects": customer_expects}
return slim # send this trimmed version back as the tool_result content
def build_system(base_prompt: str) -> str:
return (
f"{base_prompt}\n\n<case_facts>\n{json.dumps(case_facts, indent=2)}\n</case_facts>\n"
"Use case_facts as the source of truth for amounts, dates, order numbers and statuses."
)
def ask(base_prompt: str, history: list, tools: list):
return client.messages.create(
model=MODEL,
max_tokens=1024,
system=build_system(base_prompt),
tools=tools,
messages=history, # the full history, or a summary plus recent turns, on every call
)
trim_order removes the fields nobody needs before they enter the history. case_facts keeps each issue separate, keyed by order ID. Record what the customer expects ("full refund to card") as well as system values: vague summaries lose expectations first.
Fill the block from tool results, not from the model's reading of the chat. For values that change, such as an order's status, call the order system again rather than trusting an old copy. An order that shipped, came back to the depot and now waits for dispatch again is described correctly by neither an early summary nor a search over old order notes. Only the system of record knows its current state.
If the same customer comes back tomorrow, the conversation history is gone but the case facts need not be. Store the block per customer or per case in your own database and load it at the start of the next session. A summary of yesterday's chat is a weaker substitute: it carries only what the summariser chose to keep.
Measure the context before you blame the prompt
A full context rarely announces itself. The common symptom is that the agent starts picking the wrong tool, ignores an instruction it followed earlier, or quotes a stale figure, always after roughly the same number of turns. That looks like a prompt or tool description bug, so teams rewrite the schema and nothing changes. Often the cause is that accumulated tool output now outweighs the instructions.
Test fixtures hide this. A sample order record might be a few hundred tokens; a real one with notes and scan events can be several times larger, so the context fills at turn eight instead of turn forty. Check before you ship:
count = client.messages.count_tokens(
model=MODEL,
system=build_system(BASE_PROMPT),
tools=TOOLS,
messages=history_with_largest_real_tool_results,
)
print(count.input_tokens) # compare with your budget before sending
Token counting takes the same inputs as a Messages request and does not run the model. In production, log response.usage.input_tokens on every turn and alert when a conversation crosses your budget. If generation itself runs into the context window limit, current models stop with stop_reason set to model_context_window_exceeded, so handle that value instead of treating the reply as complete.
Long aggregated inputs: put key findings first
When a coordinator passes several subagent reports to a synthesis step, the guide's fix for position effects is structure:
- Open with a short key findings summary.
- Put each detailed result under an explicit section header.
- Wrap each document in its own tags with its source and date. Anthropic's prompting guidance recommends
<document>tags with<source>and other metadata, long material near the top, and the question or instructions at the end.
Upstream agents can make this easier. When the downstream agent has a small context budget, change the subagents to return structured key facts, citations and relevance scores, with dates and sources attached, instead of long prose and reasoning chains.
Which fix fits
| Situation | Choose | Why |
|---|---|---|
| Amounts and dates turn vague after several summaries | Case facts block outside the summarised history | It is rebuilt from data each turn, never rewritten by a summary |
| An order lookup returns 45 fields; the flow needs 5 | Trim the result before adding it to the conversation | Fewer tokens and less noise on every later request |
| One chat covers three orders | A separate issue record per order ID | Facts from different orders stay apart |
| Old tool results are no longer needed in a long agent run | Context editing, after extracting what you need | Frees space without losing extracted facts |
| A synthesis input runs to 30 pages | Key findings first, section headers, instructions last | Reduces missed findings in the middle |
| The downstream agent has a small context budget | Upstream agents return structured facts with metadata | The next agent gets data, not reasoning |
| Tool choice degrades after a fixed number of turns | Measure input tokens per turn, then trim or compact | Accumulated output is crowding out the instructions |
| The customer returns in a new session | Load the stored case facts at session start | History does not carry over; stored facts do |
Rules that decide exam answers
- Facts go in a block, not in the summary. If the question says numbers or dates drift after summarisation, the answer extracts them into a persistent case facts block. Asking the summariser to "keep all numbers" is still probabilistic.
- Trim tool output before it accumulates. Filtering 45 fields down to 5 at the point of entry beats telling Claude to ignore the extra fields.
- Send the history every time. The API keeps no state. An option that sends only the latest message breaks the conversation.
- A bigger context window is not the fix. It delays the problem and makes the middle of the input longer. Structure and extraction fix it. When quality drops after a set number of turns, measure tokens before rewriting prompts: correct tool choices in the early turns show the tool descriptions were not the cause.
- Position is a design choice. Key findings go at the start, details go under headers.
- Upstream agents return data, not reasoning. When the next agent is short of context, change what the previous agent sends.
Where it appears in the exam
Context Management & Reliability is a primary domain in four of the six exam scenarios: Customer Support Resolution Agent, Code Generation with Claude Code, Multi-Agent Research System and Structured Data Extraction. Case facts and tool output trimming fit the customer support scenario. Key findings placement and structured subagent output fit the research scenario. The guide's preparation exercises 1, 3 and 4 all list Domain 5 among the domains they reinforce.
Two sample questions
These are original Timo practice questions. They are not official exam questions.
Build exercise
- Build a support agent with a
lookup_ordertool that returns 40 fields. Logresponse.usage.input_tokenson every turn of a 20-turn chat about three different orders. - Add
trim_orderfrom the example above and run the same chat. Compare the token counts. - Summarise turns older than the last six with a plain "summarise this conversation" prompt. Check whether each order's amount and delivery date survive.
- Add the case facts block and repeat step 3. Ask a final question that needs the amount from the first order.
Practise this topic
- Claude Certified Architect practice exam: free, 20 questions, no sign-up
- Claude Certified Architect hub
- CCAR-F study guide: all topics
- Same topic in another exam: Context Engineering (CCDV-F)
- Previous topic: 4.6 Multi-pass review
- Next topic: 5.2 Escalation and ambiguity
Sources
- Claude Certified Architect Foundations Exam Guide, version 1.0, effective July 2026 (Anthropic), task statement 5.1
- Anthropic documentation: Context windows
- Anthropic documentation: Context editing
- Anthropic documentation: Prompting best practices
- Anthropic documentation: Token counting
By Amotion AI