Error propagation: CCAR-F task statement 5.3
CCAR-F · Context Management & Reliability (15% of the exam)
Task statement 5.3 sits in Context Management & Reliability, 15% of the CCAR-F exam. It tests how a failure in a tool or subagent travels back to the coordinator, so the coordinator can choose well: retry, try another route, carry on with partial results, or report a gap.
What the official guide covers
The Claude Certified Architect Foundations exam guide (version 1.0, effective July 2026) lists this under task statement 5.3, "Implement error propagation strategies across multi-agent systems":
| Knowledge of | Skills in |
|---|---|
| Structured error context (failure type, the query attempted, partial results, alternatives) lets the coordinator make good recovery decisions | Returning structured error context: failure type, what was attempted, partial results and possible alternatives |
| An access failure (a timeout that needs a retry decision) is different from a valid empty result (a successful query with no matches) | Reporting access failures and valid empty results differently, so the coordinator can act on each |
| Generic statuses such as "search unavailable" hide context the coordinator needs | Having subagents recover locally from transient failures and pass on only errors they cannot resolve, with what was tried and any partial results |
| Silently returning empty results as success, and ending the whole workflow on one failure, are both anti-patterns | Adding coverage annotations to synthesis output that show which findings are well supported and which areas have gaps |
How a failure reaches the coordinator
There are two layers.
Tool to agent. A tool that fails returns a tool_result with is_error: true. Anthropic's tool use docs ask for instructive messages: say what went wrong and what to try next ("Rate limit exceeded. Retry after 60 seconds."), not just "failed". Claude can then correct the input, retry or explain the problem.
Subagent to coordinator. A subagent runs in its own context. Its tool calls and results stay there, and only its final message returns to the coordinator. So whatever the coordinator needs to know about a failure must be in that final message, in a form it can act on. In Claude Code, a subagent ended by an API error does return something: a foreground subagent that already produced text returns that partial output with a note that it was cut off, and a background subagent is marked failed with its last output attached. Your own subagents should do the same for failures inside their task.
Four outcomes the coordinator must tell apart
| Outcome | Example | What the coordinator should do |
|---|---|---|
| Success | 12 relevant articles found | Use them |
| Valid empty result | The query ran; no papers match | Record "no evidence found" as a finding |
| Transient access failure, still failing after local retries | Timeout or rate limit, three attempts | Retry later, narrow the query or use another source |
| Permanent failure | Permission denied, invalid query | Do not retry as is; switch source or report a gap |
Retriable or terminal: decide before you retry
Every failure starts with one question: would sending the same request again, a little later, plausibly work? If yes, it is retriable. If not, it is terminal, and retrying only burns time and hides the real problem behind a run of identical errors. For calls to the Claude API itself, the status code tells you which:
| Response | Bucket | What to do |
|---|---|---|
| 429 rate limit | Retriable | Wait for the retry-after header, then retry. A 429 for a reached spend limit has no retry-after and keeps failing until access resumes |
| 529 overloaded, 500 internal error, 504 timeout | Retriable | Back off and retry a capped number of times |
| 400 invalid request, 401 authentication, 403 permission, 404 not found | Terminal | Fix the request or the credentials; report the error |
HTTP 200 with stop_reason: "refusal" | Not an error | Not caught by status checks. Handle it as its own outcome. Resending it unchanged usually gets another refusal: remove or rephrase the refused turn, or retry on a fallback model |
Know where retries already happen. The official SDKs retry connection errors, rate limits and server errors twice by default, with exponential backoff that honours retry-after, and max_retries changes that. If your subagent wraps the same call in its own retry loop, the attempts multiply. Pick one layer to own retries for each kind of failure.
A failure can also be partial: three of four sources answered. That is the most common case in a research system, and it is the one that the two anti-patterns handle worst.
Anti-patterns that hide a failure or overreact to it
- Suppressing the error. The subagent catches a timeout and returns
[]marked as success. The coordinator cannot tell this from "no matches", so the report states that no evidence exists. That is a false finding. - Ending everything. One subagent times out and a top-level handler stops the whole run. Four good result sets are thrown away.
- "Search unavailable". Better than silence, but the coordinator still does not know what was tried, whether retrying makes sense or what partial data exists.
- A check that fails open. A validation or screening step that errors and then lets everything through is error suppression one level up. Decide in advance whether each check blocks or passes when it cannot run, and log it either way.
Anthropic's write-up of its own multi-agent research system takes the middle path: tell the agent when a tool is failing and let it adapt, and resume from where the error happened instead of starting again.
A subagent that recovers locally and reports the rest
import time
TRANSIENT = (TimeoutError, ConnectionError)
def search_sources(query: str, sources: dict, max_attempts: int = 3) -> dict:
"""sources maps a name to a function that returns a list of results."""
found, failures = [], []
for name, search_fn in sources.items():
for attempt in range(1, max_attempts + 1):
try:
found.extend(search_fn(query)) # an empty list is a valid answer
break
except TRANSIENT as exc:
if attempt == max_attempts:
failures.append({"source": name, "failure_type": type(exc).__name__,
"retryable": True, "attempts": attempt})
else:
time.sleep(2 ** attempt) # local retry with backoff
except PermissionError:
failures.append({"source": name, "failure_type": "PermissionError",
"retryable": False, "attempts": attempt})
break # retrying will not help
if not failures:
status = "ok" if found else "empty"
else:
status = "partial" if len(failures) < len(sources) else "failed"
return {
"status": status, # ok | empty | partial | failed
"query": query,
"results": found, # partial results are kept, not discarded
"failures": failures,
"alternatives": ["narrow the date range", "search the internal archive"] if failures else [],
}
Transient failures are retried inside the subagent, so the coordinator only hears about what the subagent could not fix. An empty list from a source that answered is reported as empty, never as a failure. Partial results travel with the failure details.
The coordinator then receives something like this as the subagent's final message:
{
"status": "partial",
"query": "port congestion 2026",
"results": ["...14 items..."],
"failures": [{"source": "news_archive", "failure_type": "TimeoutError", "retryable": true, "attempts": 3}],
"alternatives": ["narrow the date range", "search the internal archive"]
}
Count what came back before you synthesise
A coordinator that sends work to 20 subagents must check that 20 results returned before it writes anything. A subagent that crashed or timed out may return nothing at all, and a synthesis step that only adds up what arrived will report "18 sections reviewed, 2 flagged" as if the job were complete. Give each dispatched unit an ID and a deadline. Before synthesis, compare the IDs sent with the IDs returned, then retry the missing units or list them as gaps. The deadline also stops one stuck subagent, such as one retrying against a rate limit with no backoff, from holding the whole run open.
Coverage annotations in the final report
The synthesis step should say what the evidence covers, so a reader does not mistake a gap for a negative finding:
Coverage
- Well supported: freight rates and port congestion (4 sources each).
- Limited: labour disputes (1 source). The news archive timed out after
3 attempts; no news coverage after March is included.
- Not covered: rail freight. No source in scope returned results (valid empty).
Subagent failures are recoverable; coordinator failures usually are not
The two levels fail differently. When one subagent fails, the coordinator can retry that unit, send it to another source, or note the gap and let the rest of the work continue. When the coordinator loses its state, the whole run fails and finished subagent work may be stranded. Design for that difference:
- Make each subagent's unit of work safe to run twice, so a retry cannot duplicate side effects.
- Checkpoint the coordinator's progress to disk so a crash resumes instead of restarting (see the manifest pattern in 5.4).
- Give every run one trace ID and pass it to each subagent, so logs from all agents can be put back together when you investigate a failure.
Which design fits
| Situation | Choose | Why |
|---|---|---|
| A search times out once | Retry inside the subagent with backoff | Transient; the coordinator does not need to know |
| It still fails after local retries | Return failure type, query, attempts, partial results, alternatives | The coordinator can pick a recovery route |
| A query succeeds with zero matches | Report empty, not an error | It is a real finding |
| One of five subagents fails permanently | Continue with four and annotate the gap | Keeps good work and stays honest |
| The coordinator gets "search unavailable" | Replace with structured error context | A generic status hides what it needs |
| A call returns 400 or 403 | Fail fast and report; no retry | The same request will fail the same way |
| Retries pile up against a rate limit | One layer owns retries, honouring retry-after | Nested retry loops multiply attempts |
| The report says "48 reviewed" but 50 units were sent | Reconcile units sent against results returned | A missing result is a silent failure, not a smaller job |
Rules that decide exam answers
- Never report failure as success. An empty result marked successful after a timeout is the most tested anti-pattern.
- Empty is not an error. A successful query with no matches is a valid result, and treating it as a failure causes pointless retries.
- Recover locally, propagate the rest. Transient errors are retried in the subagent. Only unresolved errors go up, with what was tried and partial results.
- One failure should not end the run. Continue with what succeeded and annotate the gap. Check that results returned match units sent, so a missing unit is reported, not ignored.
- Structure beats a generic message. The right option names the failure type, the attempted query, partial results and alternatives.
- Retry only what time can fix. Rate limits, overloads and timeouts are retriable; bad requests and permission errors are not.
Where it appears in the exam
Context Management & Reliability is a primary domain in four of the six exam scenarios: Customer Support Resolution Agent, Code Generation with Claude Code, Multi-Agent Research System and Structured Data Extraction. Error propagation fits the multi-agent research scenario most closely, where a coordinator depends on search and analysis subagents. The guide's preparation exercise 4 (a multi-agent research pipeline with a simulated subagent timeout) reinforces Domain 5.
Two sample questions
These are original Timo practice questions. They are not official exam questions.
Build exercise
- Write two fake search functions: one that returns results and one that raises
TimeoutErroron every call. Runsearch_sourcesand check the status ispartial. - Change the second function to return
[]and check the status isokorempty, not a failure. - Build a coordinator that receives three subagent reports, one with
status: failed. Have it write a report with a coverage section. - Replace the structured result with the string "search unavailable" and compare what the coordinator does.
Practise this topic
- Claude Certified Architect practice exam: free, 20 questions, no sign-up
- Claude Certified Architect hub
- CCAR-F study guide: all topics
- Worked example: Production Claude tool loop
- Previous topic: 5.2 Escalation and ambiguity
- Next topic: 5.4 Large codebase context
Sources
- Claude Certified Architect Foundations Exam Guide, version 1.0, effective July 2026 (Anthropic), task statement 5.3
- Anthropic documentation: Handle tool calls
- Anthropic documentation: API errors
- Claude Code documentation: Create custom subagents
- Anthropic: How we built our multi-agent research system
By Amotion AI