TimoBy Amotion AI

Error propagation: CCAR-F task statement 5.3

CCAR-F · Context Management & Reliability (15% of the exam)

Task statement 5.3 sits in Context Management & Reliability, 15% of the CCAR-F exam. It tests how a failure in a tool or subagent travels back to the coordinator, so the coordinator can choose well: retry, try another route, carry on with partial results, or report a gap.

What the official guide covers

The Claude Certified Architect Foundations exam guide (version 1.0, effective July 2026) lists this under task statement 5.3, "Implement error propagation strategies across multi-agent systems":

Knowledge ofSkills in
Structured error context (failure type, the query attempted, partial results, alternatives) lets the coordinator make good recovery decisionsReturning structured error context: failure type, what was attempted, partial results and possible alternatives
An access failure (a timeout that needs a retry decision) is different from a valid empty result (a successful query with no matches)Reporting access failures and valid empty results differently, so the coordinator can act on each
Generic statuses such as "search unavailable" hide context the coordinator needsHaving subagents recover locally from transient failures and pass on only errors they cannot resolve, with what was tried and any partial results
Silently returning empty results as success, and ending the whole workflow on one failure, are both anti-patternsAdding coverage annotations to synthesis output that show which findings are well supported and which areas have gaps

How a failure reaches the coordinator

There are two layers.

Tool to agent. A tool that fails returns a tool_result with is_error: true. Anthropic's tool use docs ask for instructive messages: say what went wrong and what to try next ("Rate limit exceeded. Retry after 60 seconds."), not just "failed". Claude can then correct the input, retry or explain the problem.

Subagent to coordinator. A subagent runs in its own context. Its tool calls and results stay there, and only its final message returns to the coordinator. So whatever the coordinator needs to know about a failure must be in that final message, in a form it can act on. In Claude Code, a subagent ended by an API error does return something: a foreground subagent that already produced text returns that partial output with a note that it was cut off, and a background subagent is marked failed with its last output attached. Your own subagents should do the same for failures inside their task.

Four outcomes the coordinator must tell apart

OutcomeExampleWhat the coordinator should do
Success12 relevant articles foundUse them
Valid empty resultThe query ran; no papers matchRecord "no evidence found" as a finding
Transient access failure, still failing after local retriesTimeout or rate limit, three attemptsRetry later, narrow the query or use another source
Permanent failurePermission denied, invalid queryDo not retry as is; switch source or report a gap

Retriable or terminal: decide before you retry

Every failure starts with one question: would sending the same request again, a little later, plausibly work? If yes, it is retriable. If not, it is terminal, and retrying only burns time and hides the real problem behind a run of identical errors. For calls to the Claude API itself, the status code tells you which:

ResponseBucketWhat to do
429 rate limitRetriableWait for the retry-after header, then retry. A 429 for a reached spend limit has no retry-after and keeps failing until access resumes
529 overloaded, 500 internal error, 504 timeoutRetriableBack off and retry a capped number of times
400 invalid request, 401 authentication, 403 permission, 404 not foundTerminalFix the request or the credentials; report the error
HTTP 200 with stop_reason: "refusal"Not an errorNot caught by status checks. Handle it as its own outcome. Resending it unchanged usually gets another refusal: remove or rephrase the refused turn, or retry on a fallback model

Know where retries already happen. The official SDKs retry connection errors, rate limits and server errors twice by default, with exponential backoff that honours retry-after, and max_retries changes that. If your subagent wraps the same call in its own retry loop, the attempts multiply. Pick one layer to own retries for each kind of failure.

A failure can also be partial: three of four sources answered. That is the most common case in a research system, and it is the one that the two anti-patterns handle worst.

Anti-patterns that hide a failure or overreact to it

  • Suppressing the error. The subagent catches a timeout and returns [] marked as success. The coordinator cannot tell this from "no matches", so the report states that no evidence exists. That is a false finding.
  • Ending everything. One subagent times out and a top-level handler stops the whole run. Four good result sets are thrown away.
  • "Search unavailable". Better than silence, but the coordinator still does not know what was tried, whether retrying makes sense or what partial data exists.
  • A check that fails open. A validation or screening step that errors and then lets everything through is error suppression one level up. Decide in advance whether each check blocks or passes when it cannot run, and log it either way.

Anthropic's write-up of its own multi-agent research system takes the middle path: tell the agent when a tool is failing and let it adapt, and resume from where the error happened instead of starting again.

A subagent that recovers locally and reports the rest

import time

TRANSIENT = (TimeoutError, ConnectionError)

def search_sources(query: str, sources: dict, max_attempts: int = 3) -> dict:
    """sources maps a name to a function that returns a list of results."""
    found, failures = [], []
    for name, search_fn in sources.items():
        for attempt in range(1, max_attempts + 1):
            try:
                found.extend(search_fn(query))   # an empty list is a valid answer
                break
            except TRANSIENT as exc:
                if attempt == max_attempts:
                    failures.append({"source": name, "failure_type": type(exc).__name__,
                                     "retryable": True, "attempts": attempt})
                else:
                    time.sleep(2 ** attempt)     # local retry with backoff
            except PermissionError:
                failures.append({"source": name, "failure_type": "PermissionError",
                                 "retryable": False, "attempts": attempt})
                break                            # retrying will not help

    if not failures:
        status = "ok" if found else "empty"
    else:
        status = "partial" if len(failures) < len(sources) else "failed"
    return {
        "status": status,            # ok | empty | partial | failed
        "query": query,
        "results": found,            # partial results are kept, not discarded
        "failures": failures,
        "alternatives": ["narrow the date range", "search the internal archive"] if failures else [],
    }

Transient failures are retried inside the subagent, so the coordinator only hears about what the subagent could not fix. An empty list from a source that answered is reported as empty, never as a failure. Partial results travel with the failure details.

The coordinator then receives something like this as the subagent's final message:

{
  "status": "partial",
  "query": "port congestion 2026",
  "results": ["...14 items..."],
  "failures": [{"source": "news_archive", "failure_type": "TimeoutError", "retryable": true, "attempts": 3}],
  "alternatives": ["narrow the date range", "search the internal archive"]
}

Count what came back before you synthesise

A coordinator that sends work to 20 subagents must check that 20 results returned before it writes anything. A subagent that crashed or timed out may return nothing at all, and a synthesis step that only adds up what arrived will report "18 sections reviewed, 2 flagged" as if the job were complete. Give each dispatched unit an ID and a deadline. Before synthesis, compare the IDs sent with the IDs returned, then retry the missing units or list them as gaps. The deadline also stops one stuck subagent, such as one retrying against a rate limit with no backoff, from holding the whole run open.

Coverage annotations in the final report

The synthesis step should say what the evidence covers, so a reader does not mistake a gap for a negative finding:

Coverage
- Well supported: freight rates and port congestion (4 sources each).
- Limited: labour disputes (1 source). The news archive timed out after
  3 attempts; no news coverage after March is included.
- Not covered: rail freight. No source in scope returned results (valid empty).

Subagent failures are recoverable; coordinator failures usually are not

The two levels fail differently. When one subagent fails, the coordinator can retry that unit, send it to another source, or note the gap and let the rest of the work continue. When the coordinator loses its state, the whole run fails and finished subagent work may be stranded. Design for that difference:

  • Make each subagent's unit of work safe to run twice, so a retry cannot duplicate side effects.
  • Checkpoint the coordinator's progress to disk so a crash resumes instead of restarting (see the manifest pattern in 5.4).
  • Give every run one trace ID and pass it to each subagent, so logs from all agents can be put back together when you investigate a failure.

Which design fits

SituationChooseWhy
A search times out onceRetry inside the subagent with backoffTransient; the coordinator does not need to know
It still fails after local retriesReturn failure type, query, attempts, partial results, alternativesThe coordinator can pick a recovery route
A query succeeds with zero matchesReport empty, not an errorIt is a real finding
One of five subagents fails permanentlyContinue with four and annotate the gapKeeps good work and stays honest
The coordinator gets "search unavailable"Replace with structured error contextA generic status hides what it needs
A call returns 400 or 403Fail fast and report; no retryThe same request will fail the same way
Retries pile up against a rate limitOne layer owns retries, honouring retry-afterNested retry loops multiply attempts
The report says "48 reviewed" but 50 units were sentReconcile units sent against results returnedA missing result is a silent failure, not a smaller job

Rules that decide exam answers

  • Never report failure as success. An empty result marked successful after a timeout is the most tested anti-pattern.
  • Empty is not an error. A successful query with no matches is a valid result, and treating it as a failure causes pointless retries.
  • Recover locally, propagate the rest. Transient errors are retried in the subagent. Only unresolved errors go up, with what was tried and partial results.
  • One failure should not end the run. Continue with what succeeded and annotate the gap. Check that results returned match units sent, so a missing unit is reported, not ignored.
  • Structure beats a generic message. The right option names the failure type, the attempted query, partial results and alternatives.
  • Retry only what time can fix. Rate limits, overloads and timeouts are retriable; bad requests and permission errors are not.

Where it appears in the exam

Context Management & Reliability is a primary domain in four of the six exam scenarios: Customer Support Resolution Agent, Code Generation with Claude Code, Multi-Agent Research System and Structured Data Extraction. Error propagation fits the multi-agent research scenario most closely, where a coordinator depends on search and analysis subagents. The guide's preparation exercise 4 (a multi-agent research pipeline with a simulated subagent timeout) reinforces Domain 5.

Two sample questions

These are original Timo practice questions. They are not official exam questions.

Question 1

A research subagent searches a news archive for "port strikes 2026". The archive times out three times. The subagent catches the exception and returns an empty result list marked as successful. The final report says no strikes were reported. What is the root problem?

Answer: D. The coordinator could not tell a timeout from "no matches", so a gap became a false finding. A adds cost without fixing the reporting, B would make every genuine empty result suspect, and C misses the point: the subagent already retried, and the problem is how it reported the outcome.

Question 2

A coordinator runs five research subagents. The patent search subagent fails with a permission error that retrying will not fix. Today a top-level handler stops the whole run when any subagent fails. What should happen instead?

Answer: B. The coordinator keeps the good work and the report states the gap and its cause. A never ends for a permanent error, C hides the gap, and D repeats the anti-pattern of ending everything on one failure.

Build exercise

  1. Write two fake search functions: one that returns results and one that raises TimeoutError on every call. Run search_sources and check the status is partial.
  2. Change the second function to return [] and check the status is ok or empty, not a failure.
  3. Build a coordinator that receives three subagent reports, one with status: failed. Have it write a report with a coverage section.
  4. Replace the structured result with the string "search unavailable" and compare what the coordinator does.

Practise this topic

Sources