TimoBy Amotion AI

Structured errors for MCP tools: CCAR-F task statement 2.2

CCAR-F · Tool Design & MCP Integration (18% of the exam)

Task statement 2.2 sits in Tool Design & MCP Integration, 18% of the CCAR-F exam. It tests how a tool tells the agent what went wrong, so the agent can choose the right next step: retry, fix the input, explain a rule to the customer, or pass the problem up.

What the official guide covers

The Claude Certified Architect Foundations exam guide (version 1.0, effective July 2026) lists this under task statement 2.2, "Implement structured error responses for MCP tools":

Knowledge ofSkills in
The MCP isError flag for reporting tool failures back to the agentReturning structured error metadata: an errorCategory (transient, validation, permission), an isRetryable boolean and a readable description
Four kinds of failure: transient (timeouts, service down), validation (bad input), business (policy violations) and permissionMarking business rule violations as not retryable, with a customer-friendly explanation the agent can pass on
Why a generic "Operation failed" stops the agent from choosing a sensible recoveryRecovering from transient failures inside a subagent, and passing up only what it cannot fix, with partial results and what it tried
Retryable vs non-retryable errors, and how metadata prevents wasted retriesTelling an access failure (needs a retry decision) apart from a valid empty result (the query worked and found nothing)

How MCP reports a tool failure

The MCP specification has two ways to report errors from a tool call:

MechanismUsed forWhat the model sees
Protocol error (a JSON-RPC error object)Unknown tool, a malformed request, a server faultClients may pass it to the model, but it is less likely to help recovery
Tool execution error (a normal result with isError: true)API failures, invalid input such as a date in the wrong format, business rule failuresClients should pass it to the model so it can correct itself

A tool execution error is an ordinary result: a content array plus isError: true. The specification does not define fields such as errorCategory or isRetryable. Those are a convention you put inside the content, and the exam expects you to use them. The flag tells the agent that the call failed; the metadata tells it what kind of failure it was.

A tool that returns structured errors (Python)

This in-process tool for the Claude Agent SDK returns each kind of failure differently. In Python the handler returns "is_error": True; the SDK sends it to Claude as a failed tool result.

import json
from typing import Any
from claude_agent_sdk import tool, create_sdk_mcp_server

def error_result(category: str, retryable: bool, message: str, **extra) -> dict[str, Any]:
    payload = {"errorCategory": category, "isRetryable": retryable, "message": message, **extra}
    return {"content": [{"type": "text", "text": json.dumps(payload)}], "is_error": True}

@tool(
    "process_refund",
    "Refunds all or part of one order. Input: order_id (ORD-123456) and amount in USD. "
    "Returns the refund ID on success, or a JSON error with errorCategory and isRetryable.",
    {"order_id": str, "amount": float},
)
async def process_refund(args: dict[str, Any]) -> dict[str, Any]:
    try:
        order = await billing.get_order(args["order_id"])       # your billing client
    except TimeoutError:
        return error_result("transient", True, "Billing service timed out. Retry once.")
    except PermissionError:
        return error_result("permission", False, "This agent cannot issue refunds for business accounts.")
    if order is None:
        return error_result("validation", False,
                            f"No order {args['order_id']}. Ask the customer to confirm the number.")
    if args["amount"] > order.refundable_amount:
        return error_result("business", False, "Amount is above the refundable balance.",
                            customerMessage=f"This order can be refunded up to ${order.refundable_amount:.2f}.")
    refund = await billing.refund(order.id, args["amount"])
    return {"content": [{"type": "text", "text": f"Refund {refund.id} issued."}]}

billing_server = create_sdk_mcp_server(name="billing", version="1.0.0", tools=[process_refund])

If the handler raised an uncaught exception instead, the SDK would still return an error result, but Claude would read only the raw exception text. Catching the error lets you write the message Claude acts on.

The same idea applies outside MCP. In a tool loop you write yourself on the Messages API, a failed tool goes back as a tool_result block with is_error: true and the explanation in its content. What you never do is drop the result or return an empty one: every tool_use block needs its matching tool_result, and an empty success reads to Claude as real data.

Three structural rules sit under this. Results are matched to calls by tool_use_id, not by position, so a wrong ID fails the next request with a validation error. All results for one assistant turn go back together in the next user message, with the tool_result blocks before any text. And a call your application declined (a person rejected it, or a policy check blocked it) still gets a result that says so; silence leaves the call unanswered. For a missing or malformed parameter, a clear is_error message is usually enough: Claude corrects the input and calls again.

Write the message for the next action

The message inside the error is a prompt. Claude reads it and decides what to do, so it should say what to change, not only what broke.

Weak messageMessage the agent can act on
"Invalid input""order_id must look like ORD-123456; received 552910. Ask the customer for the full order number."
"Forbidden""This agent cannot refund business accounts. Hand off to the billing team queue."
"Error 503""Billing service unavailable. Retry once after a short wait; if it fails again, tell the customer the refund is queued."

Each one names the category of fix: change the input, route to someone else, or wait and retry.

What the agent should do with each category

CategoryExampleisRetryableWhat the agent should do
TransientTimeout, service unavailabletrueRetry, usually once or twice; then report with what it tried
ValidationOrder number in the wrong format, missing fieldfalseDo not resend the same input; fix it or ask the user, then call again
BusinessRefund above the allowed amount, policy says nofalseDo not retry; explain the rule using the customer-friendly message, or escalate
PermissionAgent lacks access to this account typefalseDo not retry; route to someone who has access

The guide's point about wasted retries is here. Without a category, the agent cannot tell a timeout from a policy refusal, so it either retries everything (including refusals that will never succeed) or gives up on everything (including timeouts that would have worked a second later).

The test behind isRetryable is one question: would sending the same request again, a little later, plausibly work? A timeout passes that test. A bad order number, a policy refusal and a missing permission do not; they fail the same way until something changes. When you cannot tell, mark the error as not retryable and report it. A wrong "not retryable" surfaces quickly and gets fixed. A wrong "retryable" hammers the service and hides the real cause behind a wall of identical failures.

Decide where retries live

Retries can happen at two layers, and stacking them multiplies attempts.

  • The API call layer. The official Anthropic SDKs already retry connection errors, rate limits and server errors with exponential backoff, honouring the retry-after header when present. The client's max_retries setting controls this.
  • The tool layer. Your tool code or subagent decides whether to retry a failing downstream service. Keep it bounded (a fixed number of attempts with a growing, slightly randomised wait, using the service's own retry-after value when it gives one), never a tight loop that resends at once. Retry only the errors classed as transient; a validation or permission failure goes straight back to the agent.

If you wrap your own retry loop around an SDK call that is already retrying, one rate limit turns into many extra requests against the same limit. Pick one owner for each kind of failure.

Empty results are not errors

A query that runs and finds nothing is a success. Return it as a normal result, for example {"matches": [], "searched": "orders since 1 March"}. A query that never ran because the database was unreachable is a failure. Return it with isError: true and a transient category.

If both cases come back as an empty list, the agent will tell a customer "you have no orders" when the truth is "I could not check". That is the most common way this topic appears in questions.

Errors in multi-agent systems

In a coordinator and subagent design, the subagent handles what it can. It retries a transient failure locally. If that still fails, it returns a structured error to the coordinator with the failure type, what it attempted and any partial results. The coordinator then decides whether to retry with a different query, use another source or continue with a gap noted in the report. Hiding the failure (returning success with no data) or stopping the whole workflow both take that decision away from the coordinator.

Rules that decide exam answers

  • Use isError: true for tool execution failures. Put the explanation in the result content so Claude can read it and react.
  • Categorise every error. Transient, validation, business and permission errors need different responses; "Operation failed" supports none of them.
  • Business rule violations are not retryable. Mark them isRetryable: false and include a message the agent can give the customer.
  • Never return a failure as an empty success. Empty results and access failures must look different to the agent.
  • Recover locally, escalate with context. Subagents retry transient errors themselves and pass up only what they cannot fix, with partial results and what was tried.
  • Bounded retries, one owner. Retry only transient failures, with a cap and backoff, at one layer. Retrying a validation or business error, or retrying on top of the SDK's own retries, wastes calls.

Where it appears in the exam

Tool Design & MCP Integration is a primary domain in three of the six exam scenarios: Customer Support Resolution Agent, Multi-Agent Research System and Developer Productivity with Claude. Refund refusals and escalation fit the customer support scenario. Subagent failures and partial results fit the research scenario. The guide's first practice exercise asks you to add exactly these error fields to your tools.

Two sample questions

These are original Timo practice questions. They are not official exam questions.

Question 1

A support agent's MCP tools return isError: true with the text "Operation failed" for every problem. Logs show the agent calling process_refund three times in a row when a refund breaks company policy, then telling the customer the system is down. What should the team change?

Answer: B. The agent cannot tell a policy refusal from an outage because every error looks the same; categories and a retryable flag fix that. A also stops useful retries of real timeouts. C hides failures from the agent entirely. D slows down retries that should never happen.

Question 2

A research subagent's search_case_records tool returns an empty list both when no records match and when the records database cannot be reached. During an outage, the final report states that "no prior cases exist" for several clients. What change fixes the cause?

Answer: C. The tool must make access failures and valid empty results look different, so the agent can retry or report a gap. A makes every genuine "no records" answer doubtful. B stops the whole run when recovery or partial results were possible. D doubles the cost and still cannot tell an outage from no data.

Build exercise

  1. Build an Agent SDK tool with @tool and the error_result helper above. Simulate a timeout, a bad order number, a policy refusal and a permission failure.
  2. Ask the agent to handle each case and note what it does: retry, ask for input, explain the rule or escalate.
  3. Replace the structured errors with a plain "Operation failed" and repeat. Count wasted retries and wrong messages to the customer.
  4. Make a search tool return an empty list during a simulated outage. Then return a transient error instead and compare what the agent tells the user.

Practise this topic

Sources