Structured errors for MCP tools: CCAR-F task statement 2.2
CCAR-F · Tool Design & MCP Integration (18% of the exam)
Task statement 2.2 sits in Tool Design & MCP Integration, 18% of the CCAR-F exam. It tests how a tool tells the agent what went wrong, so the agent can choose the right next step: retry, fix the input, explain a rule to the customer, or pass the problem up.
What the official guide covers
The Claude Certified Architect Foundations exam guide (version 1.0, effective July 2026) lists this under task statement 2.2, "Implement structured error responses for MCP tools":
| Knowledge of | Skills in |
|---|---|
The MCP isError flag for reporting tool failures back to the agent | Returning structured error metadata: an errorCategory (transient, validation, permission), an isRetryable boolean and a readable description |
| Four kinds of failure: transient (timeouts, service down), validation (bad input), business (policy violations) and permission | Marking business rule violations as not retryable, with a customer-friendly explanation the agent can pass on |
| Why a generic "Operation failed" stops the agent from choosing a sensible recovery | Recovering from transient failures inside a subagent, and passing up only what it cannot fix, with partial results and what it tried |
| Retryable vs non-retryable errors, and how metadata prevents wasted retries | Telling an access failure (needs a retry decision) apart from a valid empty result (the query worked and found nothing) |
How MCP reports a tool failure
The MCP specification has two ways to report errors from a tool call:
| Mechanism | Used for | What the model sees |
|---|---|---|
Protocol error (a JSON-RPC error object) | Unknown tool, a malformed request, a server fault | Clients may pass it to the model, but it is less likely to help recovery |
Tool execution error (a normal result with isError: true) | API failures, invalid input such as a date in the wrong format, business rule failures | Clients should pass it to the model so it can correct itself |
A tool execution error is an ordinary result: a content array plus isError: true. The specification does not define fields such as errorCategory or isRetryable. Those are a convention you put inside the content, and the exam expects you to use them. The flag tells the agent that the call failed; the metadata tells it what kind of failure it was.
A tool that returns structured errors (Python)
This in-process tool for the Claude Agent SDK returns each kind of failure differently. In Python the handler returns "is_error": True; the SDK sends it to Claude as a failed tool result.
import json
from typing import Any
from claude_agent_sdk import tool, create_sdk_mcp_server
def error_result(category: str, retryable: bool, message: str, **extra) -> dict[str, Any]:
payload = {"errorCategory": category, "isRetryable": retryable, "message": message, **extra}
return {"content": [{"type": "text", "text": json.dumps(payload)}], "is_error": True}
@tool(
"process_refund",
"Refunds all or part of one order. Input: order_id (ORD-123456) and amount in USD. "
"Returns the refund ID on success, or a JSON error with errorCategory and isRetryable.",
{"order_id": str, "amount": float},
)
async def process_refund(args: dict[str, Any]) -> dict[str, Any]:
try:
order = await billing.get_order(args["order_id"]) # your billing client
except TimeoutError:
return error_result("transient", True, "Billing service timed out. Retry once.")
except PermissionError:
return error_result("permission", False, "This agent cannot issue refunds for business accounts.")
if order is None:
return error_result("validation", False,
f"No order {args['order_id']}. Ask the customer to confirm the number.")
if args["amount"] > order.refundable_amount:
return error_result("business", False, "Amount is above the refundable balance.",
customerMessage=f"This order can be refunded up to ${order.refundable_amount:.2f}.")
refund = await billing.refund(order.id, args["amount"])
return {"content": [{"type": "text", "text": f"Refund {refund.id} issued."}]}
billing_server = create_sdk_mcp_server(name="billing", version="1.0.0", tools=[process_refund])
If the handler raised an uncaught exception instead, the SDK would still return an error result, but Claude would read only the raw exception text. Catching the error lets you write the message Claude acts on.
The same idea applies outside MCP. In a tool loop you write yourself on the Messages API, a failed tool goes back as a tool_result block with is_error: true and the explanation in its content. What you never do is drop the result or return an empty one: every tool_use block needs its matching tool_result, and an empty success reads to Claude as real data.
Three structural rules sit under this. Results are matched to calls by tool_use_id, not by position, so a wrong ID fails the next request with a validation error. All results for one assistant turn go back together in the next user message, with the tool_result blocks before any text. And a call your application declined (a person rejected it, or a policy check blocked it) still gets a result that says so; silence leaves the call unanswered. For a missing or malformed parameter, a clear is_error message is usually enough: Claude corrects the input and calls again.
Write the message for the next action
The message inside the error is a prompt. Claude reads it and decides what to do, so it should say what to change, not only what broke.
| Weak message | Message the agent can act on |
|---|---|
| "Invalid input" | "order_id must look like ORD-123456; received 552910. Ask the customer for the full order number." |
| "Forbidden" | "This agent cannot refund business accounts. Hand off to the billing team queue." |
| "Error 503" | "Billing service unavailable. Retry once after a short wait; if it fails again, tell the customer the refund is queued." |
Each one names the category of fix: change the input, route to someone else, or wait and retry.
What the agent should do with each category
| Category | Example | isRetryable | What the agent should do |
|---|---|---|---|
| Transient | Timeout, service unavailable | true | Retry, usually once or twice; then report with what it tried |
| Validation | Order number in the wrong format, missing field | false | Do not resend the same input; fix it or ask the user, then call again |
| Business | Refund above the allowed amount, policy says no | false | Do not retry; explain the rule using the customer-friendly message, or escalate |
| Permission | Agent lacks access to this account type | false | Do not retry; route to someone who has access |
The guide's point about wasted retries is here. Without a category, the agent cannot tell a timeout from a policy refusal, so it either retries everything (including refusals that will never succeed) or gives up on everything (including timeouts that would have worked a second later).
The test behind isRetryable is one question: would sending the same request again, a little later, plausibly work? A timeout passes that test. A bad order number, a policy refusal and a missing permission do not; they fail the same way until something changes. When you cannot tell, mark the error as not retryable and report it. A wrong "not retryable" surfaces quickly and gets fixed. A wrong "retryable" hammers the service and hides the real cause behind a wall of identical failures.
Decide where retries live
Retries can happen at two layers, and stacking them multiplies attempts.
- The API call layer. The official Anthropic SDKs already retry connection errors, rate limits and server errors with exponential backoff, honouring the
retry-afterheader when present. The client'smax_retriessetting controls this. - The tool layer. Your tool code or subagent decides whether to retry a failing downstream service. Keep it bounded (a fixed number of attempts with a growing, slightly randomised wait, using the service's own retry-after value when it gives one), never a tight loop that resends at once. Retry only the errors classed as transient; a validation or permission failure goes straight back to the agent.
If you wrap your own retry loop around an SDK call that is already retrying, one rate limit turns into many extra requests against the same limit. Pick one owner for each kind of failure.
Empty results are not errors
A query that runs and finds nothing is a success. Return it as a normal result, for example {"matches": [], "searched": "orders since 1 March"}. A query that never ran because the database was unreachable is a failure. Return it with isError: true and a transient category.
If both cases come back as an empty list, the agent will tell a customer "you have no orders" when the truth is "I could not check". That is the most common way this topic appears in questions.
Errors in multi-agent systems
In a coordinator and subagent design, the subagent handles what it can. It retries a transient failure locally. If that still fails, it returns a structured error to the coordinator with the failure type, what it attempted and any partial results. The coordinator then decides whether to retry with a different query, use another source or continue with a gap noted in the report. Hiding the failure (returning success with no data) or stopping the whole workflow both take that decision away from the coordinator.
Rules that decide exam answers
- Use
isError: truefor tool execution failures. Put the explanation in the result content so Claude can read it and react. - Categorise every error. Transient, validation, business and permission errors need different responses; "Operation failed" supports none of them.
- Business rule violations are not retryable. Mark them
isRetryable: falseand include a message the agent can give the customer. - Never return a failure as an empty success. Empty results and access failures must look different to the agent.
- Recover locally, escalate with context. Subagents retry transient errors themselves and pass up only what they cannot fix, with partial results and what was tried.
- Bounded retries, one owner. Retry only transient failures, with a cap and backoff, at one layer. Retrying a validation or business error, or retrying on top of the SDK's own retries, wastes calls.
Where it appears in the exam
Tool Design & MCP Integration is a primary domain in three of the six exam scenarios: Customer Support Resolution Agent, Multi-Agent Research System and Developer Productivity with Claude. Refund refusals and escalation fit the customer support scenario. Subagent failures and partial results fit the research scenario. The guide's first practice exercise asks you to add exactly these error fields to your tools.
Two sample questions
These are original Timo practice questions. They are not official exam questions.
Build exercise
- Build an Agent SDK tool with
@tooland theerror_resulthelper above. Simulate a timeout, a bad order number, a policy refusal and a permission failure. - Ask the agent to handle each case and note what it does: retry, ask for input, explain the rule or escalate.
- Replace the structured errors with a plain "Operation failed" and repeat. Count wasted retries and wrong messages to the customer.
- Make a search tool return an empty list during a simulated outage. Then return a transient error instead and compare what the agent tells the user.
Practise this topic
- Claude Certified Architect practice exam: free, 20 questions, no sign-up
- Claude Certified Architect hub
- CCAR-F study guide: all topics
- Same topic in another exam: Tool Implementation (CCDV-F)
- Previous topic: 2.1 Tool interface design
- Next topic: 2.3 Tool distribution and tool choice
Sources
- Claude Certified Architect Foundations Exam Guide, version 1.0, effective July 2026 (Anthropic), task statement 2.2
- Model Context Protocol: Tools (error handling)
- Claude Agent SDK documentation: Give Claude custom tools
- Anthropic documentation: Handle tool calls
- Anthropic documentation: Errors
By Amotion AI