TimoBy Amotion AI

Tool interface design: CCAR-F task statement 2.1

CCAR-F · Tool Design & MCP Integration (18% of the exam)

Task statement 2.1 sits in Tool Design & MCP Integration, 18% of the CCAR-F exam. It tests one skill: writing tool names and descriptions that let Claude pick the right tool every time, and spotting what breaks that choice when it goes wrong.

What the official guide covers

The Claude Certified Architect Foundations exam guide (version 1.0, effective July 2026) lists this under task statement 2.1, "Design effective tool interfaces with clear descriptions and boundaries":

Knowledge ofSkills in
Tool descriptions are the main thing Claude uses to choose a tool; thin descriptions make choices between similar tools unreliableWriting descriptions that separate each tool's purpose, inputs, outputs and when to use it instead of a similar tool
Good descriptions include input formats, example queries, edge cases and boundariesRenaming tools and rewriting descriptions to remove overlap, for example turning a vague analyze_content into a web-specific extract_web_results
Overlapping or near-identical descriptions send calls to the wrong toolSplitting a generic tool into purpose-specific tools, each with a defined input and output contract
Wording in the system prompt can tie keywords to the wrong toolReviewing system prompts for keyword-driven instructions that override good tool descriptions

How Claude chooses a tool

Every request carries a list of tool definitions. Each one carries three fields: name, description and input_schema. Claude reads all of them, compares them with the user's request and picks one. Your code does not route the call. The description is the only explanation Claude gets.

Anthropic's tool documentation calls detailed descriptions "by far the most important factor in tool performance" and asks for at least three to four sentences per tool. A good description says:

  1. What the tool does and what data it reaches.
  2. When to use it, and when not to (name the similar tool to use instead). If the call is only safe after another step, say so: "call only after get_record has confirmed the current value".
  3. What each parameter means and what format it expects.
  4. What the tool returns, and what it does not return.

Parameter descriptions matter too. A parameter called user is ambiguous; user_id with a format example is not.

Parameters, examples and results are part of the interface

The description gets Claude to the right tool. The rest of the definition decides whether the call that follows is a good one.

  • Required means required. Put a field in required only when the call makes no sense without it. If an order lookup marks customer_email as required but the customer only gave an order number, Claude has to supply an email it does not have. Leave such fields optional and say in the description what happens when they are missing.
  • Show the format when it matters. Tool definitions accept an optional input_examples array of sample inputs. Each example must validate against the input_schema, or the request is rejected. Use it for date formats, ID patterns and nested objects, where a sentence of description is easy to misread. It does not apply to Anthropic's server tools such as web search.
  • Enforce the shape when a bad argument is costly. Setting strict: true on a tool definition makes the API guarantee that Claude's input matches the input_schema, so missing fields and wrong types never reach your code. It checks shape, not meaning: a well-formed but wrong order number still gets through.
  • Same input shape, sharper boundary. When two tools both take a single query string, the schema gives Claude nothing to tell them apart. The description is the only signal left, so both tools need a sentence saying when not to use them.
  • Return what the next step needs. A tool's output is also part of its contract. Return meaningful names and stable identifiers rather than internal codes, and only the fields Claude needs to decide what to do next. Bloated results use context and bury the answer.

A thin definition and a fixed one (Python)

Two tools that both say "Searches for information" give Claude nothing to choose between. Here is the fixed version:

TOOLS = [
    {
        "name": "search_policy_docs",
        "description": (
            "Searches the internal policy library (refunds, returns, warranties, shipping). "
            "Use it when the customer asks what the company allows or requires. "
            "Do not use it for the status of a specific order; use get_order_status for that. "
            "Input is a short natural-language question, for example 'refund window for opened items'. "
            "Returns up to five policy passages, each with its document title and last-updated date. "
            "Returns nothing about individual customers or orders."
        ),
        "input_schema": {
            "type": "object",
            "properties": {
                "question": {"type": "string", "description": "Plain-language policy question"}
            },
            "required": ["question"],
        },
    },
    {
        "name": "get_order_status",
        "description": (
            "Looks up one order in the order system by its order number. "
            "Use it when the customer asks where an order is, when it ships or whether it was delivered. "
            "Input is the order number in the format ORD-123456; ask the customer for it if it is missing. "
            "Returns status, carrier, tracking number and expected delivery date. "
            "Does not return refund rules; use search_policy_docs for those."
        ),
        "input_schema": {
            "type": "object",
            "properties": {
                "order_id": {"type": "string", "description": "Order number, e.g. ORD-123456"}
            },
            "required": ["order_id"],
        },
    },
]

response = client.messages.create(
    model=MODEL, max_tokens=1024, tools=TOOLS,
    messages=[{"role": "user", "content": "Where is order ORD-552910?"}],
)

Each description names the other tool and says when to switch. That boundary sentence is what stops misrouting.

When to rename, split or merge

Overlap has three fixes, and the exam expects you to match the fix to the cause.

SituationChooseWhy
Two tools have near-identical descriptions and Claude mixes them upRewrite both descriptions with boundaries; rename if the names also overlapClaude can only choose between tools it can tell apart
A name hides what the tool really does (analyze_content only processes web search results)Rename it to say its job (extract_web_results) and narrow the descriptionThe name is read alongside the description
One generic tool does several jobs with different outputs (analyze_document extracts, summarises and fact-checks)Split it into purpose-specific tools (extract_data_points, summarize_content, verify_claim_against_source)Each tool gets one input and output contract Claude can rely on
Several tools are steps of one workflow on the same resource (create, review, merge a pull request)Keep them together as one tool with an action parameter, as Anthropic's docs suggestFewer, clearer tools reduce selection mistakes
Tools from several services share a libraryPrefix names with the service (github_list_prs, jira_search)The prefix makes the source obvious
Two tools take identical parameters and keep getting swappedAdd a "do not use this when..." sentence to both descriptionsWith matching schemas, the description is the only signal
You keep adding sentences to keep two near-identical tools apart, and they still collideCombine them into a single tool that takes a type or source parameterSome pairs are one job pretending to be two

Splitting and merging are not opposites. The test is the same: one tool should have one clear contract. Split when one tool returns different kinds of output depending on how it is called. Merge when separate tools are steps on the same object and would otherwise look alike.

Purpose-specific does not mean narrow. Claude Code works well with general tools such as Read, Grep and Bash because each has one clear contract and Claude can combine them for jobs nobody wrote a tool for. The thing to avoid is a tool whose output changes shape depending on how it is called, not a tool that is broadly useful.

Boundary sentences that point to earlier turns ("only use this if search_kb already ran for this question") work only if your application sends the full conversation history on each request. If your code trims old turns, Claude cannot check the condition.

Check that it is a description problem

Not every wrong call comes from the definitions. Two patterns point elsewhere:

  • Right for several turns, then wrong. The definitions did not change; the context did. As large tool results pile up, accuracy drops (Anthropic calls this context rot). Trim or clear old tool output instead of rewriting descriptions (see 5.1 Conversation context).
  • A correct pick, then an API validation error. A tool_result whose tool_use_id does not match the call is a bug in your loop's message handling, not in routing.

Rewrite descriptions when the misrouting shows up from the first turn and the two tools read alike.

Check the system prompt too

A well-written description can still lose to the system prompt. An instruction such as "Whenever the customer mentions a policy, search the knowledge base first" ties the word "policy" to one tool. When a customer says "my policy number is 4471, where is my order?", Claude may follow the keyword and call the wrong tool.

Review the system prompt for instructions keyed on words rather than on the task. Rewrite them to describe the goal ("answer questions about company rules from the policy library") and let the tool descriptions do the routing.

Rules that decide exam answers

  • Fix the description first, once context is ruled out. When Claude picks the wrong tool from the first turn and the descriptions are short, rewrite them before trying few-shot examples, routing layers or classifiers. Wrong picks that start only after many turns point to a crowded context.
  • State the boundary, not just the purpose. "Use this for X; for Y use other_tool" is what separates two similar tools.
  • Rename when the name misleads. A rewritten description under a misleading name still invites the wrong call.
  • Split tools that do several jobs. A generic tool with mixed outputs is the cause, not Claude's judgement.
  • Look for keyword traps in the system prompt. If descriptions are good and misrouting continues, an instruction tied to a keyword is the likely cause.
  • Required only when the call is meaningless without it. Over-required schemas push Claude to fill fields it has no value for; optional fields let it leave them out.

Where it appears in the exam

Tool Design & MCP Integration is a primary domain in three of the six exam scenarios: Customer Support Resolution Agent, Multi-Agent Research System and Developer Productivity with Claude. The guide's practice exercises on building a multi-tool agent and a multi-agent research pipeline also reinforce this domain, including writing descriptions for tools with similar functions.

Two sample questions

These are original Timo practice questions. They are not official exam questions.

Question 1

A support agent has two tools with clear, distinct descriptions: search_policy_docs and get_order_status. The system prompt says "If the customer uses the word policy, always call search_policy_docs first." Customers with insurance add-ons often write "my policy number is ..." when asking about deliveries, and in those cases the agent searches the policy library instead of checking the order. What should the architect change?

Answer: D. The descriptions are already good; the keyword-based instruction overrides them, so removing it fixes the cause. A patches a few phrasings and leaves the trap in place. B hides two different jobs behind one vague tool. C makes tool selection worse, not better.

Question 2

A research system has two tools named search and lookup. Both descriptions read "Finds relevant information for the query." search queries the public web; lookup queries the company wiki. Claude sends questions about internal processes to the web and gets outdated public answers. What is the best fix?

Answer: A. The names and descriptions overlap completely, so Claude has no basis for choosing; clear names and boundaries fix that. B removes a tool the system needs for public research. C biases every query toward the wiki, including public questions. D changes who holds the tools, not how they are described.

Build exercise

  1. Define two tools with one-line, near-identical descriptions. Send 20 mixed questions through the Messages API and log which tool Claude calls for each.
  2. Rewrite both descriptions with purpose, inputs, outputs and a boundary sentence naming the other tool. Run the same 20 questions and compare the log.
  3. Build one generic tool with a mode parameter (extract, summarise, verify). Split it into three purpose-specific tools and note how the calls change.
  4. Add a keyword-based line to the system prompt that conflicts with a description. Find a question that triggers the wrong tool, then reword the line and test again.

Practise this topic

Sources