TimoBy Amotion AI

Escalation and ambiguity: CCAR-F task statement 5.2

CCAR-F · Context Management & Reliability (15% of the exam)

Task statement 5.2 sits in Context Management & Reliability, 15% of the CCAR-F exam. It tests when an agent should hand a case to a person, when it should resolve the case itself, and what it should do when it cannot tell which customer or which request it is dealing with.

What the official guide covers

The Claude Certified Architect Foundations exam guide (version 1.0, effective July 2026) lists this under task statement 5.2, "Design effective escalation and ambiguity resolution patterns":

Knowledge ofSkills in
The right escalation triggers: the customer asks for a human, a policy exception or gap (not just a complex case), or no meaningful progressWriting explicit escalation criteria, with few-shot examples, into the system prompt
Escalating at once when a customer demands a person, versus offering to resolve a straightforward issueHonouring a request for a human immediately, without investigating first
Sentiment and self-reported confidence scores are poor proxies for how complex a case really isAcknowledging frustration and offering a fix when the issue is within the agent's ability, escalating if the customer asks again
Multiple customer matches call for more identifiers, not a heuristic choiceEscalating when policy is unclear or silent on the request, such as a competitor price match when policy covers only own-site adjustments
Telling the agent to ask for more identifiers when a lookup returns several matches

The three triggers

TriggerExampleWhat the agent does
The customer asks for a person"Put me through to someone."Escalates now. No lookups first, no attempt to talk them out of it.
Policy exception or gapA competitor price match when policy covers only price changes on your own siteEscalates. It does not stretch the policy to fit.
No meaningful progressTwo different approaches have failed, or the fix needs a system the agent cannot reachEscalates with a summary of what was tried.

The third trigger matches Anthropic's general advice on agents: they should pause for human feedback at checkpoints or when they hit a blocker, rather than keep trying.

Things that are not triggers on their own: a long case that policy fully covers, a customer who is upset, or a low confidence number from the model. The guide's sample material makes the same point from the other side: an agent that escalates simple cases and tries to handle policy exceptions itself has its boundaries the wrong way round.

Upset is not the same as "asked for a human"

Two messages look alike but need different handling:

  • "This is the third damaged parcel. I'm fed up." The issue (a damage replacement with photos) is within the agent's ability. Acknowledge the frustration, offer the replacement, and escalate only if the customer then asks for a person.
  • "I want to talk to a real person." This is an explicit request. Escalate immediately, even if one lookup would solve it.

Why sentiment and confidence scores mislead

Sentiment measures tone, not difficulty. A calm customer may need a policy exception; an angry one may have a two-minute fix. A confidence score the model reports about itself has the same flaw: it is not calibrated against outcomes, so an agent can be confidently wrong on hard cases. Clear criteria in the prompt fix the real cause, which is that the agent does not know where the boundary is.

Write the criteria into the system prompt, with examples

Anthropic's prompting guidance recommends a small set of examples (three to five) that are relevant and diverse, wrapped in <example> tags so Claude can tell them apart from instructions. Explaining the reason behind a rule also helps Claude apply it to cases the examples do not cover.

<escalation_policy>
Call escalate_to_human when any of these is true:
1. The customer asks for a person, a manager or a human agent. Escalate at once.
   Do not look anything up first.
2. The request needs something our refund and returns policy does not cover,
   or covers only for a different situation. Do not stretch the policy to fit.
3. You have tried two different approaches and the issue is still not resolved.

Do not escalate only because the customer is upset or the case has many steps.
If the customer is upset but has not asked for a person, acknowledge it and
offer the fix. Escalate if they ask for a person after that.

If get_customer returns more than one match, ask for one more identifier
(email address, order number or postcode). Never choose between matches yourself.
Why: a wrong match means we refund or disclose data on the wrong account.
</escalation_policy>

<examples>
<example>
Customer: "I've waited a week. Just put me through to someone."
Action: escalate_to_human(reason="customer_request"). No lookups first.
</example>
<example>
Customer: "Another shop sells this for $20 less. Match it."
Action: Policy covers price changes on our own site only. Escalate with
reason="policy_gap", including the competitor price the customer gave.
</example>
<example>
Customer: "Third damaged parcel this year. I'm furious. Photos attached."
Action: Damage replacement with photos is covered. Apologise and arrange the
replacement. Escalate only if the customer asks for a person.
</example>
<example>
get_customer("Sam Lee") returns 3 accounts.
Action: "I found more than one account under that name. Could you give me the
email address on the account or an order number?"
</example>
</examples>

For hard numeric limits, such as refunds above a set amount, use a hook in code rather than a prompt rule; see 1.5 Agent SDK hooks.

Make the handoff something a person can act on

An escalation is only as good as what reaches the person who picks it up. A reviewer needs three things: the inputs the agent worked from, what the agent tried or proposed, and why the case was flagged. Without the reason, every handoff looks the same and gets the same quick glance. Without the inputs, the person has to reopen the transcript or ask the customer to repeat themselves.

Put those fields in the tool schema, so the agent cannot call it without them:

{
  "name": "escalate_to_human",
  "description": "Hand the conversation to a human agent. Call this when the customer asks for a person, when policy does not cover the request, or after two different approaches have failed. Do not call it only because the customer is upset.",
  "input_schema": {
    "type": "object",
    "properties": {
      "reason": {
        "type": "string",
        "enum": ["customer_request", "policy_gap", "no_progress"]
      },
      "customer_id": {"type": "string", "description": "Verified ID, or 'unverified'"},
      "order_ids": {"type": "array", "items": {"type": "string"}},
      "customer_wants": {"type": "string", "description": "The request in the customer's own words"},
      "attempts": {"type": "array", "items": {"type": "string"}, "description": "What was tried and the result of each"},
      "summary": {"type": "string", "description": "Two or three sentences a person can act on"}
    },
    "required": ["reason", "customer_id", "customer_wants", "summary"]
  }
}

The enum keeps reasons countable, so you can report how many escalations were real requests, policy gaps or stalls. The attempts list is also how the agent knows it has reached "two approaches failed": record each attempt in the case state as it happens, rather than asking the model to remember. For a customer_request escalation the list is usually empty, which is correct: the agent should not have tried anything first.

Escalating too much is also a failure

Sending every unclear case to a person feels safe, but it has two costs. First contact resolution falls, and the human queue grows until reviewers stop reading each item and approve or close on habit. Oversight that covers everything ends up checking nothing. The three triggers exist to keep the queue small enough that each escalation gets real attention. The answer to "the agent gets some cases wrong" is sharper criteria and better examples, not a lower bar for handing off.

Several matching customers: ask, do not guess

When get_customer returns three accounts for one name, the agent must not pick the most recent, the one with the most orders or the one nearest the delivery address. Each of those is a guess that sometimes refunds or discloses data on the wrong account. Asking for one more identifier costs the customer a single reply and removes the guess.

Which response fits

SituationChooseWhy
The customer explicitly asks for a personEscalate immediatelyThe guide says honour it without investigating first
The customer is frustrated; the issue is within policyAcknowledge, offer the fix, escalate if they ask againResolves the case while respecting a repeated request
Policy is silent or unclear on the requestEscalate with a policy gap reasonThe agent should not invent policy
Two approaches failed with no progressEscalate with what was triedMore attempts waste the customer's time
A lookup returns several customersAsk for another identifierA heuristic choice can act on the wrong account
Agent escalates easy cases and handles exceptions itselfExplicit criteria plus few-shot examplesFixes unclear boundaries, the actual cause
Human agents complain they must reread every transcriptRequired handoff fields: reason, inputs, attempts, summaryThe person sees why the case reached them

Rules that decide exam answers

  • An explicit request for a person wins. Escalate at once. Options that investigate first or try to persuade the customer to stay are wrong.
  • Escalate on policy gaps, not on complexity. A long case that policy covers stays with the agent. A short request that policy does not cover goes to a person.
  • Sentiment and self-rated confidence are distractors. Neither measures whether the case needs a human.
  • Clarify, never guess, on multiple matches. Ask for another identifier before any action on an account.
  • Fix boundaries in the prompt first. Explicit criteria with examples come before classifiers or new infrastructure.
  • Escalate with context, and not by default. A handoff carries the reason, the inputs and what was tried. "Escalate everything uncertain" is a wrong answer, because a flooded queue gets rubber-stamped.

Where it appears in the exam

Context Management & Reliability is a primary domain in four of the six exam scenarios: Customer Support Resolution Agent, Code Generation with Claude Code, Multi-Agent Research System and Structured Data Extraction. Escalation questions fit the customer support scenario most closely, where the agent has an escalate_to_human tool and a target of high first-contact resolution while knowing when to escalate. The guide's preparation exercise 1 (a multi-tool agent with escalation logic) reinforces Domain 5.

Two sample questions

These are original Timo practice questions. They are not official exam questions.

Question 1

A customer's first message about a late delivery reads: "I want to speak to a real person now." The agent could probably solve the problem with one lookup_order call. What should the agent do?

Answer: B. An explicit request for a person is honoured immediately, without investigating first. A delays the handoff the customer asked for, C overrides an explicit request, and D uses tone, which does not decide whether a person is needed.

Question 2

get_customer sometimes returns two or three accounts for one name. The agent currently picks the account with the most recent order, and last week it refunded the wrong customer. What should the team change?

Answer: D. Several matches need clarification, not a heuristic. A is another guess, B relies on a self-reported score that is not calibrated, and C sends a person a case the customer could settle in one reply.

Build exercise

  1. Give a support agent four tools (get_customer, lookup_order, process_refund, escalate_to_human) and the policy and examples above in its system prompt.
  2. Write 12 test messages: three explicit requests for a person, three frustrated but simple cases, three policy gaps and three names with multiple matches. Record the agent's first action for each.
  3. Remove the <examples> block and run the 12 messages again. Count how many first actions change.
  4. Make escalate_to_human require reason, customer_id and summary. Read five handoffs and check that a person could act on each without opening the transcript.

Practise this topic

Sources