TimoBy Amotion AI

Evaluation, Testing and Optimization: CCAR-P domain 4 study guide

CCAR-P · Evaluation, Testing & Optimization (16% of the exam)

Domain 4 is Evaluation, Testing & Optimization, 16% of the CCAR-P exam. It tests whether you can prove a Claude system works, find the cause when it does not, and make it cheaper or faster without breaking it.

What the official guide covers

The Claude Certified Architect Professional exam guide (version 1.0, effective July 2026) lists six tasks under Domain 4, "Evaluation, Testing & Optimization":

What the guide listsWhat it means in practice
Define evaluation metrics (accuracy, latency, cost, safety, security)Set a measurable target for each dimension before building
Design evaluation datasets and test frameworks using mixed methodsBuild test sets from real cases and grade them with code, models and people
Conduct A/B testing and iterative improvementChange one thing, compare against a baseline, keep or roll back
Diagnose system issues (prompt failure, hallucinations, model mismatch)Tell the cause from the symptom and check the right component first
Optimise token usage, latency and cost-performanceCut tokens and time while holding quality on the test set
Monitor performance with logging and observability toolsKeep measuring in production, not only before launch

Defining metrics

Anthropic's guidance says success criteria should be specific, measurable, achievable and relevant, and most systems need several. Decide how to measure each dimension.

DimensionMeasured byExample target
AccuracyCode grader against labelled answers, or a model grader with a rubric98% of invoice fields exact
LatencyTime to first token and total time, at the 50th and 95th percentilep95 under 6 seconds
CostTokens per completed task, including cache reads and writes, plus review time30% fewer tokens than baseline
SafetyModel grader for policy breaches, calibrated against human labelsNo harmful outputs in the red-team set
SecurityPrompt injection and data leakage testsNo injected instruction followed

Write the criteria and the first test cases before production code. That forces "summarise claims accurately" into something testable (which fields, what pass rate, which failures count) while the design is still cheap to change, and it gives you the gate every later change must pass. Take thresholds from the business requirement, not from whatever the prototype scored.

Building datasets and graders

Build the dataset from what the system will really see:

  • Real traffic, labelled by the people who own the outcome.
  • Known failures from incidents and complaints, kept as permanent tests.
  • Edge cases and adversarial inputs, such as injected instructions.
  • Both directions: cases where Claude should act and where it should not.

Anthropic's guidance favours volume: many auto-graded cases beat a few hand-graded ones. Mix three graders:

GraderStrengthWeaknessUse for
CodeFast, exact, repeatableMisses valid variationsFields, formats, totals, whether an action happened
ModelScales to open-ended outputNeeds a clear rubric and calibrationTone, completeness, faithfulness to sources
HumanHighest judgementSlow and costlyCalibrating model graders; high-stakes samples

Climb this ladder from the top: use code wherever the check is unambiguous, even when the stakes are high (a leaked customer ID is a string match). One requirement can need two graders: code checks whether a forbidden action happened; a model grades whether a summary is faithful. Give model graders a detailed rubric and a small fixed set of verdict labels, let them reason before scoring, and consider a different model from the one being graded. Before trusting a model grader, compare its verdicts with human labels on a sample; an uncalibrated grader gives confident scores you cannot rely on.

Two dataset traps. A single-turn set says nothing about whether a chat system keeps facts straight across turns, so multi-turn behaviour needs its own set of whole conversations. And a test set that predates a prompt change keeps passing while measuring behaviour the system no longer has: update the cases whenever the prompt, retrieval or model changes.

For agents, Anthropic's evals article adds: grade the outcome, not the path; run several trials per task; and separate capability tests, expected to fail at first, from regression tests, which should almost always pass. Read transcripts to see whether a failure is the agent's or the grader's.

A/B testing and iteration

  1. Pin the model ID and prompt version for the baseline.
  2. Change one thing: the prompt, the model tier, the effort level, the retrieval settings.
  3. Run both versions on the same offline test set, with several trials per case.
  4. If the change passes every gate, send a small share of live traffic to it and compare the same metrics, plus sampled human review.
  5. Keep it, or roll back. Add any new failures to the test set.

A live test needs a written hypothesis that names the change, the primary metric, the minimum gain and the limits on other metrics (for example, "task success up at least 5 points with p95 latency unchanged"). Assign users or sessions at random and keep each one in the same group. Set the sample size and the primary metric before starting; LLM outputs vary more than ordinary software, so samples run larger, and 50 sessions per arm rarely separate a real gain from noise. Set the confidence you need from the cost of a wrong call, not from the effect you expect: a small prompt change to a medical summariser still needs a large sample. A statistically significant result still has to be large enough to justify maintaining the new version, and no secondary metric may get worse.

SituationTest patternTrade-off
Enough traffic, a bad output is tolerable and reversibleLive A/B splitReal user signals, but some users see the worse version
A single bad output is unacceptable, or traffic is lowShadow test: the new version gets a copy of live requests, users see only the current versionNo user exposure, but scoring relies on offline graders

Diagnosing problems

SymptomLikely causeCheck first
Same wrong format or missed rule across many inputsPrompt failureThe instructions and examples for that rule
Confident details not found in the sourceHallucinationWhether the source was in context, and whether Claude may say it does not know
Wrong answers after documents change, model unchangedRetrievalChunking, re-indexing and which chunks were returned
Simple cases pass, multi-step reasoning cases failModel mismatchTier and effort level for that slice
Answers cut off mid-sentenceOutput limitstop_reason of max_tokens
Quality changes after a deployment with no prompt changeConfiguration driftWhich model ID and prompt version are live

Anthropic's techniques for reducing hallucinations include allowing Claude to say it does not know, asking it to extract direct quotes before answering, requiring citations for each claim, restricting it to the documents provided, and comparing several runs for inconsistencies.

Optimising cost and latency

  • Tokens: cache stable prefixes, send retrieved passages instead of whole documents, make tool results concise.
  • Latency: route simple requests to a smaller tier, lower effort where tests allow, stream responses, run independent tool calls in parallel.
  • Cost on non-urgent work: move overnight or bulk jobs, including large evaluation runs, to the Message Batches API (see 4.5 Message Batches API).

Model cost and latency before you build, from three inputs: call volume from the business owner, tokens per request and model tier. Use the distribution, not the average: a small share of long documents can dominate spend, and SLAs break on the p95, not the median. Then test what happens if volume doubles or the long tail grows.

Plan reliability with the cost model: retries with backoff next to each call, a fallback tier in the orchestration layer, and a circuit breaker at the service boundary.

Every optimisation goes through the same test set and gates.

Monitoring in production

In production, log for each request the model ID, prompt version, token usage (including cache reads and writes), latency, stop_reason and tool calls. Score a sample of live traffic with your offline graders, and alert when scores, errors or token use move. When a metric moves, work out which of three causes it is, because each has a different fix: the inputs changed (data drift), outputs on stable inputs changed (behaviour drift), or a new model or prompt version went live. Break metrics down per request and per segment, since an average can look healthy while a small slice fails or eats the budget. Map each technical metric to the business measure the sponsor reads, such as task success to first-contact resolution. For the monitoring design itself, see Integration.

Example: an evaluation plan

evaluation_plan: invoice-extraction-v3
change_under_test: "Move field extraction from Opus tier to Sonnet tier, effort low"
baseline: "Opus tier, prompt extraction@v14, pinned model IDs"

dataset:
  - production_sample: 400 invoices from the last 30 days, labelled by accounts payable
  - known_failures: 60 invoices from the incident log
  - adversarial: 40 invoices with instructions hidden in the notes field

graders:
  field_accuracy: code, exact match per field against labels
  line_items_sum_to_total: code
  notes_summary_quality: model, rubric scored 1 to 5, calibrated on 50 human-scored items
  injection_followed: code, fails if any output field is outside the schema

trials_per_case: 3      # report the share of cases that pass all 3 trials

gates:
  field_accuracy: ">= 98% and no more than 0.5 points below baseline"
  p95_latency_seconds: "<= 6"
  tokens_per_invoice: "report; target 30% below baseline"
  injection_followed: "0 cases"

rollout:
  online_ab: "10% of traffic for 2 weeks; 50 outputs per week reviewed by people"
  rollback_trigger: "human-reviewed field accuracy below 97% in any week"
  after_rollout: "add every new failure to known_failures"

Rules that decide exam answers

  • Test on real cases. An option that builds the dataset from production traffic and known failures beats one that adds more synthetic cases.
  • One change at a time against a pinned baseline. Otherwise the result cannot be explained.
  • Calibrate model graders with people. Human review is for checking graders and high-stakes samples, not for grading everything.
  • Diagnose before you upgrade. A bigger model does not fix missing context, stale retrieval or a broken prompt.
  • Grade outcomes, and run several trials. One pass of an agent proves little about reliability.
  • Optimisations must pass the same gates. A cheaper system that fails the accuracy gate is not an improvement.

Where it appears in the exam

Evaluation, Testing & Optimization carries 16% of the CCAR-P exam. The guide's own sample item for this domain describes a system that starts giving wrong answers after a change and asks where to look first. Expect symptoms to diagnose, test designs to choose between, and cost or latency targets that must be met without losing quality.

Two sample questions

These are original Timo practice questions. They are not official exam questions.

Question 1

A team released a new extraction prompt after it scored higher on their test set. Production complaints about missing fields then rose. The test set was built six months ago from synthetic invoices. What should the architect do next?

Answer: B. The test set no longer matches real traffic, so the score misled the team; real cases fix the measurement and stay as regression tests. A repeats the same unrepresentative test, C changes something else before finding the cause, and D cannot scale and does not fix the dataset.

Question 2

A ticket summariser sometimes includes order numbers that never appear in the ticket. It uses no retrieval: the input is the ticket text only. The model ID and prompt have not changed recently. What should the architect do first?

Answer: A. The details are invented, so grounding the summary in quoted source text and testing for unsupported numbers addresses the cause and measures the fix. B does not tell Claude to stay within the ticket, C may encourage more order numbers, and D changes a retrieval step this system does not have.

Build exercise

  1. Write success criteria for one Claude feature at your workplace, one line each for accuracy, latency, cost, safety and security.
  2. Collect 30 real inputs and 10 known failures, label the expected results, and write one code grader and one model grader with a rubric.
  3. Run the current version three times per case and record the share of cases that pass all three trials.
  4. Write an evaluation plan like the one above for one change you want to make, with gates and a rollback trigger.

Practise this topic

Sources