Evaluation, Testing and Optimization: CCAR-P domain 4 study guide
CCAR-P · Evaluation, Testing & Optimization (16% of the exam)
Domain 4 is Evaluation, Testing & Optimization, 16% of the CCAR-P exam. It tests whether you can prove a Claude system works, find the cause when it does not, and make it cheaper or faster without breaking it.
What the official guide covers
The Claude Certified Architect Professional exam guide (version 1.0, effective July 2026) lists six tasks under Domain 4, "Evaluation, Testing & Optimization":
| What the guide lists | What it means in practice |
|---|---|
| Define evaluation metrics (accuracy, latency, cost, safety, security) | Set a measurable target for each dimension before building |
| Design evaluation datasets and test frameworks using mixed methods | Build test sets from real cases and grade them with code, models and people |
| Conduct A/B testing and iterative improvement | Change one thing, compare against a baseline, keep or roll back |
| Diagnose system issues (prompt failure, hallucinations, model mismatch) | Tell the cause from the symptom and check the right component first |
| Optimise token usage, latency and cost-performance | Cut tokens and time while holding quality on the test set |
| Monitor performance with logging and observability tools | Keep measuring in production, not only before launch |
Defining metrics
Anthropic's guidance says success criteria should be specific, measurable, achievable and relevant, and most systems need several. Decide how to measure each dimension.
| Dimension | Measured by | Example target |
|---|---|---|
| Accuracy | Code grader against labelled answers, or a model grader with a rubric | 98% of invoice fields exact |
| Latency | Time to first token and total time, at the 50th and 95th percentile | p95 under 6 seconds |
| Cost | Tokens per completed task, including cache reads and writes, plus review time | 30% fewer tokens than baseline |
| Safety | Model grader for policy breaches, calibrated against human labels | No harmful outputs in the red-team set |
| Security | Prompt injection and data leakage tests | No injected instruction followed |
Write the criteria and the first test cases before production code. That forces "summarise claims accurately" into something testable (which fields, what pass rate, which failures count) while the design is still cheap to change, and it gives you the gate every later change must pass. Take thresholds from the business requirement, not from whatever the prototype scored.
Building datasets and graders
Build the dataset from what the system will really see:
- Real traffic, labelled by the people who own the outcome.
- Known failures from incidents and complaints, kept as permanent tests.
- Edge cases and adversarial inputs, such as injected instructions.
- Both directions: cases where Claude should act and where it should not.
Anthropic's guidance favours volume: many auto-graded cases beat a few hand-graded ones. Mix three graders:
| Grader | Strength | Weakness | Use for |
|---|---|---|---|
| Code | Fast, exact, repeatable | Misses valid variations | Fields, formats, totals, whether an action happened |
| Model | Scales to open-ended output | Needs a clear rubric and calibration | Tone, completeness, faithfulness to sources |
| Human | Highest judgement | Slow and costly | Calibrating model graders; high-stakes samples |
Climb this ladder from the top: use code wherever the check is unambiguous, even when the stakes are high (a leaked customer ID is a string match). One requirement can need two graders: code checks whether a forbidden action happened; a model grades whether a summary is faithful. Give model graders a detailed rubric and a small fixed set of verdict labels, let them reason before scoring, and consider a different model from the one being graded. Before trusting a model grader, compare its verdicts with human labels on a sample; an uncalibrated grader gives confident scores you cannot rely on.
Two dataset traps. A single-turn set says nothing about whether a chat system keeps facts straight across turns, so multi-turn behaviour needs its own set of whole conversations. And a test set that predates a prompt change keeps passing while measuring behaviour the system no longer has: update the cases whenever the prompt, retrieval or model changes.
For agents, Anthropic's evals article adds: grade the outcome, not the path; run several trials per task; and separate capability tests, expected to fail at first, from regression tests, which should almost always pass. Read transcripts to see whether a failure is the agent's or the grader's.
A/B testing and iteration
- Pin the model ID and prompt version for the baseline.
- Change one thing: the prompt, the model tier, the effort level, the retrieval settings.
- Run both versions on the same offline test set, with several trials per case.
- If the change passes every gate, send a small share of live traffic to it and compare the same metrics, plus sampled human review.
- Keep it, or roll back. Add any new failures to the test set.
A live test needs a written hypothesis that names the change, the primary metric, the minimum gain and the limits on other metrics (for example, "task success up at least 5 points with p95 latency unchanged"). Assign users or sessions at random and keep each one in the same group. Set the sample size and the primary metric before starting; LLM outputs vary more than ordinary software, so samples run larger, and 50 sessions per arm rarely separate a real gain from noise. Set the confidence you need from the cost of a wrong call, not from the effect you expect: a small prompt change to a medical summariser still needs a large sample. A statistically significant result still has to be large enough to justify maintaining the new version, and no secondary metric may get worse.
| Situation | Test pattern | Trade-off |
|---|---|---|
| Enough traffic, a bad output is tolerable and reversible | Live A/B split | Real user signals, but some users see the worse version |
| A single bad output is unacceptable, or traffic is low | Shadow test: the new version gets a copy of live requests, users see only the current version | No user exposure, but scoring relies on offline graders |
Diagnosing problems
| Symptom | Likely cause | Check first |
|---|---|---|
| Same wrong format or missed rule across many inputs | Prompt failure | The instructions and examples for that rule |
| Confident details not found in the source | Hallucination | Whether the source was in context, and whether Claude may say it does not know |
| Wrong answers after documents change, model unchanged | Retrieval | Chunking, re-indexing and which chunks were returned |
| Simple cases pass, multi-step reasoning cases fail | Model mismatch | Tier and effort level for that slice |
| Answers cut off mid-sentence | Output limit | stop_reason of max_tokens |
| Quality changes after a deployment with no prompt change | Configuration drift | Which model ID and prompt version are live |
Anthropic's techniques for reducing hallucinations include allowing Claude to say it does not know, asking it to extract direct quotes before answering, requiring citations for each claim, restricting it to the documents provided, and comparing several runs for inconsistencies.
Optimising cost and latency
- Tokens: cache stable prefixes, send retrieved passages instead of whole documents, make tool results concise.
- Latency: route simple requests to a smaller tier, lower effort where tests allow, stream responses, run independent tool calls in parallel.
- Cost on non-urgent work: move overnight or bulk jobs, including large evaluation runs, to the Message Batches API (see 4.5 Message Batches API).
Model cost and latency before you build, from three inputs: call volume from the business owner, tokens per request and model tier. Use the distribution, not the average: a small share of long documents can dominate spend, and SLAs break on the p95, not the median. Then test what happens if volume doubles or the long tail grows.
Plan reliability with the cost model: retries with backoff next to each call, a fallback tier in the orchestration layer, and a circuit breaker at the service boundary.
Every optimisation goes through the same test set and gates.
Monitoring in production
In production, log for each request the model ID, prompt version, token usage (including cache reads and writes), latency, stop_reason and tool calls. Score a sample of live traffic with your offline graders, and alert when scores, errors or token use move. When a metric moves, work out which of three causes it is, because each has a different fix: the inputs changed (data drift), outputs on stable inputs changed (behaviour drift), or a new model or prompt version went live. Break metrics down per request and per segment, since an average can look healthy while a small slice fails or eats the budget. Map each technical metric to the business measure the sponsor reads, such as task success to first-contact resolution. For the monitoring design itself, see Integration.
Example: an evaluation plan
evaluation_plan: invoice-extraction-v3
change_under_test: "Move field extraction from Opus tier to Sonnet tier, effort low"
baseline: "Opus tier, prompt extraction@v14, pinned model IDs"
dataset:
- production_sample: 400 invoices from the last 30 days, labelled by accounts payable
- known_failures: 60 invoices from the incident log
- adversarial: 40 invoices with instructions hidden in the notes field
graders:
field_accuracy: code, exact match per field against labels
line_items_sum_to_total: code
notes_summary_quality: model, rubric scored 1 to 5, calibrated on 50 human-scored items
injection_followed: code, fails if any output field is outside the schema
trials_per_case: 3 # report the share of cases that pass all 3 trials
gates:
field_accuracy: ">= 98% and no more than 0.5 points below baseline"
p95_latency_seconds: "<= 6"
tokens_per_invoice: "report; target 30% below baseline"
injection_followed: "0 cases"
rollout:
online_ab: "10% of traffic for 2 weeks; 50 outputs per week reviewed by people"
rollback_trigger: "human-reviewed field accuracy below 97% in any week"
after_rollout: "add every new failure to known_failures"
Rules that decide exam answers
- Test on real cases. An option that builds the dataset from production traffic and known failures beats one that adds more synthetic cases.
- One change at a time against a pinned baseline. Otherwise the result cannot be explained.
- Calibrate model graders with people. Human review is for checking graders and high-stakes samples, not for grading everything.
- Diagnose before you upgrade. A bigger model does not fix missing context, stale retrieval or a broken prompt.
- Grade outcomes, and run several trials. One pass of an agent proves little about reliability.
- Optimisations must pass the same gates. A cheaper system that fails the accuracy gate is not an improvement.
Where it appears in the exam
Evaluation, Testing & Optimization carries 16% of the CCAR-P exam. The guide's own sample item for this domain describes a system that starts giving wrong answers after a change and asks where to look first. Expect symptoms to diagnose, test designs to choose between, and cost or latency targets that must be met without losing quality.
Two sample questions
These are original Timo practice questions. They are not official exam questions.
Build exercise
- Write success criteria for one Claude feature at your workplace, one line each for accuracy, latency, cost, safety and security.
- Collect 30 real inputs and 10 known failures, label the expected results, and write one code grader and one model grader with a rubric.
- Run the current version three times per case and record the share of cases that pass all three trials.
- Write an evaluation plan like the one above for one change you want to make, with gates and a rollback trigger.
Practise this topic
- Claude Certified Architect Professional practice exam: free, 20 questions, no sign-up
- Claude Certified Architect hub
- CCAR-P study guide: all topics
- Worked example: Professional output validation
- Previous topic: Integration
- Next topic: Governance, Safety and Risk
Sources
- Claude Certified Architect, Professional Exam Guide, version 1.0, effective July 2026 (Anthropic), Domain 4: Evaluation, Testing & Optimization
- Anthropic documentation: Define success criteria and build evaluations
- Anthropic: Demystifying evals for AI agents
- Anthropic documentation: Reduce hallucinations
By Amotion AI