Human review and confidence: CCAR-F task statement 5.5
CCAR-F · Context Management & Reliability (15% of the exam)
Task statement 5.5 sits in Context Management & Reliability, 15% of the CCAR-F exam. It tests how to decide which outputs a person checks in an extraction system: measure accuracy per segment, calibrate confidence scores on labelled data, route the uncertain cases to reviewers, and keep sampling the cases you no longer review.
What the official guide covers
The Claude Certified Architect Foundations exam guide (version 1.0, effective July 2026) lists this under task statement 5.5, "Design human review workflows and confidence calibration":
| Knowledge of | Skills in |
|---|---|
| A high overall accuracy figure (such as 97%) can hide poor results on particular document types or fields | Analysing accuracy by document type and field to confirm every segment performs before reducing human review |
| Stratified random sampling measures error rates in high-confidence extractions and finds new error patterns | Running stratified random samples of high-confidence extractions to keep measuring error rates and spot new patterns |
| Field-level confidence scores, calibrated on labelled validation sets, can direct review effort | Having the model output a confidence score per field, then setting review thresholds with labelled validation sets |
| Accuracy must be checked by document type and field before high-confidence extractions are automated | Sending low-confidence extractions, and those from ambiguous or contradictory source documents, to people, using limited reviewer time where it counts |
Why one accuracy number is not enough
Take an invoice pipeline measured on 5,000 labelled documents.
| Segment | Share of documents | Field accuracy |
|---|---|---|
| Typed PDF invoices | 80% | 99.1% |
| Scanned invoices | 15% | 93.5% |
| Handwritten delivery notes | 5% | 81.0% |
| All documents | 100% | about 97% |
The overall figure looks ready to automate. One segment in twenty is wrong almost one time in five. The same happens across fields: total may be near perfect while due_date fails on a date format one supplier uses. Before reducing review anywhere, break accuracy down by document type and by field, and automate segment by segment.
Anthropic's evaluation guidance points the same way: success criteria should be specific and measurable, most use cases need several criteria rather than one, and test sets should mirror the real mix of tasks, edge cases included.
Field-level confidence, then calibration
Ask for a confidence value next to each extracted field, not one score per document. With structured outputs or a tool schema, each field becomes an object such as {"value": "2026-03-12", "confidence": 0.82}.
A raw confidence number is not a probability. On its own it says little about how often the field is right. Calibrate it: run the model over a labelled validation set, then, for each document type and field, find the lowest score at which accuracy meets your target. That becomes the threshold for that segment. A segment that never meets the target gets no threshold and stays fully reviewed.
import random
from collections import defaultdict
def calibrate(validation, target=0.98, min_rows=50,
steps=(0.5, 0.6, 0.7, 0.8, 0.9, 0.95)):
"""validation: dicts with doc_type, field, confidence, correct (bool), from labelled data.
Returns {(doc_type, field): threshold or None}."""
segments = defaultdict(list)
for row in validation:
segments[(row["doc_type"], row["field"])].append(row)
thresholds = {}
for segment, rows in segments.items():
thresholds[segment] = None # None means always review
for t in steps:
above = [r for r in rows if r["confidence"] >= t]
if len(above) >= min_rows and sum(r["correct"] for r in above) / len(above) >= target:
thresholds[segment] = t
break
return thresholds
def route(extraction, thresholds, sources_conflict=False):
if sources_conflict: # contradictory documents go to a person
return "human_review"
for field, item in extraction["fields"].items():
t = thresholds.get((extraction["doc_type"], field))
if t is None or item["confidence"] < t:
return "human_review"
return "auto_accept"
def audit_sample(accepted, per_stratum=20, seed=None):
"""Stratified random sample of auto-accepted extractions, per document type."""
rng = random.Random(seed)
strata = defaultdict(list)
for item in accepted:
strata[item["doc_type"]].append(item)
return [x for items in strata.values() for x in rng.sample(items, min(per_stratum, len(items)))]
Routing with limited reviewer time
Reviewers are the scarce resource, so put the queue in order:
- Extractions where the source documents are ambiguous or contradict each other (an invoice total that does not match its line items, two dates for one delivery). Confidence does not matter here; a person decides.
- Fields below their calibrated threshold, high-impact fields first (amounts before reference notes).
- Segments with no threshold yet.
- The audit sample from auto-accepted work.
Stakes decide where review goes; confidence decides how much
Confidence estimates how likely an output is to be wrong. It says nothing about what a wrong output costs. Two other properties set that: how expensive the error is if it goes through, and how easily it can be undone. Combine them into one rule: send an output to a person before it takes effect when confidence is low and the action is either hard to reverse or costly when wrong. Let confident, reversible, low-cost outputs through. When the signals disagree, cost and reversibility win, because they decide the damage. And the rule only holds if the confidence score has been calibrated first.
Where the person sits is a separate choice:
| Placement | How it works | Fits |
|---|---|---|
| Review before the action | Nothing takes effect until a person approves | Irreversible or high-cost outputs, such as a payment amount sent to the bank |
| Audit after the action | Output is used at once; a person checks it later | Reversible, lower-cost outputs where speed matters |
| Sampled review | A stratified random share is checked | Measuring quality of auto-accepted work; it guards the system, not each item |
What the reviewer must see
Review fails in two ways. Volume: if more items arrive than a person can read, they stop reading and approve to keep up. Missing context: if the queue shows only the extracted value and an approve button, the reviewer has nothing to check it against. Give each item three things: the source (the page image or text span the value came from), the extracted output, and the reason it was flagged ("due_date confidence 0.62, threshold 0.85 for scanned invoices", or "total does not match line items"). The flag reason tells the reviewer what to look at, and it is also data: count reasons each week to see which segment needs a better prompt or threshold.
Valid is not the same as correct
Schema validation and structured outputs guarantee shape: the date parses, the total is a number, required fields are present. They cannot tell you the value is the right one. A delivery note that says "ordered 3 March, delivered 12 April" can yield a perfectly valid but wrong order_date. Only labelled examples, calibrated thresholds and review catch that. When you find such a case, add it to the labelled set so every later prompt or model change is tested against it.
If you use a model to grade the audit sample instead of a person, treat that grader like the extractor: run it on items people have already labelled and check how often it agrees with them before you trust its scores.
Keep measuring after you automate
Once a segment is auto-accepted, nobody looks at it, which is exactly when a new supplier template or a scanner change can start producing errors. A stratified random sample (a fixed number drawn at random from each document type, or each supplier group) keeps measuring the error rate in every segment, including small ones. A plain random sample of all documents would mostly draw typed PDFs and could miss a small segment completely. Reviewing the first or newest items is not random at all.
Treat the labelled results from these audits as new validation data, and recalibrate thresholds when a segment's error rate moves.
Thresholds also belong to one prompt and one model. A score of 0.85 from today's prompt does not mean what it meant before the prompt was reworded or the model was swapped, so rerun the labelled set and recalibrate after every such change. Keep that set representative of current traffic: one built from the dozen documents the team knew best, and never updated, keeps passing while the live system fails on inputs it never contained.
Which design fits
| Situation | Choose | Why |
|---|---|---|
| Overall accuracy is high and the team wants to cut review | Break down accuracy by document type and field first | The average can hide a weak segment |
| The model reports confidence per field | Calibrate thresholds per segment on labelled data | Raw scores are not error rates |
| Sources contradict each other | Send to a person regardless of confidence | The model cannot settle a real conflict in the source |
| A segment has been automated | Stratified random audit sample | Catches drift and new error patterns |
| Reviewer time is short | Prioritise conflicts and low-confidence, high-impact fields | Time goes where errors are likely and costly |
| Reviewers approve nearly everything without reading | Fewer items, routed by stakes, each shown with source and flag reason | Volume and missing context turn review into a rubber stamp |
| The prompt or model changed after thresholds were set | Rerun the labelled set and recalibrate before trusting the thresholds | Calibration holds only for the version it was measured on |
| Every extraction passes schema validation but some values are wrong | Labelled test cases for the failing pattern, plus review routing | Validation checks shape, not correctness |
Rules that decide exam answers
- Segment before you automate. An option that automates on the overall figure alone is wrong.
- Calibrate, do not trust raw scores. A confidence threshold only means something after it is checked against labelled data, and it must be checked again after any prompt or model change.
- Keep sampling what you automated. High-confidence output still needs a stratified random audit.
- Conflict beats confidence. Ambiguous or contradictory sources go to a person even when the score is high.
- Spend review time by risk. Reviewing everything, or a fixed slice such as the oldest items, wastes the scarcest resource. Cost and reversibility set the stakes; confidence only estimates the chance of error.
- A reviewer needs the source and the reason. An option that sends only the output to an approval queue fails even when the queue is small.
Where it appears in the exam
Context Management & Reliability is a primary domain in four of the six exam scenarios: Customer Support Resolution Agent, Code Generation with Claude Code, Multi-Agent Research System and Structured Data Extraction. Review routing and calibration fit the Structured Data Extraction scenario. The guide's preparation exercise 3 ends with a human review routing step using field-level confidence and accuracy by document type, and lists Domain 5 among the domains it reinforces.
Two sample questions
These are original Timo practice questions. They are not official exam questions.
Build exercise
- Label 200 extractions from a mix of document types: for each field record the model's value, its confidence and whether it was correct.
- Compute accuracy overall, then by document type and by field. Write down any segment more than five points below the average.
- Run
calibrateon the labelled rows and list which segments get a threshold and which stay fully reviewed. - Route a fresh set of 50 extractions with
route, then draw anaudit_samplefrom the auto-accepted ones and check them by hand.
Practise this topic
- Claude Certified Architect practice exam: free, 20 questions, no sign-up
- Claude Certified Architect hub
- CCAR-F study guide: all topics
- Worked example: Professional output validation
- Same topic in another exam: Governance, Safety and Risk (CCAR-P) and Governance, Risk and Responsible Use (CCAO-F)
- Previous topic: 5.4 Large codebase context
- Next topic: 5.6 Information provenance
Sources
- Claude Certified Architect Foundations Exam Guide, version 1.0, effective July 2026 (Anthropic), task statement 5.5
- Anthropic documentation: Define success criteria and build evaluations
- Anthropic documentation: Structured outputs
By Amotion AI