TimoBy Amotion AI

Human review and confidence: CCAR-F task statement 5.5

CCAR-F · Context Management & Reliability (15% of the exam)

Task statement 5.5 sits in Context Management & Reliability, 15% of the CCAR-F exam. It tests how to decide which outputs a person checks in an extraction system: measure accuracy per segment, calibrate confidence scores on labelled data, route the uncertain cases to reviewers, and keep sampling the cases you no longer review.

What the official guide covers

The Claude Certified Architect Foundations exam guide (version 1.0, effective July 2026) lists this under task statement 5.5, "Design human review workflows and confidence calibration":

Knowledge ofSkills in
A high overall accuracy figure (such as 97%) can hide poor results on particular document types or fieldsAnalysing accuracy by document type and field to confirm every segment performs before reducing human review
Stratified random sampling measures error rates in high-confidence extractions and finds new error patternsRunning stratified random samples of high-confidence extractions to keep measuring error rates and spot new patterns
Field-level confidence scores, calibrated on labelled validation sets, can direct review effortHaving the model output a confidence score per field, then setting review thresholds with labelled validation sets
Accuracy must be checked by document type and field before high-confidence extractions are automatedSending low-confidence extractions, and those from ambiguous or contradictory source documents, to people, using limited reviewer time where it counts

Why one accuracy number is not enough

Take an invoice pipeline measured on 5,000 labelled documents.

SegmentShare of documentsField accuracy
Typed PDF invoices80%99.1%
Scanned invoices15%93.5%
Handwritten delivery notes5%81.0%
All documents100%about 97%

The overall figure looks ready to automate. One segment in twenty is wrong almost one time in five. The same happens across fields: total may be near perfect while due_date fails on a date format one supplier uses. Before reducing review anywhere, break accuracy down by document type and by field, and automate segment by segment.

Anthropic's evaluation guidance points the same way: success criteria should be specific and measurable, most use cases need several criteria rather than one, and test sets should mirror the real mix of tasks, edge cases included.

Field-level confidence, then calibration

Ask for a confidence value next to each extracted field, not one score per document. With structured outputs or a tool schema, each field becomes an object such as {"value": "2026-03-12", "confidence": 0.82}.

A raw confidence number is not a probability. On its own it says little about how often the field is right. Calibrate it: run the model over a labelled validation set, then, for each document type and field, find the lowest score at which accuracy meets your target. That becomes the threshold for that segment. A segment that never meets the target gets no threshold and stays fully reviewed.

import random
from collections import defaultdict

def calibrate(validation, target=0.98, min_rows=50,
              steps=(0.5, 0.6, 0.7, 0.8, 0.9, 0.95)):
    """validation: dicts with doc_type, field, confidence, correct (bool), from labelled data.
    Returns {(doc_type, field): threshold or None}."""
    segments = defaultdict(list)
    for row in validation:
        segments[(row["doc_type"], row["field"])].append(row)
    thresholds = {}
    for segment, rows in segments.items():
        thresholds[segment] = None                     # None means always review
        for t in steps:
            above = [r for r in rows if r["confidence"] >= t]
            if len(above) >= min_rows and sum(r["correct"] for r in above) / len(above) >= target:
                thresholds[segment] = t
                break
    return thresholds

def route(extraction, thresholds, sources_conflict=False):
    if sources_conflict:                               # contradictory documents go to a person
        return "human_review"
    for field, item in extraction["fields"].items():
        t = thresholds.get((extraction["doc_type"], field))
        if t is None or item["confidence"] < t:
            return "human_review"
    return "auto_accept"

def audit_sample(accepted, per_stratum=20, seed=None):
    """Stratified random sample of auto-accepted extractions, per document type."""
    rng = random.Random(seed)
    strata = defaultdict(list)
    for item in accepted:
        strata[item["doc_type"]].append(item)
    return [x for items in strata.values() for x in rng.sample(items, min(per_stratum, len(items)))]

Routing with limited reviewer time

Reviewers are the scarce resource, so put the queue in order:

  1. Extractions where the source documents are ambiguous or contradict each other (an invoice total that does not match its line items, two dates for one delivery). Confidence does not matter here; a person decides.
  2. Fields below their calibrated threshold, high-impact fields first (amounts before reference notes).
  3. Segments with no threshold yet.
  4. The audit sample from auto-accepted work.

Stakes decide where review goes; confidence decides how much

Confidence estimates how likely an output is to be wrong. It says nothing about what a wrong output costs. Two other properties set that: how expensive the error is if it goes through, and how easily it can be undone. Combine them into one rule: send an output to a person before it takes effect when confidence is low and the action is either hard to reverse or costly when wrong. Let confident, reversible, low-cost outputs through. When the signals disagree, cost and reversibility win, because they decide the damage. And the rule only holds if the confidence score has been calibrated first.

Where the person sits is a separate choice:

PlacementHow it worksFits
Review before the actionNothing takes effect until a person approvesIrreversible or high-cost outputs, such as a payment amount sent to the bank
Audit after the actionOutput is used at once; a person checks it laterReversible, lower-cost outputs where speed matters
Sampled reviewA stratified random share is checkedMeasuring quality of auto-accepted work; it guards the system, not each item

What the reviewer must see

Review fails in two ways. Volume: if more items arrive than a person can read, they stop reading and approve to keep up. Missing context: if the queue shows only the extracted value and an approve button, the reviewer has nothing to check it against. Give each item three things: the source (the page image or text span the value came from), the extracted output, and the reason it was flagged ("due_date confidence 0.62, threshold 0.85 for scanned invoices", or "total does not match line items"). The flag reason tells the reviewer what to look at, and it is also data: count reasons each week to see which segment needs a better prompt or threshold.

Valid is not the same as correct

Schema validation and structured outputs guarantee shape: the date parses, the total is a number, required fields are present. They cannot tell you the value is the right one. A delivery note that says "ordered 3 March, delivered 12 April" can yield a perfectly valid but wrong order_date. Only labelled examples, calibrated thresholds and review catch that. When you find such a case, add it to the labelled set so every later prompt or model change is tested against it.

If you use a model to grade the audit sample instead of a person, treat that grader like the extractor: run it on items people have already labelled and check how often it agrees with them before you trust its scores.

Keep measuring after you automate

Once a segment is auto-accepted, nobody looks at it, which is exactly when a new supplier template or a scanner change can start producing errors. A stratified random sample (a fixed number drawn at random from each document type, or each supplier group) keeps measuring the error rate in every segment, including small ones. A plain random sample of all documents would mostly draw typed PDFs and could miss a small segment completely. Reviewing the first or newest items is not random at all.

Treat the labelled results from these audits as new validation data, and recalibrate thresholds when a segment's error rate moves.

Thresholds also belong to one prompt and one model. A score of 0.85 from today's prompt does not mean what it meant before the prompt was reworded or the model was swapped, so rerun the labelled set and recalibrate after every such change. Keep that set representative of current traffic: one built from the dozen documents the team knew best, and never updated, keeps passing while the live system fails on inputs it never contained.

Which design fits

SituationChooseWhy
Overall accuracy is high and the team wants to cut reviewBreak down accuracy by document type and field firstThe average can hide a weak segment
The model reports confidence per fieldCalibrate thresholds per segment on labelled dataRaw scores are not error rates
Sources contradict each otherSend to a person regardless of confidenceThe model cannot settle a real conflict in the source
A segment has been automatedStratified random audit sampleCatches drift and new error patterns
Reviewer time is shortPrioritise conflicts and low-confidence, high-impact fieldsTime goes where errors are likely and costly
Reviewers approve nearly everything without readingFewer items, routed by stakes, each shown with source and flag reasonVolume and missing context turn review into a rubber stamp
The prompt or model changed after thresholds were setRerun the labelled set and recalibrate before trusting the thresholdsCalibration holds only for the version it was measured on
Every extraction passes schema validation but some values are wrongLabelled test cases for the failing pattern, plus review routingValidation checks shape, not correctness

Rules that decide exam answers

  • Segment before you automate. An option that automates on the overall figure alone is wrong.
  • Calibrate, do not trust raw scores. A confidence threshold only means something after it is checked against labelled data, and it must be checked again after any prompt or model change.
  • Keep sampling what you automated. High-confidence output still needs a stratified random audit.
  • Conflict beats confidence. Ambiguous or contradictory sources go to a person even when the score is high.
  • Spend review time by risk. Reviewing everything, or a fixed slice such as the oldest items, wastes the scarcest resource. Cost and reversibility set the stakes; confidence only estimates the chance of error.
  • A reviewer needs the source and the reason. An option that sends only the output to an approval queue fails even when the queue is small.

Where it appears in the exam

Context Management & Reliability is a primary domain in four of the six exam scenarios: Customer Support Resolution Agent, Code Generation with Claude Code, Multi-Agent Research System and Structured Data Extraction. Review routing and calibration fit the Structured Data Extraction scenario. The guide's preparation exercise 3 ends with a human review routing step using field-level confidence and accuracy by document type, and lists Domain 5 among the domains it reinforces.

Two sample questions

These are original Timo practice questions. They are not official exam questions.

Question 1

An invoice extraction pipeline scores 97% field accuracy on 5,000 labelled documents. The team plans to stop human review for every extraction the model marks above 0.9 confidence. What should they do first?

Answer: D. The overall figure can hide a weak segment, and 0.9 means nothing until it is checked against labelled data per segment. A automates more on the same untested number, B adds text but no calibration, and C confirms the average without looking inside it.

Question 2

After automating typed invoices, a team reviews the 200 oldest auto-accepted extractions each week and finds almost no errors. Two months later, a supplier's new invoice layout causes wrong totals for weeks before anyone notices. Which change would have caught this sooner?

Answer: A. Sampling each supplier group at random measures every segment, including the one that changed. B is still not random and may miss a small supplier, C does not help when the model is confident and wrong, and D only catches large totals.

Build exercise

  1. Label 200 extractions from a mix of document types: for each field record the model's value, its confidence and whether it was correct.
  2. Compute accuracy overall, then by document type and by field. Write down any segment more than five points below the average.
  3. Run calibrate on the labelled rows and list which segments get a threshold and which stay fully reviewed.
  4. Route a fresh set of 50 extractions with route, then draw an audit_sample from the auto-accepted ones and check them by hand.

Practise this topic

Sources