Multi-pass review: CCAR-F task statement 4.6
CCAR-F · Prompt Engineering & Structured Output (20% of the exam)
Task statement 4.6 sits in Prompt Engineering & Structured Output, 20% of the CCAR-F exam. It tests how you structure a review so that it catches problems: a separate Claude instance that has not seen the generator's reasoning, and large reviews split into focused per-file passes plus a pass that looks across files.
What the official guide covers
The Claude Certified Architect Foundations exam guide (version 1.0, effective July 2026) lists this under task statement 4.6, "Design multi-instance and multi-pass review architectures":
| Knowledge of | Skills in |
|---|---|
| The limit of self-review: a model keeps the reasoning from generation, so in the same session it is less likely to question its own decisions | Using a second, independent Claude instance to review generated code without the generator's reasoning |
| Independent review instances, with no prior reasoning in context, catch subtle issues better than self-review instructions or extended thinking | Splitting large multi-file reviews into per-file passes for local issues and separate integration passes for data flow across files |
| Multi-pass review: per-file local passes plus cross-file integration passes, to avoid diluted attention and contradictory findings | Running verification passes where the model reports a confidence with each finding, so findings can be routed for review |
Why self-review misses things
When you ask the same conversation to review what it just wrote, the review starts with the whole generation in context: the plan, the assumptions and the reasons for each choice. The reviewer reads the code through those same assumptions. If an assumption was wrong, the review inherits it. Asking it to "be critical" or to think longer does not remove that context.
An independent review is a new request. Its messages list holds only what a reviewer needs: the requirements, the code and the review criteria. It has no access to the generator's transcript. It can also use its own system prompt written for reviewing, not generating.
In Claude Code the same rule applies. A subagent starts with a fresh context window and does not see the conversation history, so a review subagent given the diff and the criteria is independent. A fork is not: it inherits the whole conversation so far, generator reasoning included. Use a fresh subagent or a new session for review.
Anthropic describes the same structure in two places. Its prompting guide calls the most common chaining pattern self-correction: generate a draft, have Claude review it against criteria, then refine, with each step a separate API call. Its engineering post "Building effective agents" describes an evaluator-optimizer workflow, where one call generates and another evaluates and gives feedback.
Why one pass over a large change fails
A single pass over 20 changed files spreads Claude's attention across all of them. Typical symptoms: detailed comments on some files and almost none on others, obvious bugs missed, and the same pattern flagged in one file and approved in another.
The fix is to split the review by what each pass needs to see:
- Per-file passes. One call per file, looking only for local issues: logic errors, unsafe input handling, broken error paths. These calls are independent, so they can run in parallel.
- An integration pass. One call that sees the changed interfaces and the per-file findings, and looks only at how the files work together: a changed function signature with callers not updated, a value sent in cents and read as dollars, a renamed field.
"Building effective agents" gives the reason: for tasks with several considerations, models generally perform better when each one gets a separate call with focused attention.
A per-file and integration review (Python)
import anthropic
client = anthropic.Anthropic()
MODEL = "your-model-id"
FILE_REVIEW = (
"You review one file from a pull request. Report only issues inside this "
"file: logic errors, unsafe input handling, broken error paths. For each "
"finding give location, issue, severity and confidence (high, medium, low)."
)
INTEGRATION_REVIEW = (
"You review how the changed files work together. Report only cross-file "
"issues: changed signatures whose callers were not updated, data passed "
"in a shape or unit the receiver does not expect, renamed fields. Do not "
"repeat local findings."
)
def ask(system, content):
# Each call is a new conversation: no generator reasoning, no earlier passes.
response = client.messages.create(
model=MODEL,
max_tokens=4096,
system=system,
messages=[{"role": "user", "content": content}],
)
return next(b.text for b in response.content if b.type == "text")
def review_pull_request(changed_files, interface_summary):
local = {
path: ask(FILE_REVIEW, f'<file path="{path}">\n{code}\n</file>')
for path, code in changed_files.items()
}
findings = "\n".join(
f'<findings file="{path}">\n{text}\n</findings>' for path, text in local.items()
)
integration = ask(
INTEGRATION_REVIEW,
f"<changed_interfaces>\n{interface_summary}\n</changed_interfaces>\n{findings}",
)
return local, integration
The per-file calls can run concurrently in production. Each ask call starts a fresh conversation, so the reviewer is independent of whatever generated the code.
Check coverage before the integration pass. Every changed file must come back with findings or an explicit "no findings". If one call times out or fails, rerun it or report the gap: a summary built over 20 of 22 files reads as complete and is not.
Routing findings by confidence
The guide's third skill adds a verification pass where the model reports a confidence with each finding, so you can route them: high-confidence findings posted directly, low-confidence ones sent to a person or a second check. Treat the confidence as a routing signal, not as truth. Before relying on a threshold, label a sample of findings yourself and check that high-confidence findings really are right more often. Task statement 5.5 covers calibration in more depth.
Write the verification pass the way you would write a grader. Anthropic's guide to building evaluations gives three tips for model-based grading, and each applies here:
- A clear rubric. State what makes a finding valid, for example "the reviewer can name the input that triggers the bug".
- A fixed set of verdicts. Ask for confirmed, likely or unlikely, not a free-form score that drifts from run to run.
- Reasoning before the verdict. Have the pass explain its check first, then give the label.
Confidence also does not set the stakes. It estimates how likely a finding is to be wrong. What a missed bug would cost, and how easily it could be undone, decides which findings must reach a person. A low-confidence finding in payment code still goes to a reviewer; a confident finding about a log message does not need one.
Other ways to split a review
Splitting by file is the split the exam guide names, but two related workflow patterns from "Building effective agents" use the same idea:
- One pass per concern, run in parallel. A security pass, a correctness pass and a data-handling pass, each with its own criteria, then one call that merges and deduplicates the findings. Each prompt can be tested and improved on its own.
- A focused revision pass. When generated output keeps breaking a few rules (a forbidden library, a missing log line), a second call whose only job is to find and fix those violations beats adding more rules to the first prompt.
| Situation | Choose | Why |
|---|---|---|
| Generated code needs a review before merge | A second instance with only the requirements, code and criteria | It does not share the generator's assumptions |
| The plan is to ask the same session to "review carefully" | Replace with an independent instance | Same context, same blind spots |
| A 20-file review gives uneven, contradictory findings | Per-file passes plus an integration pass | Each pass has a narrow focus |
| A bug crosses files, such as a changed signature | The integration pass | Per-file passes cannot see the caller |
| Someone proposes a model with a larger context window | Split the review instead | More room does not fix diluted attention |
| Too many findings for people to check | Confidence per finding, thresholds checked on labelled samples | Routes human effort where it is needed |
| Security and style findings crowd each other out in one prompt | Separate passes per concern, then a merge step | Each pass has one set of criteria |
Rules that decide exam answers
- A fresh instance beats "review your own work". The reviewer must not carry the generator's reasoning, so a forked subagent that inherits the conversation does not count.
- More thinking in the same session is not independence. The guide ranks an independent instance above self-review instructions and extended thinking.
- Split by file, then integrate. Per-file passes find local issues; a separate pass finds cross-file issues. Confirm every file was reviewed before merging the findings.
- A bigger context window is not the fix. The problem is attention, not space.
- Confidence routes findings; it does not replace criteria. Check the scores against labelled examples before trusting a threshold.
- Stakes decide what a person must see. Confidence says how likely a finding is wrong; cost and reversibility say how much that matters.
Where it appears in the exam
Prompt Engineering & Structured Output is a primary domain in two of the six exam scenarios: Claude Code for Continuous Integration and Structured Data Extraction. Multi-pass review fits the CI scenario most closely, where Claude runs automated code reviews and gives feedback on pull requests.
Two sample questions
These are original Timo practice questions. They are not official exam questions.
Build exercise
- Ask Claude to write a function from a short spec, then ask the same conversation to review it. Save the findings.
- Start a new conversation with only the spec, the code and three review criteria. Compare its findings with step 1.
- Take a change that touches about ten files. Run one pass over all of them, then the per-file plus integration review above. Compare depth and contradictions.
- Add a confidence to each finding, label 30 findings yourself as right or wrong, and check whether high-confidence findings are right more often before you route on them.
Practise this topic
- Claude Certified Architect practice exam: free, 20 questions, no sign-up
- Claude Certified Architect hub
- CCAR-F study guide: all topics
- Previous topic: 4.5 Message Batches API
- Next topic: 5.1 Context across long conversations
Sources
- Claude Certified Architect Foundations Exam Guide, version 1.0, effective July 2026 (Anthropic), task statement 4.6
- Anthropic documentation: Prompting best practices
- Anthropic: Building effective agents
- Anthropic documentation: Define success criteria and build evaluations
- Claude Code documentation: Create custom subagents
By Amotion AI