TimoBy Amotion AI

Multi-pass review: CCAR-F task statement 4.6

CCAR-F · Prompt Engineering & Structured Output (20% of the exam)

Task statement 4.6 sits in Prompt Engineering & Structured Output, 20% of the CCAR-F exam. It tests how you structure a review so that it catches problems: a separate Claude instance that has not seen the generator's reasoning, and large reviews split into focused per-file passes plus a pass that looks across files.

What the official guide covers

The Claude Certified Architect Foundations exam guide (version 1.0, effective July 2026) lists this under task statement 4.6, "Design multi-instance and multi-pass review architectures":

Knowledge ofSkills in
The limit of self-review: a model keeps the reasoning from generation, so in the same session it is less likely to question its own decisionsUsing a second, independent Claude instance to review generated code without the generator's reasoning
Independent review instances, with no prior reasoning in context, catch subtle issues better than self-review instructions or extended thinkingSplitting large multi-file reviews into per-file passes for local issues and separate integration passes for data flow across files
Multi-pass review: per-file local passes plus cross-file integration passes, to avoid diluted attention and contradictory findingsRunning verification passes where the model reports a confidence with each finding, so findings can be routed for review

Why self-review misses things

When you ask the same conversation to review what it just wrote, the review starts with the whole generation in context: the plan, the assumptions and the reasons for each choice. The reviewer reads the code through those same assumptions. If an assumption was wrong, the review inherits it. Asking it to "be critical" or to think longer does not remove that context.

An independent review is a new request. Its messages list holds only what a reviewer needs: the requirements, the code and the review criteria. It has no access to the generator's transcript. It can also use its own system prompt written for reviewing, not generating.

In Claude Code the same rule applies. A subagent starts with a fresh context window and does not see the conversation history, so a review subagent given the diff and the criteria is independent. A fork is not: it inherits the whole conversation so far, generator reasoning included. Use a fresh subagent or a new session for review.

Anthropic describes the same structure in two places. Its prompting guide calls the most common chaining pattern self-correction: generate a draft, have Claude review it against criteria, then refine, with each step a separate API call. Its engineering post "Building effective agents" describes an evaluator-optimizer workflow, where one call generates and another evaluates and gives feedback.

Why one pass over a large change fails

A single pass over 20 changed files spreads Claude's attention across all of them. Typical symptoms: detailed comments on some files and almost none on others, obvious bugs missed, and the same pattern flagged in one file and approved in another.

The fix is to split the review by what each pass needs to see:

  • Per-file passes. One call per file, looking only for local issues: logic errors, unsafe input handling, broken error paths. These calls are independent, so they can run in parallel.
  • An integration pass. One call that sees the changed interfaces and the per-file findings, and looks only at how the files work together: a changed function signature with callers not updated, a value sent in cents and read as dollars, a renamed field.

"Building effective agents" gives the reason: for tasks with several considerations, models generally perform better when each one gets a separate call with focused attention.

A per-file and integration review (Python)

import anthropic

client = anthropic.Anthropic()
MODEL = "your-model-id"

FILE_REVIEW = (
    "You review one file from a pull request. Report only issues inside this "
    "file: logic errors, unsafe input handling, broken error paths. For each "
    "finding give location, issue, severity and confidence (high, medium, low)."
)
INTEGRATION_REVIEW = (
    "You review how the changed files work together. Report only cross-file "
    "issues: changed signatures whose callers were not updated, data passed "
    "in a shape or unit the receiver does not expect, renamed fields. Do not "
    "repeat local findings."
)

def ask(system, content):
    # Each call is a new conversation: no generator reasoning, no earlier passes.
    response = client.messages.create(
        model=MODEL,
        max_tokens=4096,
        system=system,
        messages=[{"role": "user", "content": content}],
    )
    return next(b.text for b in response.content if b.type == "text")

def review_pull_request(changed_files, interface_summary):
    local = {
        path: ask(FILE_REVIEW, f'<file path="{path}">\n{code}\n</file>')
        for path, code in changed_files.items()
    }
    findings = "\n".join(
        f'<findings file="{path}">\n{text}\n</findings>' for path, text in local.items()
    )
    integration = ask(
        INTEGRATION_REVIEW,
        f"<changed_interfaces>\n{interface_summary}\n</changed_interfaces>\n{findings}",
    )
    return local, integration

The per-file calls can run concurrently in production. Each ask call starts a fresh conversation, so the reviewer is independent of whatever generated the code.

Check coverage before the integration pass. Every changed file must come back with findings or an explicit "no findings". If one call times out or fails, rerun it or report the gap: a summary built over 20 of 22 files reads as complete and is not.

Routing findings by confidence

The guide's third skill adds a verification pass where the model reports a confidence with each finding, so you can route them: high-confidence findings posted directly, low-confidence ones sent to a person or a second check. Treat the confidence as a routing signal, not as truth. Before relying on a threshold, label a sample of findings yourself and check that high-confidence findings really are right more often. Task statement 5.5 covers calibration in more depth.

Write the verification pass the way you would write a grader. Anthropic's guide to building evaluations gives three tips for model-based grading, and each applies here:

  • A clear rubric. State what makes a finding valid, for example "the reviewer can name the input that triggers the bug".
  • A fixed set of verdicts. Ask for confirmed, likely or unlikely, not a free-form score that drifts from run to run.
  • Reasoning before the verdict. Have the pass explain its check first, then give the label.

Confidence also does not set the stakes. It estimates how likely a finding is to be wrong. What a missed bug would cost, and how easily it could be undone, decides which findings must reach a person. A low-confidence finding in payment code still goes to a reviewer; a confident finding about a log message does not need one.

Other ways to split a review

Splitting by file is the split the exam guide names, but two related workflow patterns from "Building effective agents" use the same idea:

  • One pass per concern, run in parallel. A security pass, a correctness pass and a data-handling pass, each with its own criteria, then one call that merges and deduplicates the findings. Each prompt can be tested and improved on its own.
  • A focused revision pass. When generated output keeps breaking a few rules (a forbidden library, a missing log line), a second call whose only job is to find and fix those violations beats adding more rules to the first prompt.
SituationChooseWhy
Generated code needs a review before mergeA second instance with only the requirements, code and criteriaIt does not share the generator's assumptions
The plan is to ask the same session to "review carefully"Replace with an independent instanceSame context, same blind spots
A 20-file review gives uneven, contradictory findingsPer-file passes plus an integration passEach pass has a narrow focus
A bug crosses files, such as a changed signatureThe integration passPer-file passes cannot see the caller
Someone proposes a model with a larger context windowSplit the review insteadMore room does not fix diluted attention
Too many findings for people to checkConfidence per finding, thresholds checked on labelled samplesRoutes human effort where it is needed
Security and style findings crowd each other out in one promptSeparate passes per concern, then a merge stepEach pass has one set of criteria

Rules that decide exam answers

  • A fresh instance beats "review your own work". The reviewer must not carry the generator's reasoning, so a forked subagent that inherits the conversation does not count.
  • More thinking in the same session is not independence. The guide ranks an independent instance above self-review instructions and extended thinking.
  • Split by file, then integrate. Per-file passes find local issues; a separate pass finds cross-file issues. Confirm every file was reviewed before merging the findings.
  • A bigger context window is not the fix. The problem is attention, not space.
  • Confidence routes findings; it does not replace criteria. Check the scores against labelled examples before trusting a threshold.
  • Stakes decide what a person must see. Confidence says how likely a finding is wrong; cost and reversibility say how much that matters.

Where it appears in the exam

Prompt Engineering & Structured Output is a primary domain in two of the six exam scenarios: Claude Code for Continuous Integration and Structured Data Extraction. Multi-pass review fits the CI scenario most closely, where Claude runs automated code reviews and gives feedback on pull requests.

Two sample questions

These are original Timo practice questions. They are not official exam questions.

Question 1

A developer asks Claude Code to write a data migration script, then in the same session asks it to "review the script carefully for bugs". The review rarely finds anything, but QA later finds a bug that deletes rows when a column is empty. What change would most improve the review?

Answer: C. A new instance reviews the script without the generator's assumptions, which is where the missed case came from. A and D keep the same context and blind spots, and B repeats the same flawed review three times.

Question 2

A pull request changes 22 files in an ordering service. Per-file review passes catch local issues well, but no pass flags that create_order now returns amounts in cents while invoice.py still treats them as dollars. What should the team add?

Answer: B. The bug sits between two files, so it needs a pass whose job is cross-file data flow. A brings back the diluted attention the split removed, C still looks at one file at a time, and D moves the review work onto developers.

Build exercise

  1. Ask Claude to write a function from a short spec, then ask the same conversation to review it. Save the findings.
  2. Start a new conversation with only the spec, the code and three review criteria. Compare its findings with step 1.
  3. Take a change that touches about ten files. Run one pass over all of them, then the per-file plus integration review above. Compare depth and contradictions.
  4. Add a confidence to each finding, label 30 findings yourself as right or wrong, and check whether high-confidence findings are right more often before you route on them.

Practise this topic

Sources