TimoBy Amotion AI

Iterative refinement: CCAR-F task statement 3.5

CCAR-F · Claude Code Configuration & Workflows (20% of the exam)

Task statement 3.5 sits in Claude Code Configuration & Workflows, 20% of the CCAR-F exam. It tests how you get from a first attempt to a correct result: which kind of feedback to give Claude when a prose instruction is not working, and whether to send several problems together or one at a time.

What the official guide covers

The Claude Certified Architect Foundations exam guide (version 1.0, effective July 2026) lists this under task statement 3.5, "Apply iterative refinement techniques for progressive improvement":

Knowledge ofSkills in
Concrete input/output examples are the clearest way to show a transformation when prose is read inconsistentlyGiving 2 or 3 input/output examples when a prose description produces inconsistent results
Test-driven iteration: write the tests first, then share test failures to guide each roundWriting tests for expected behaviour, edge cases and performance before implementation, then sharing the failures
The interview pattern: Claude asks questions that surface things the developer had not considered, before it buildsUsing the interview pattern in unfamiliar domains to surface design questions such as cache invalidation and failure modes
Interacting problems go in one message; independent problems can be fixed one after anotherGiving specific test cases with an example input and expected output to fix edge cases, such as null values in a migration script
Describing interacting issues together in one detailed message, and working through independent issues in sequence

Show the transformation with examples

A prose instruction such as "normalise the phone numbers" leaves choices open: which format, what to do with extensions, what to do with rubbish input. Claude may choose differently on each run. Two or three examples remove the ambiguity, because they show the exact output for the cases that varied.

Convert phone numbers to E.164. Examples:

Input: "(415) 555-0132", country US   ->  "+14155550132"
Input: "020 7946 0018", country GB    ->  "+442079460018"
Input: "ext. 22 only", country US     ->  null  (no full number: return null, do not guess)

Pick examples from the cases where the output was inconsistent, and include one edge case. Anthropic's prompting guidance says good examples are relevant to the real task and varied enough that Claude does not copy an unintended pattern. For designing few-shot prompts in production, see 4.2 Few-shot prompting.

Test-driven iteration

A test gives Claude a pass or fail signal it can read itself. Claude writes code, runs the check, reads the result and tries again until it passes. You stop being the only check.

  1. Write the tests first: expected behaviour, edge cases and any performance limit.
  2. Run them and confirm they fail.
  3. Ask Claude to make them pass and to run them after each change. Tell it not to edit the tests.
  4. When something still fails, paste the exact failure output. Do not describe it in your own words.
# test_migrate_customers.py, written before migrate_customers.py exists
import pytest
from migrate_customers import migrate_row

def test_full_row_is_normalised():
    row = {"id": 7, "email": "A@Example.com", "phone": "(415) 555-0132", "country": "US"}
    assert migrate_row(row) == {"id": 7, "email": "a@example.com", "phone": "+14155550132"}

def test_null_phone_stays_null():
    row = {"id": 8, "email": "b@example.com", "phone": None, "country": "US"}
    assert migrate_row(row)["phone"] is None

def test_missing_email_raises():
    with pytest.raises(ValueError):
        migrate_row({"id": 9, "email": None, "phone": None, "country": "US"})

The second test is the guide's "specific test case for an edge case". "Handle nulls properly" is vague; {"phone": None} must give "phone": None is not.

A useful order when Claude writes the tests too: show it the relevant files, ask it to list the test cases it would write, choose the ones that matter, have it write those, and only then ask for the code.

Make the test loop run without reminders

"Run the tests after each change" is a request. On a long run, Claude may skip it, or report success without having run anything. Two mechanisms turn the request into a gate.

A Stop hook. The Stop event fires when Claude tries to end its turn. Exit code 2 refuses: Claude keeps working and receives your stderr text as the reason. Exit code 1 does not block, so a hook that exits 1 on failure lets Claude stop anyway.

#!/usr/bin/env bash
# .claude/hooks/tests-must-pass.sh
if ! out=$(python -m pytest -q 2>&1); then
  echo "Tests are failing. Fix the code, do not edit the tests:" >&2
  echo "$out" | tail -n 30 >&2
  exit 2
fi
exit 0
{
  "hooks": {
    "Stop": [
      { "hooks": [ { "type": "command", "command": "\"$CLAUDE_PROJECT_DIR\"/.claude/hooks/tests-must-pass.sh" } ] }
    ]
  }
}

Put the JSON in .claude/settings.json so the whole team gets the gate. The hook input includes stop_hook_active, which your script can check if you want it to give up after one retry.

A goal. /goal all tests in tests/billing pass and no file under tests/ is modified keeps Claude working across turns until a separate model judges the condition met. That evaluator reads only the conversation; it does not run commands. Write the condition so Claude's own output proves it, such as a test run that appears in the transcript.

Green tests are not the whole check. A test can be loosened until it passes. Read the diff of the test files, or make "no test was weakened" part of the condition.

The interview pattern

In an unfamiliar area you do not know which questions matter. Ask Claude to interview you before it writes anything. Claude Code has an AskUserQuestion tool for this.

I want to add a cache in front of the pricing service. Before you write any code,
interview me with the AskUserQuestion tool. Focus on the hard parts: how entries
are invalidated when prices change, what happens when the pricing service is down,
and how stale a price may be. When we have covered everything, write the spec to SPEC.md.

Then start a fresh session and implement from SPEC.md. The new session has clean context and a written spec to check against.

One message or one at a time?

Problems interact when fixing one changes the other. A renamed CSV column and the validator that checks that column are one problem in two places. Fix them separately and each fix can break the other. Send them together, in one detailed message that explains how they connect. Problems that do not touch each other, such as a typo in a log line and a slow query, are easier to fix and check one at a time.

SituationDo thisWhy
Prose instructions give different output on each runAdd 2 or 3 input/output examplesExamples define the format the prose left open
The behaviour can be testedWrite tests first and share the failuresEach round has a clear pass or fail signal
You are new to the domain and do not know the design questionsInterview pattern, then a specSurfaces issues such as invalidation and failure modes before code exists
One edge case keeps failingGive the exact input and expected output as a testA specific case beats "handle edge cases better"
Two issues where fixing one affects the otherOne detailed message covering bothSeparate fixes can undo each other
Unrelated issuesFix them one at a timeEach fix is small and easy to check
Claude says it is done but did not run the testsStop hook that exits 2 on failure, or a /goal with a test conditionThe check runs whether or not anyone asks
Lint or type errors keep appearing after editsPostToolUse hook on edits that runs the checker and exits 2 on errorsClaude reads the errors after each edit; the hook cannot undo the edit, only report it
Claude went down a wrong path a few prompts ago/rewind to the last good checkpoint, then re-promptRemoves the bad attempt instead of arguing it back
You have corrected the same thing twice and it is still wrong/clear and start again with a better promptThe context is full of failed attempts

Rules that decide exam answers

  • Inconsistent output from prose means add examples. Rewriting the prose, adding capitals or saying "be consistent" does not define the missing format.
  • Tests first, then failures as feedback. The exact failing test output is more useful to Claude than a summary of what went wrong.
  • Interview before building in an unfamiliar domain. Let Claude raise the questions you have not thought of, then build from the agreed spec.
  • Interacting issues together; independent issues apart. The deciding question is whether one fix changes the other.
  • Edge cases need concrete cases. Give the input and the expected output, not a general request to be careful.
  • A gate beats a reminder. When tests must run every time, use a Stop hook that exits 2, not a line in the prompt or CLAUDE.md.

Where it appears in the exam

Claude Code Configuration & Workflows is a primary domain in three of the six exam scenarios: Code Generation with Claude Code, Developer Productivity with Claude, and Claude Code for Continuous Integration. Refinement questions fit the Code Generation scenario most closely, where the team uses Claude Code for code generation, refactoring and debugging.

Two sample questions

These are original Timo practice questions. They are not official exam questions.

Question 1

A developer asks Claude Code to "convert our legacy date strings to ISO 8601". Across three runs, Claude reads "03/04/2024" as 3 April once and as 4 March twice, and handles "2024-1-5" differently each time. The developer has already rewritten the instruction in prose twice. What should they do next?

Answer: C. Examples define the transformation that the prose leaves open, including which way to read "03/04/2024". A repeats the vague instruction more loudly. B adds a planning step but still no definition of the format. D finds inconsistent results after the fact instead of removing the cause.

Question 2

Claude is changing an order export. Two problems remain: the CSV header now says quantity instead of qty, and the downstream validator still checks for qty, so every file fails validation. Separately, a log message has a typo. How should the developer give feedback?

Answer: A. The header and the validator interact, so fixing them separately risks one change undoing the other; the typo is independent and easy to check on its own. B and C split the interacting pair. D mixes an unrelated fix into the coupled change, which makes the result harder to check.

Build exercise

  1. Ask Claude to write a slugify(title) function from a one-line prose description. Run it in three fresh sessions and compare the outputs for the same five titles.
  2. Add three input/output examples, including one edge case such as a title with only punctuation. Repeat and compare.
  3. Write the pytest file above for a small migration function, confirm the tests fail, then ask Claude to make them pass without editing the tests. Paste any failure output back exactly.
  4. Start a new feature with the interview prompt and let Claude write SPEC.md. Count the questions it raised that you had not planned for, then implement from the spec in a fresh session.

Practise this topic

Sources