TimoBy Amotion AI

Few-shot prompting: CCAR-F task statement 4.2

CCAR-F · Prompt Engineering & Structured Output (20% of the exam)

Task statement 4.2 sits in Prompt Engineering & Structured Output, 20% of the CCAR-F exam. It tests when to stop adding instructions and add a few worked examples instead, and how to pick examples that teach Claude a judgment it can apply to cases you did not show.

What the official guide covers

The Claude Certified Architect Foundations exam guide (version 1.0, effective July 2026) lists this under task statement 4.2, "Apply few-shot prompting to improve output consistency and quality":

Knowledge ofSkills in
Few-shot examples are the most effective way to get consistent, actionable output when detailed instructions alone give inconsistent resultsWriting 2 to 4 targeted examples for ambiguous cases, each showing why one action was chosen over plausible alternatives
Examples show how to handle ambiguous cases, such as which tool fits an unclear request or a gap in branch-level test coverageWriting examples that fix the output format (location, issue, severity, suggested fix)
Examples let Claude apply the same judgment to new patterns, not only to the cases listedWriting examples that separate acceptable code patterns from real issues, to cut false positives while keeping generalisation
Examples reduce hallucination in extraction, for example with informal measurements and varied document structuresWriting examples for varied document structures (inline citations or bibliographies, a methodology section or details spread through the text) so required fields are not returned empty

How examples change the output

Claude reads examples as a pattern to follow. Two things transfer from them:

  • The format. If every example shows a finding as location, issue, severity and suggested fix, Claude writes findings that way. Anthropic's prompting guide calls examples one of the most reliable ways to steer format, tone and structure.
  • The decision rule. If one example flags a pattern and another accepts a similar one, Claude works out what separates them. That is how examples generalise to cases you did not list.

Anthropic's guide gives three tests for good examples: relevant (close to your real inputs), diverse (covering edge cases, varied enough that Claude does not copy an accidental pattern) and structured (wrapped in <example> tags so Claude can tell them apart from instructions). It suggests 3 to 5 examples for general format steering. The exam guide's skill line is narrower: 2 to 4 targeted examples for the ambiguous cases. Both point the same way: a few examples, chosen with care.

Choose the hard cases and show the reasoning

An example of an obvious case teaches little, because Claude already handles it. Spend examples on the boundary: requests where two tools both look right, code that looks risky but is acceptable in your codebase, documents where the data sits in an unusual place.

Each example should say why the chosen action beat the plausible alternative. A bare label ("not a finding") shows the answer. A reason ("not a finding: the broad except is at the top of a worker loop and logs the full traceback") shows the rule, and the rule is what Claude carries to the next case.

A few-shot block for code review findings

Report each finding in exactly this format. The examples show where the
line falls between a real issue and an acceptable pattern.

<examples>
<example>
<code>
def get_user(user_id):
    return db.execute(f"SELECT * FROM users WHERE id = {user_id}")
</code>
<finding>
location: users.py, get_user
issue: SQL is built from a parameter, so user_id can inject SQL.
severity: critical
suggested_fix: db.execute("SELECT * FROM users WHERE id = %s", (user_id,))
</finding>
</example>

<example>
<code>
while True:
    try:
        handle(queue.get())
    except Exception:
        log.exception("job failed")
</code>
<finding>
none. A broad except is acceptable here: it is the top of a worker loop,
it logs the full traceback, and stopping the worker would drop every
later job. Flag broad excepts only when they hide the error.
</finding>
</example>

<example>
<code>
def parse_amount(text):
    try:
        return Decimal(text)
    except Exception:
        return None
</code>
<finding>
location: billing.py, parse_amount
issue: Every parse error becomes None with no log, so bad input is
silently treated as a missing amount.
severity: major
suggested_fix: Catch InvalidOperation only and log the rejected text.
</finding>
</example>
</examples>

The second and third examples look alike on the surface. Together they teach the actual rule (a broad except is a problem when it hides the error), so Claude can apply it to a broad except it has never seen.

Examples for extraction from varied documents

The same method fixes extraction gaps. If papers with a numbered bibliography extract well but papers with inline author-year citations come back with an empty citations list, the instructions are not the problem: Claude has not seen what a citation looks like in the second layout. Add one example per layout, each showing the correct output. The same applies to informal values: one example that turns "about two and a half metres" into a number with a unit shows the rule more clearly than a paragraph of instructions.

Where good examples come from

  • From your evaluation runs. Outputs that scored highest on your test set are real inputs with checked answers. Use them, and add a line after each saying why it is right.
  • Balanced across the cases you will see. If all three examples are findings, Claude leans towards reporting. Include "no finding" cases too. An unbalanced set steers the output in a way no single response shows; it only appears across many runs.
  • Kept up to date. When the team's conventions change, an old example keeps teaching the old rule. Review examples when you review the criteria.
  • Checked by Claude. Anthropic's prompting guide suggests asking Claude to judge your examples for relevance and variety, or to draft more from your first few. A person still confirms each answer.

When to leave examples out

Start without examples. If the task is well specified (sort tickets into five named categories) and the output already passes your tests, examples add tokens and upkeep for no gain. Add examples when showing the format or the judgment is easier than describing it. Ask for step-by-step reasoning instead when the answer depends on several conditions that interact.

Examples are part of the prompt, not training. They are sent, and counted as input tokens, on every call, and they stop working the moment you remove them. Nothing is learned permanently. That is also why the right set can change when you change models: re-run your tests on the new model before you keep, add or drop examples.

SituationChooseWhy
Task is clearly specified and output already passes your testsNo examplesNothing to fix; examples add cost and upkeep
Output format drifts even though the prompt describes it2 or 3 examples in the exact formatExamples steer format more reliably than descriptions
Clear cases are handled well, borderline cases splitExamples of the borderline cases, each with the reasonTeaches the decision, not just the answer
The reviewer flags a pattern the team acceptsAn example of that pattern marked as no finding, with whyCuts false positives and still generalises
Required fields come back empty for one document layoutOne example from each layoutShows where the data sits
Output must match a JSON schema every timeStructured output with a schema (4.3)Examples improve consistency; they do not guarantee it
A rule must hold on every callCode, such as a hook or validationExamples are guidance, not enforcement

Rules that decide exam answers

  • When format drifts despite clear instructions, add examples. More paragraphs of instructions is the tempting wrong answer.
  • Spend examples on ambiguous cases. Ten easy examples teach less than three borderline ones.
  • Show the reasoning, not just the label. The reason is what lets Claude handle new patterns.
  • Vary and balance the examples. Near-identical examples teach Claude to copy surface details, and a set with only findings pushes it towards reporting.
  • Do not force a value to fix an empty field. Making a field required or adding "never leave this empty" pushes Claude to invent data; an example from the missed layout fixes the cause.
  • Examples guide; they do not train. An option claiming that examples permanently teach the model, or that they make calls cheaper, is wrong: they cost tokens on every call and act only while present.

Where it appears in the exam

Prompt Engineering & Structured Output is a primary domain in two of the six exam scenarios: Claude Code for Continuous Integration and Structured Data Extraction. Few-shot examples for review findings and false positives fit the CI scenario. Examples for varied document formats fit the extraction scenario.

Two sample questions

These are original Timo practice questions. They are not official exam questions.

Question 1

A support agent has two tools, search_help_articles and open_support_ticket, and both have detailed descriptions. Requests such as "my CSV export keeps failing, can someone look at it?" go to one tool on some days and the other on others. The system prompt already has two paragraphs on when to use each tool. What should the team add?

Answer: D. Targeted examples at the boundary, with the reasoning, teach the judgment Claude is missing. A adds more of the instructions that already fail, B removes Claude's understanding of the request, and C spends examples on cases Claude already gets right.

Question 2

An extraction pipeline pulls citations from research papers. Papers with a numbered bibliography extract correctly, but papers that use inline author-year citations return an empty citations list, although the citations are there. The prompt already says "extract every citation". What is the best fix?

Answer: B. An example of each layout shows Claude what an inline citation looks like and where to find it. A and C push Claude to invent citations for papers that truly have none, and D repeats a prompt that has already failed.

Build exercise

  1. Write a review prompt that only describes the output format. Run it on five code samples and note how the format and the decisions vary.
  2. Add three examples in <example> tags: one real issue in full format, one acceptable pattern marked as no finding with the reason, and one borderline case.
  3. Run the same five samples again, then one sample with a pattern none of the examples shows. Check that Claude applies the rule, not just the examples.
  4. Take six documents in two layouts. Count the empty required fields, add one example per layout, and count again.

Practise this topic

Sources