TimoBy Amotion AI

Prompts with explicit criteria: CCAR-F task statement 4.1

CCAR-F · Prompt Engineering & Structured Output (20% of the exam)

Task statement 4.1 sits in Prompt Engineering & Structured Output, 20% of the CCAR-F exam. It tests one skill: replacing vague instructions such as "be conservative" with criteria that say exactly what Claude should report, what it should skip and how to grade what it reports.

What the official guide covers

The Claude Certified Architect Foundations exam guide (version 1.0, effective July 2026) lists this under task statement 4.1, "Design prompts with explicit criteria to improve precision and reduce false positives":

Knowledge ofSkills in
Explicit criteria work better than vague instructions: "flag a comment only when it contradicts what the code does" beats "check that comments are accurate"Writing review criteria that name what to report (bugs, security) and what to skip (minor style, local patterns), instead of filtering by confidence
General instructions such as "be conservative" or "only report high-confidence findings" do not raise precision the way specific categories doTurning off categories with many false positives for a while, to win back developer trust while you fix the prompts for those categories
False positives in one category make developers distrust the accurate categories tooDefining each severity level with concrete code examples so findings are classified the same way every time

Why "be conservative" does not work

A vague instruction leaves the boundary to Claude. "Only report high-confidence findings" does not say which kinds of issue count, so each run draws the line in a different place. Fewer comments appear, but wrong ones still get through and right ones get cut.

Anthropic's prompting guide describes Claude as a brilliant new employee who does not know your norms. A new reviewer told to "be conservative" would ask: conservative about what? Explicit criteria answer that question in the prompt:

  • Categories to report, each with the test that makes something reportable.
  • Categories to skip, named, so Claude does not have to guess.
  • Severity levels, each defined by an example.
  • The reason for the rule. The prompting guide notes that explaining why an instruction matters helps Claude apply it to cases you did not list.

Anthropic's guide to evaluations makes the same point about success criteria: they should be specific and measurable, such as "accurate sentiment classification" rather than "good performance". A review prompt is a set of success criteria for each comment.

Criteria say what counts as a finding. For reviews that need judgment, you can also add process steps that say how to check: "first trace where the value comes from, then decide whether it is trusted input". Steps help when a finding depends on several conditions. Criteria are needed in every review prompt.

Find what the prompt leaves unsaid before adding words

When a review prompt misfires, the reflex is to add another paragraph or repeat the instruction in capitals. Usually the prompt is silent on something, and Claude fills that silence with its own guess, a different guess on different runs. Find the gap first, then add the one piece that fills it.

What you seeWhat the prompt does not sayWhat to add
Comments on things the team does not care aboutWhich categories are out of scopeA skip list
The same issue gets different labels across runsWhat each label meansSeverity definitions with examples
Right findings in a shape the parser rejectsThe exact output formatAn output contract or a schema (4.3)
Correct on common code, wrong on one unusual patternA rule for that patternA criterion that names it, or an example of it (4.2)

If three rewrites in a row have not helped, stop adding text and check which of these pieces is missing.

Vague instruction or explicit criterion?

VagueExplicit
"Check that comments are accurate.""Flag a comment only when it describes behaviour the code does not have, for example a docstring that says the function returns None when it raises an exception."
"Be conservative.""Report bugs that change behaviour and security issues. Skip naming, formatting and patterns already used elsewhere in this repository."
"Only report high-confidence findings.""Report a possible null value only when you can name the code path where it is null."
"Rate the severity.""Use critical, major or minor as defined below, with one code example for each."

A review prompt with report, skip and severity rules

This is the kind of system prompt the exam expects in a CI review scenario. Each rule is a category or a test, not a level of confidence.

You review pull requests for a Python payments service. Developers read every
comment you post, so a wrong comment costs them time and trust.

<report>
- Bugs: code whose behaviour differs from what the function name, docstring
  or tests say it should do.
- Security: SQL built from strings, missing permission checks on an endpoint,
  secrets or tokens written in code or logs.
- Comments that contradict the code: only when the comment describes
  behaviour the code does not have.
</report>

<skip>
- Naming, formatting and import order (the linter handles these).
- Patterns used the same way in at least two other files in this repository.
- Suggestions to add comments or docstrings.
</skip>

<severity>
critical: can lose money or expose data in production.
  Example: query = f"SELECT * FROM payments WHERE id = {payment_id}"
major: wrong result for some valid inputs, no data exposure.
  Example: total = sum(items[1:])   # skips the first item
minor: correct today, likely to break on a reasonable change.
  Example: catching Exception and returning None in a parser
</severity>

If nothing meets the report criteria, say "No findings." Do not add
low-severity findings to fill the review.

False positives and developer trust

Developers do not judge a review tool category by category. If most comments about performance turn out to be wrong, they start skipping the whole review, including the one comment about a real security hole. That is why the guide treats precision as a trust problem, not only a quality score.

The fix is to measure each category separately. Track which findings developers accept and which they dismiss. When one category is mostly dismissed, turn it off, keep the accurate categories running, rewrite the criteria for the noisy one and test the new version on past pull requests with known outcomes before you turn it back on.

Change one thing at a time when you test, and read the results per category, not only the overall rate. An average can stay flat while one category gets better and another gets worse.

SituationChooseWhy
Reviews are full of style commentsAdd a skip list that names style and local patternsThe boundary is now in the prompt, not left to Claude
One category is mostly false positives and developers ignore everythingTurn that category off for now, fix its criteria, test, then turn it back onProtects trust in the accurate categories
The same issue gets "major" one day and "minor" the nextDefine each severity level with a code exampleClaude compares against a fixed reference
Someone proposes asking Claude for a confidence score and hiding low scoresWrite categorical criteria insteadA confidence score does not define what is in scope

Rules that decide exam answers

  • Categories beat confidence. When an option filters findings by self-reported confidence and another defines what to report and skip, the criteria option is right.
  • "Be conservative" is not a criterion. Any option that only makes the instruction louder ("be very conservative", "only report issues you are sure of") leaves the boundary undefined.
  • Say what to skip, not only what to report. Naming the categories to ignore is what removes the noise.
  • Turn off a noisy category instead of living with it. Keeping a mostly wrong category running damages trust in the categories that work.
  • Define severity with examples. A short code example for each level gives consistent labels; adjectives alone do not.
  • Find the gap before rewriting. Rewording or repeating an instruction does not add the missing category, definition or format, and neither does moving to a larger model: a more capable model still has to guess at whatever the prompt leaves unsaid.

Where it appears in the exam

Prompt Engineering & Structured Output is a primary domain in two of the six exam scenarios: Claude Code for Continuous Integration and Structured Data Extraction. This task statement fits the CI scenario most closely, because the guide describes it as designing prompts that give actionable feedback and keep false positives low.

Two sample questions

These are original Timo practice questions. They are not official exam questions.

Question 1

A CI review bot posts about 25 comments on each pull request, most of them about naming and comment wording. Developers now dismiss all of its comments, including a correct one about SQL built from user input. The prompt says "Review this PR carefully. Be conservative and only report issues you are confident about." What change will most improve precision?

Answer: C. Explicit report and skip categories define the boundary, so the style noise stops and the security finding stays. A repeats the vague instruction more strongly, B is the confidence-based filtering the guide says does not work, and D keeps the same undefined boundary.

Question 2

A review tool reports findings in five categories. Developers accept most security and correctness findings, but dismiss most performance findings, and a survey shows they have started skipping whole reviews. Rewriting the performance criteria will take a week. What should the team do now?

Answer: A. Removing the noisy category restores trust in the four accurate ones while the fix is built. B still posts the wrong findings, C throws away the categories that work, and D relies on Claude's own label instead of better criteria.

Build exercise

  1. Collect 10 past pull requests where you know the real issues. Run a review with the prompt "Review this code. Be conservative." Count the findings and how many are false positives.
  2. Rewrite the prompt with <report> and <skip> lists like the one above. Run it on the same pull requests and compare false positives per category.
  3. Add severity definitions with one code example each. Run the same pull request three times and check whether each finding gets the same severity every time.
  4. Find the category with the most false positives. Turn it off, rewrite its criteria, test it on the 10 pull requests, and only then turn it back on.

Practise this topic

Sources