Prompts with explicit criteria: CCAR-F task statement 4.1
CCAR-F · Prompt Engineering & Structured Output (20% of the exam)
Task statement 4.1 sits in Prompt Engineering & Structured Output, 20% of the CCAR-F exam. It tests one skill: replacing vague instructions such as "be conservative" with criteria that say exactly what Claude should report, what it should skip and how to grade what it reports.
What the official guide covers
The Claude Certified Architect Foundations exam guide (version 1.0, effective July 2026) lists this under task statement 4.1, "Design prompts with explicit criteria to improve precision and reduce false positives":
| Knowledge of | Skills in |
|---|---|
| Explicit criteria work better than vague instructions: "flag a comment only when it contradicts what the code does" beats "check that comments are accurate" | Writing review criteria that name what to report (bugs, security) and what to skip (minor style, local patterns), instead of filtering by confidence |
| General instructions such as "be conservative" or "only report high-confidence findings" do not raise precision the way specific categories do | Turning off categories with many false positives for a while, to win back developer trust while you fix the prompts for those categories |
| False positives in one category make developers distrust the accurate categories too | Defining each severity level with concrete code examples so findings are classified the same way every time |
Why "be conservative" does not work
A vague instruction leaves the boundary to Claude. "Only report high-confidence findings" does not say which kinds of issue count, so each run draws the line in a different place. Fewer comments appear, but wrong ones still get through and right ones get cut.
Anthropic's prompting guide describes Claude as a brilliant new employee who does not know your norms. A new reviewer told to "be conservative" would ask: conservative about what? Explicit criteria answer that question in the prompt:
- Categories to report, each with the test that makes something reportable.
- Categories to skip, named, so Claude does not have to guess.
- Severity levels, each defined by an example.
- The reason for the rule. The prompting guide notes that explaining why an instruction matters helps Claude apply it to cases you did not list.
Anthropic's guide to evaluations makes the same point about success criteria: they should be specific and measurable, such as "accurate sentiment classification" rather than "good performance". A review prompt is a set of success criteria for each comment.
Criteria say what counts as a finding. For reviews that need judgment, you can also add process steps that say how to check: "first trace where the value comes from, then decide whether it is trusted input". Steps help when a finding depends on several conditions. Criteria are needed in every review prompt.
Find what the prompt leaves unsaid before adding words
When a review prompt misfires, the reflex is to add another paragraph or repeat the instruction in capitals. Usually the prompt is silent on something, and Claude fills that silence with its own guess, a different guess on different runs. Find the gap first, then add the one piece that fills it.
| What you see | What the prompt does not say | What to add |
|---|---|---|
| Comments on things the team does not care about | Which categories are out of scope | A skip list |
| The same issue gets different labels across runs | What each label means | Severity definitions with examples |
| Right findings in a shape the parser rejects | The exact output format | An output contract or a schema (4.3) |
| Correct on common code, wrong on one unusual pattern | A rule for that pattern | A criterion that names it, or an example of it (4.2) |
If three rewrites in a row have not helped, stop adding text and check which of these pieces is missing.
Vague instruction or explicit criterion?
| Vague | Explicit |
|---|---|
| "Check that comments are accurate." | "Flag a comment only when it describes behaviour the code does not have, for example a docstring that says the function returns None when it raises an exception." |
| "Be conservative." | "Report bugs that change behaviour and security issues. Skip naming, formatting and patterns already used elsewhere in this repository." |
| "Only report high-confidence findings." | "Report a possible null value only when you can name the code path where it is null." |
| "Rate the severity." | "Use critical, major or minor as defined below, with one code example for each." |
A review prompt with report, skip and severity rules
This is the kind of system prompt the exam expects in a CI review scenario. Each rule is a category or a test, not a level of confidence.
You review pull requests for a Python payments service. Developers read every
comment you post, so a wrong comment costs them time and trust.
<report>
- Bugs: code whose behaviour differs from what the function name, docstring
or tests say it should do.
- Security: SQL built from strings, missing permission checks on an endpoint,
secrets or tokens written in code or logs.
- Comments that contradict the code: only when the comment describes
behaviour the code does not have.
</report>
<skip>
- Naming, formatting and import order (the linter handles these).
- Patterns used the same way in at least two other files in this repository.
- Suggestions to add comments or docstrings.
</skip>
<severity>
critical: can lose money or expose data in production.
Example: query = f"SELECT * FROM payments WHERE id = {payment_id}"
major: wrong result for some valid inputs, no data exposure.
Example: total = sum(items[1:]) # skips the first item
minor: correct today, likely to break on a reasonable change.
Example: catching Exception and returning None in a parser
</severity>
If nothing meets the report criteria, say "No findings." Do not add
low-severity findings to fill the review.
False positives and developer trust
Developers do not judge a review tool category by category. If most comments about performance turn out to be wrong, they start skipping the whole review, including the one comment about a real security hole. That is why the guide treats precision as a trust problem, not only a quality score.
The fix is to measure each category separately. Track which findings developers accept and which they dismiss. When one category is mostly dismissed, turn it off, keep the accurate categories running, rewrite the criteria for the noisy one and test the new version on past pull requests with known outcomes before you turn it back on.
Change one thing at a time when you test, and read the results per category, not only the overall rate. An average can stay flat while one category gets better and another gets worse.
| Situation | Choose | Why |
|---|---|---|
| Reviews are full of style comments | Add a skip list that names style and local patterns | The boundary is now in the prompt, not left to Claude |
| One category is mostly false positives and developers ignore everything | Turn that category off for now, fix its criteria, test, then turn it back on | Protects trust in the accurate categories |
| The same issue gets "major" one day and "minor" the next | Define each severity level with a code example | Claude compares against a fixed reference |
| Someone proposes asking Claude for a confidence score and hiding low scores | Write categorical criteria instead | A confidence score does not define what is in scope |
Rules that decide exam answers
- Categories beat confidence. When an option filters findings by self-reported confidence and another defines what to report and skip, the criteria option is right.
- "Be conservative" is not a criterion. Any option that only makes the instruction louder ("be very conservative", "only report issues you are sure of") leaves the boundary undefined.
- Say what to skip, not only what to report. Naming the categories to ignore is what removes the noise.
- Turn off a noisy category instead of living with it. Keeping a mostly wrong category running damages trust in the categories that work.
- Define severity with examples. A short code example for each level gives consistent labels; adjectives alone do not.
- Find the gap before rewriting. Rewording or repeating an instruction does not add the missing category, definition or format, and neither does moving to a larger model: a more capable model still has to guess at whatever the prompt leaves unsaid.
Where it appears in the exam
Prompt Engineering & Structured Output is a primary domain in two of the six exam scenarios: Claude Code for Continuous Integration and Structured Data Extraction. This task statement fits the CI scenario most closely, because the guide describes it as designing prompts that give actionable feedback and keep false positives low.
Two sample questions
These are original Timo practice questions. They are not official exam questions.
Build exercise
- Collect 10 past pull requests where you know the real issues. Run a review with the prompt "Review this code. Be conservative." Count the findings and how many are false positives.
- Rewrite the prompt with
<report>and<skip>lists like the one above. Run it on the same pull requests and compare false positives per category. - Add severity definitions with one code example each. Run the same pull request three times and check whether each finding gets the same severity every time.
- Find the category with the most false positives. Turn it off, rewrite its criteria, test it on the 10 pull requests, and only then turn it back on.
Practise this topic
- Claude Certified Architect practice exam: free, 20 questions, no sign-up
- Claude Certified Architect hub
- CCAR-F study guide: all topics
- Same topic in another exam: Prompt Engineering (CCDV-F)
- Previous topic: 3.6 Claude Code in CI/CD
- Next topic: 4.2 Few-shot prompting
Sources
- Claude Certified Architect Foundations Exam Guide, version 1.0, effective July 2026 (Anthropic), task statement 4.1
- Anthropic documentation: Prompting best practices
- Anthropic documentation: Define success criteria and build evaluations
By Amotion AI