Output Evaluation and Validation: CCAO-F domain 2 study guide
CCAO-F · Output Evaluation and Validation (21% of the exam)
Output Evaluation and Validation is domain 2 of the Claude Certified Associate Foundations exam and, at 21%, the largest domain. It tests one habit: you check Claude's output against a source before anyone relies on it, and decide how much checking it needs.
What the official guide covers
The Claude Certified Associate Foundations exam guide (version 1.0, effective July 2026) lists six tasks under domain 2, Output Evaluation and Validation:
| What the guide lists | What it means in practice |
|---|---|
| Evaluate outputs for accuracy and completeness | Check each claim against the source and confirm every part of the request was answered |
| Identify hallucinations, inconsistencies and biases | Spot invented details, statements that contradict the source or each other, and one-sided framing |
| Apply fact-checking and validation techniques | Open the cited source, ask for supporting quotes, recheck numbers, compare with an independent source |
| Determine when human review or more verification is needed | Match the level of review to the stakes and the audience |
| Edit, adapt, refine and compare outputs for the intended audience | Rewrite for the reader, and compare two versions against stated criteria |
| Organise information and choose output formats (artifacts, inline, structured data) | Decide whether the result belongs in the chat, in an artifact or in a table or data format |
Check against three references, every time
Compare the output with three fixed references:
- Your request. Tick off each part you asked for, not only the easy ones.
- The source material. Trace specific claims back to the documents you supplied.
- The standards of your field. A figure with no unit or a citation you cannot find fails, however fluent it reads.
Review accuracy and completeness separately. An output can be correct line by line and still omit the factor that changes the decision; gaps are harder to spot.
Then give the output one of three verdicts and note why:
| Verdict | When it applies | Example |
|---|---|---|
| Ready to use | Meets the request, matches the sources, clears your field's standards, and the stakes allow it | Three options to shorten an internal approval process, for a team discussion |
| Needs revision | Close, with a specific gap you can name | A price summary that dropped the minimum order quantity shown in the supplier's PDF; re-prompt from the source |
| Needs a person | High stakes, errors or uncertainty mean it cannot go out on Claude's draft | A gap analysis against a regulation that was never uploaded, so Claude worked from memory |
Decide the stakes first, then the depth of review; over-checking a low-stakes draft wastes the time saved. Polish does not change the stakes: a clean draft bound for a regulator needs a person, while a rough internal brainstorm needs a light edit.
Anthropic's AI Fluency framework calls this skill Discernment and asks you to judge three things: the product (is the output accurate and fit for use), the process (did Claude's reasoning skip a step or rest on a weak assumption) and the performance (did it behave as asked, for example flagging gaps instead of filling them).
What goes wrong in Claude's output
Anthropic's help centre warns that Claude can hallucinate details that look authoritative, and should not be your single source of truth for high-stakes advice. Look for these failures:
| Failure | What it looks like | How you catch it |
|---|---|---|
| Invented specifics | A clause number, statistic, quote, name or link that does not exist | Find each one in the source or open the link |
| Out-of-date facts | A rule, price or org chart that has since changed | Check the date and version of your source |
| Claimed actions | "I have emailed the supplier" when no email tool was connected | Check the connected tools; without a connector, nothing was sent |
| Inconsistency | The summary says one thing, the table says another | Compare every number and claim across sections |
| Omission | Three of the four questions answered, or only the first half of a document used | Tick off each part of the original request |
| Misreading a source | A web result summarised without the context that changes its meaning | Open the original page, as the help centre advises |
| Bias | One side's view only, stereotyped examples, a recommendation that favours one group without reason | Ask who is missing and whether the evidence supports the framing |
| Agreeing with your framing | You asked "why is supplier A the best choice?" and Claude found reasons | Ask the question neutrally, or ask for the case against |
| Contradiction across a long document | Page 2 gives one market size, page 8 builds on another | Do a separate pass that compares every figure across the whole document |
Be most suspicious of precision without a source and of certainty on questions that should be hedged, such as legal or date-sensitive ones.
Fact-checking techniques that work
- Trace to the source. For every figure, date, name and citation, find the place it came from. If you cannot find it, treat it as unverified.
- Ask for quotes first. Ask Claude to quote the passage that supports each claim, then check that the quote exists in the document. Anthropic's prompting guidance recommends grounding answers in quotes for long documents.
- Recalculate. Add up totals, recheck percentages and confirm that comparisons ("fastest growth", "lowest cost") match the data.
- Use a second source. For external or senior audiences, confirm key facts independently. A second answer from Claude is not an independent source.
- Run it twice and compare. Where two runs agree, confidence rises; where they differ, you have found the claims that need a person to check.
Build the checks into the prompt
Checking is cheaper when the prompt makes errors easy to find. Anthropic's guidance recommends letting Claude say it does not know and grounding answers in quotes:
Answer only from the attached supplier contract. Do not use general
knowledge.
Before answering, quote the clauses that bear on each question.
After each claim, give the clause number in brackets.
If the contract does not cover a question, list it under
"Not covered" instead of answering it.
Every claim now points to a place you can check. Curate the inputs too: remove duplicate copies, label each file's role, and leave out what the question does not need. A muddled summary from near-duplicate drafts is fixed by curating the files, not by a bigger model.
Review checklist
Use this before anything Claude helped write goes out.
OUTPUT REVIEW CHECKLIST
Task: ____________ Reader: ____________ Stakes: low / medium / high
Accuracy
[ ] Every figure, date and name found in the source
[ ] Every citation or link opened and checked
[ ] Totals and comparisons recalculated
[ ] No claim that an action was taken unless a connected tool did it
Completeness
[ ] Every part of the request answered
[ ] Gaps marked, not filled with guesses
Consistency and balance
[ ] Summary matches the body and any tables
[ ] Other viewpoints or affected groups considered
[ ] Wording does not stereotype or favour a group without reason
Fit for the reader
[ ] Tone, length and terms suit the audience
[ ] Format suits where it will be used (email, slide, sheet)
Sign-off
[ ] Human review level matches the stakes table
[ ] Reviewer name: ____________
When a person must review
| Situation | Review needed | Why |
|---|---|---|
| Internal brainstorm or first draft for yourself | Your own read-through | Low stakes; you will rework it anyway |
| Content for customers, the public or senior leaders | Full fact-check and owner sign-off | Errors reach people who act on them |
| Legal, financial, medical, HR or regulatory content | Review by a qualified specialist | Claude does not replace professional judgement |
| Decisions that affect a specific person | A person makes and owns the decision | Accountability stays with people |
Four questions set the level: how costly is an error, can it be undone, who will see it, and does a rule govern it? A board pack that reads cleanly still trips three. Fix your never-send-without-review list in advance: client deliverables, reported figures, regulated data, public or legal statements.
Iterating is not escalating. If rounds of prompting have stopped improving a high-stakes draft, another prompt will not supply the missing judgement; pass it to a colleague or specialist.
Adapt and compare for the reader
Before, a technical summary written for the IT team:
The SSO migration completed on schedule. 4% of users hit SAML assertion
errors due to clock skew on legacy IdP nodes; remediation is in progress.
After, the same facts for all staff:
The new single sign-on is live. A small number of people may see a login
error this week. If you do, wait five minutes and try again. IT is fixing
the cause.
The facts stayed; the vocabulary, detail and call to action changed. When comparing two versions, score them against criteria set in advance, not on which reads best.
Choosing the output format
| The output is | Use | Why |
|---|---|---|
| A short answer or a few lines you will read once | Inline in the chat | Nothing to reuse or share |
| A document, slide outline, diagram or small tool you will edit or share | An artifact | Substantial, self-contained content shown beside the chat, private by default and publishable |
| Data going into a spreadsheet or another system | A table or structured data such as CSV | Columns and rows paste cleanly and can be checked line by line |
| Totals, percentages or charts that will be reported | Code execution: Claude writes and runs code on your file in a sandbox | The figure is calculated from the rows, not written as likely-looking text |
Code execution makes a figure traceable, not automatically right: Claude writes the code, so spot-check what was calculated. When quality matters, compare two drafts and edit the stronger one.
Rules that decide exam answers
- Verify against the source, not against Claude. A confident tone or a self-rated confidence score is not evidence; asking for one is the tempting wrong option.
- Specific details are the riskiest. Citation numbers, statistics, quotes and links are the most likely to be invented, so check those first.
- Review scales with the stakes, not the polish. Legal, financial or external content gets specialist review however clean it looks; low-risk internal drafts do not need full review.
- Rewording is not checking. Making an output more formal, shorter or clearer does nothing for its accuracy.
- A person owns the result. Claude can draft and summarise; a named person approves what goes out and any decision about people.
- Compute figures that matter. Totals and percentages that will be reported come from code execution or a recalculation, not from prose.
Where it appears in the exam
Domain 2 carries 21% of the exam, the highest weight, so expect around 12 or 13 of the 60 items. Questions describe an output about to be shared and ask what to check, what the error is, how much review it needs, or which format suits the reader.
Two sample questions
These are original Timo practice questions. They are not official exam questions.
Build exercise
- Ask Claude to summarise a document you know well. Run the review checklist above and record every error.
- Ask Claude to quote the passage behind each claim in the summary. Check that each quote exists word for word.
- Ask for one message in versions for two audiences. Score each against four criteria written down first.
Practise this topic
- Claude Certified Associate practice exam: free, 20 questions, no sign-up
- CCAO-F study guide: all topics
- Worked example: Professional output validation
- Same topic in another exam: Output Handling (CCDV-F), the developer angle
- Previous topic: Prompting and Task Execution
- Next topic: Product and Model Selection
Sources
- Claude Certified Associate Foundations Exam Guide, version 1.0, effective July 2026 (Anthropic), domain 2: Output Evaluation and Validation
- Claude Help Center: Claude is providing incorrect or misleading responses. What's going on?
- Claude Help Center: Claude is producing links that don't work and falsely claiming that it has sent emails or produced external documents
- Claude Help Center: What are artifacts and how do I use them?
- Anthropic documentation: Prompting best practices
By Amotion AI