TimoBy Amotion AI

Output Evaluation and Validation: CCAO-F domain 2 study guide

CCAO-F · Output Evaluation and Validation (21% of the exam)

Output Evaluation and Validation is domain 2 of the Claude Certified Associate Foundations exam and, at 21%, the largest domain. It tests one habit: you check Claude's output against a source before anyone relies on it, and decide how much checking it needs.

What the official guide covers

The Claude Certified Associate Foundations exam guide (version 1.0, effective July 2026) lists six tasks under domain 2, Output Evaluation and Validation:

What the guide listsWhat it means in practice
Evaluate outputs for accuracy and completenessCheck each claim against the source and confirm every part of the request was answered
Identify hallucinations, inconsistencies and biasesSpot invented details, statements that contradict the source or each other, and one-sided framing
Apply fact-checking and validation techniquesOpen the cited source, ask for supporting quotes, recheck numbers, compare with an independent source
Determine when human review or more verification is neededMatch the level of review to the stakes and the audience
Edit, adapt, refine and compare outputs for the intended audienceRewrite for the reader, and compare two versions against stated criteria
Organise information and choose output formats (artifacts, inline, structured data)Decide whether the result belongs in the chat, in an artifact or in a table or data format

Check against three references, every time

Compare the output with three fixed references:

  1. Your request. Tick off each part you asked for, not only the easy ones.
  2. The source material. Trace specific claims back to the documents you supplied.
  3. The standards of your field. A figure with no unit or a citation you cannot find fails, however fluent it reads.

Review accuracy and completeness separately. An output can be correct line by line and still omit the factor that changes the decision; gaps are harder to spot.

Then give the output one of three verdicts and note why:

VerdictWhen it appliesExample
Ready to useMeets the request, matches the sources, clears your field's standards, and the stakes allow itThree options to shorten an internal approval process, for a team discussion
Needs revisionClose, with a specific gap you can nameA price summary that dropped the minimum order quantity shown in the supplier's PDF; re-prompt from the source
Needs a personHigh stakes, errors or uncertainty mean it cannot go out on Claude's draftA gap analysis against a regulation that was never uploaded, so Claude worked from memory

Decide the stakes first, then the depth of review; over-checking a low-stakes draft wastes the time saved. Polish does not change the stakes: a clean draft bound for a regulator needs a person, while a rough internal brainstorm needs a light edit.

Anthropic's AI Fluency framework calls this skill Discernment and asks you to judge three things: the product (is the output accurate and fit for use), the process (did Claude's reasoning skip a step or rest on a weak assumption) and the performance (did it behave as asked, for example flagging gaps instead of filling them).

What goes wrong in Claude's output

Anthropic's help centre warns that Claude can hallucinate details that look authoritative, and should not be your single source of truth for high-stakes advice. Look for these failures:

FailureWhat it looks likeHow you catch it
Invented specificsA clause number, statistic, quote, name or link that does not existFind each one in the source or open the link
Out-of-date factsA rule, price or org chart that has since changedCheck the date and version of your source
Claimed actions"I have emailed the supplier" when no email tool was connectedCheck the connected tools; without a connector, nothing was sent
InconsistencyThe summary says one thing, the table says anotherCompare every number and claim across sections
OmissionThree of the four questions answered, or only the first half of a document usedTick off each part of the original request
Misreading a sourceA web result summarised without the context that changes its meaningOpen the original page, as the help centre advises
BiasOne side's view only, stereotyped examples, a recommendation that favours one group without reasonAsk who is missing and whether the evidence supports the framing
Agreeing with your framingYou asked "why is supplier A the best choice?" and Claude found reasonsAsk the question neutrally, or ask for the case against
Contradiction across a long documentPage 2 gives one market size, page 8 builds on anotherDo a separate pass that compares every figure across the whole document

Be most suspicious of precision without a source and of certainty on questions that should be hedged, such as legal or date-sensitive ones.

Fact-checking techniques that work

  1. Trace to the source. For every figure, date, name and citation, find the place it came from. If you cannot find it, treat it as unverified.
  2. Ask for quotes first. Ask Claude to quote the passage that supports each claim, then check that the quote exists in the document. Anthropic's prompting guidance recommends grounding answers in quotes for long documents.
  3. Recalculate. Add up totals, recheck percentages and confirm that comparisons ("fastest growth", "lowest cost") match the data.
  4. Use a second source. For external or senior audiences, confirm key facts independently. A second answer from Claude is not an independent source.
  5. Run it twice and compare. Where two runs agree, confidence rises; where they differ, you have found the claims that need a person to check.

Build the checks into the prompt

Checking is cheaper when the prompt makes errors easy to find. Anthropic's guidance recommends letting Claude say it does not know and grounding answers in quotes:

Answer only from the attached supplier contract. Do not use general
knowledge.
Before answering, quote the clauses that bear on each question.
After each claim, give the clause number in brackets.
If the contract does not cover a question, list it under
"Not covered" instead of answering it.

Every claim now points to a place you can check. Curate the inputs too: remove duplicate copies, label each file's role, and leave out what the question does not need. A muddled summary from near-duplicate drafts is fixed by curating the files, not by a bigger model.

Review checklist

Use this before anything Claude helped write goes out.

OUTPUT REVIEW CHECKLIST
Task: ____________  Reader: ____________  Stakes: low / medium / high

Accuracy
[ ] Every figure, date and name found in the source
[ ] Every citation or link opened and checked
[ ] Totals and comparisons recalculated
[ ] No claim that an action was taken unless a connected tool did it

Completeness
[ ] Every part of the request answered
[ ] Gaps marked, not filled with guesses

Consistency and balance
[ ] Summary matches the body and any tables
[ ] Other viewpoints or affected groups considered
[ ] Wording does not stereotype or favour a group without reason

Fit for the reader
[ ] Tone, length and terms suit the audience
[ ] Format suits where it will be used (email, slide, sheet)

Sign-off
[ ] Human review level matches the stakes table
[ ] Reviewer name: ____________

When a person must review

SituationReview neededWhy
Internal brainstorm or first draft for yourselfYour own read-throughLow stakes; you will rework it anyway
Content for customers, the public or senior leadersFull fact-check and owner sign-offErrors reach people who act on them
Legal, financial, medical, HR or regulatory contentReview by a qualified specialistClaude does not replace professional judgement
Decisions that affect a specific personA person makes and owns the decisionAccountability stays with people

Four questions set the level: how costly is an error, can it be undone, who will see it, and does a rule govern it? A board pack that reads cleanly still trips three. Fix your never-send-without-review list in advance: client deliverables, reported figures, regulated data, public or legal statements.

Iterating is not escalating. If rounds of prompting have stopped improving a high-stakes draft, another prompt will not supply the missing judgement; pass it to a colleague or specialist.

Adapt and compare for the reader

Before, a technical summary written for the IT team:

The SSO migration completed on schedule. 4% of users hit SAML assertion
errors due to clock skew on legacy IdP nodes; remediation is in progress.

After, the same facts for all staff:

The new single sign-on is live. A small number of people may see a login
error this week. If you do, wait five minutes and try again. IT is fixing
the cause.

The facts stayed; the vocabulary, detail and call to action changed. When comparing two versions, score them against criteria set in advance, not on which reads best.

Choosing the output format

The output isUseWhy
A short answer or a few lines you will read onceInline in the chatNothing to reuse or share
A document, slide outline, diagram or small tool you will edit or shareAn artifactSubstantial, self-contained content shown beside the chat, private by default and publishable
Data going into a spreadsheet or another systemA table or structured data such as CSVColumns and rows paste cleanly and can be checked line by line
Totals, percentages or charts that will be reportedCode execution: Claude writes and runs code on your file in a sandboxThe figure is calculated from the rows, not written as likely-looking text

Code execution makes a figure traceable, not automatically right: Claude writes the code, so spot-check what was calculated. When quality matters, compare two drafts and edit the stronger one.

Rules that decide exam answers

  • Verify against the source, not against Claude. A confident tone or a self-rated confidence score is not evidence; asking for one is the tempting wrong option.
  • Specific details are the riskiest. Citation numbers, statistics, quotes and links are the most likely to be invented, so check those first.
  • Review scales with the stakes, not the polish. Legal, financial or external content gets specialist review however clean it looks; low-risk internal drafts do not need full review.
  • Rewording is not checking. Making an output more formal, shorter or clearer does nothing for its accuracy.
  • A person owns the result. Claude can draft and summarise; a named person approves what goes out and any decision about people.
  • Compute figures that matter. Totals and percentages that will be reported come from code execution or a recalculation, not from prose.

Where it appears in the exam

Domain 2 carries 21% of the exam, the highest weight, so expect around 12 or 13 of the 60 items. Questions describe an output about to be shared and ask what to check, what the error is, how much review it needs, or which format suits the reader.

Two sample questions

These are original Timo practice questions. They are not official exam questions.

Question 1

A finance assistant pastes a table of quarterly sales by region and asks Claude for a short commentary for regional managers. The commentary says the North region "grew fastest this quarter", but the table shows South with the larger increase. The other figures look right. What is the best next step?

Answer: D. One wrong comparison shows the output was not checked against its source, so every claim needs checking before it goes to managers. A sends a known error, B relies on self-reported confidence, and C may remove one error without checking for others.

Question 2

A marketing team uses Claude to draft about 30 social posts a week from an approved message guide. This week it also drafts a press release announcing a partnership, with financial figures and a quote from the CEO. Which review plan fits best?

Answer: B. Review should match the stakes: routine posts from an approved guide need spot checks, while public figures and a named quote need full checking and sign-off. A puts the effort in the wrong place, C ignores the risk of invented figures, and D lets Claude check its own work.

Build exercise

  1. Ask Claude to summarise a document you know well. Run the review checklist above and record every error.
  2. Ask Claude to quote the passage behind each claim in the summary. Check that each quote exists word for word.
  3. Ask for one message in versions for two audiences. Score each against four criteria written down first.

Practise this topic

Sources