Output handling: CCDV-F study guide
CCDV-F · Prompt and Context Engineering, topic weight 2.6% of the exam
Output Handling sits in Prompt and Context Engineering, which is 11.0% of the CCDV-F exam, and carries 2.6% on its own. It tests what your code does with Claude's response before anything downstream uses it: get it in a predictable structure, check it, parse it without crashing, and refuse to trust an answer just because it sounds certain.
What the official guide covers
The Claude Certified Developer Foundations exam guide (version 1.0, effective July 2026) describes this topic as patterns for producing, validating and consuming Claude output.
| What the guide lists | What it means in practice |
|---|---|
| Structured output patterns | JSON outputs with a schema, or tools with strict: true, instead of format instructions alone |
| Response validation | Check stop_reason, then the structure, then the values against your business rules |
| Defensive parsing | Read content blocks by type, catch parse and validation errors, and never pass unchecked output on |
| Skepticism toward confident output | A fluent, certain answer can still be wrong; verify it against the source or a system of record |
Getting structured output
| Pattern | How you ask for it | What it guarantees |
|---|---|---|
| JSON outputs | output_config={"format": {"type": "json_schema", "schema": ...}} | The text response is valid JSON that matches the schema |
| Strict tool use | "strict": True on a tool definition | The tool's input matches its input_schema |
Forced tool call without strict | tool_choice={"type": "tool", "name": ...} | That the tool is called, not that its input matches the schema; add strict: true for that. Not every model supports forced tool use |
| Format instructions only | "Reply in JSON with these fields..." | Nothing; parse with full error handling |
| Prefilling the reply | A partial assistant message | Rejected with a 400 error on newer models; use structured outputs |
Structured outputs use constrained decoding, so the schema is followed. Some details still need your attention:
- Objects in the schema need
"additionalProperties": false. - JSON Schema support is limited. Constraints such as
minimum,maximum,minLengthandmaxLengthare not enforced; the SDK helpers move them into field descriptions. Check them in code. - Capitalisation of string
enumandconstvalues is not guaranteed. Normalise before you compare. - Citations cannot be combined with
output_config.format; the request returns a 400 error. - The first request with a new schema has extra latency while its grammar compiles. A compiled grammar stays cached for 24 hours after its last use, so a stable schema pays this once, while schemas generated per request pay it every time.
- The API adds a short system prompt describing the format, so input tokens rise slightly. Changing
output_config.formatalso invalidates the prompt cache for that conversation.
The CCAR-F guide covers schema design in 4.3 Structured output with JSON schemas. This page is about what your code does next.
When schema-valid output still fails
| Failure | What you see | Handle it by |
|---|---|---|
stop_reason is max_tokens | Truncated JSON that does not parse | Treat as incomplete; raise max_tokens and retry |
stop_reason is refusal | A refusal that may not match the schema, returned with HTTP 200 | Do not parse and do not retry blindly; your HTTP error handling will not catch it, so log it and follow your fallback |
| Thinking blocks come first | content[0] is not the text block | Select blocks where type == "text" |
| Valid structure, wrong value | A total that does not add up, an invented ID | Business rule checks against the source |
The last row is the one the exam cares about most. A schema guarantees shape, not truth.
A defensive parser (Python)
import json
import anthropic
from pydantic import BaseModel, ValidationError
client = anthropic.Anthropic()
class Invoice(BaseModel):
invoice_number: str
currency: str
line_total: float
tax: float
grand_total: float
INVOICE_SCHEMA = {
"type": "object",
"properties": {
"invoice_number": {"type": "string"},
"currency": {"type": "string"},
"line_total": {"type": "number"},
"tax": {"type": "number"},
"grand_total": {"type": "number"},
},
"required": ["invoice_number", "currency", "line_total", "tax", "grand_total"],
"additionalProperties": False,
}
class OutputError(Exception):
pass
def parse_invoice(response, source_text: str) -> Invoice:
# 1. Was the response complete?
if response.stop_reason in ("max_tokens", "refusal"):
raise OutputError(f"incomplete output: {response.stop_reason}")
# 2. Read text blocks by type, never by position.
text = "".join(b.text for b in response.content if b.type == "text")
# 3. Parse and validate the structure.
try:
invoice = Invoice.model_validate(json.loads(text))
except (json.JSONDecodeError, ValidationError) as err:
raise OutputError(f"bad structure: {err}") from err
# 4. Check the values, not just the shape.
invoice.currency = invoice.currency.upper()
if abs(invoice.line_total + invoice.tax - invoice.grand_total) > 0.01:
raise OutputError("totals do not add up")
if invoice.invoice_number not in source_text:
raise OutputError("invoice number does not appear in the source document")
return invoice
response = client.messages.create(
model=MODEL,
max_tokens=2048,
messages=[{"role": "user", "content": f"Extract the invoice fields.\n<invoice>\n{source_text}\n</invoice>"}],
output_config={"format": {"type": "json_schema", "schema": INVOICE_SCHEMA}},
)
invoice = parse_invoice(response, source_text)
When parse_invoice raises, do not pass a partial result on. Retry once with the error message added to the conversation (the CCAR-F guide covers this in 4.4 Validation and retry loops), and send repeated failures to a person. Log every failure type so you can see which one is growing.
Consuming a streamed response
Streaming changes when output is safe to use. The response arrives as events: message_start, then for each content block a content_block_start, one or more content_block_delta events and a content_block_stop, then message_delta (which carries the stop_reason) and finally message_stop.
- Show text as it arrives, but act on nothing partial. A
tool_useblock's input arrives asinput_json_deltafragments. They are not valid JSON until that block'scontent_block_stop, so parse and run the tool only then. - Commit the turn only after
message_stop. If your loop appends the assistant turn whenever it stops reading, a dropped connection writes a half-builttool_useblock into history, and the error appears on the next request, far from its cause. - On interruption, discard partial tool and thinking blocks. The streaming docs state they cannot be partially recovered. Retry from the last complete turn; only text can be continued.
- Fine-grained tool streaming skips validation. With
eager_input_streamingon a tool, input arrives unbuffered and can be invalid JSON even after the block closes, or cut mid-parameter bymax_tokens. Guard the parse, and return unparseable input to Claude as atool_resultwithis_error: true.
The SDK's stream helper can assemble the final message for you, which removes most of this bookkeeping.
Skepticism toward confident output
Claude states a wrong answer in the same fluent, confident tone as a right one. Tone is not evidence. Anthropic's guidance on reducing hallucinations gives techniques you can build into the application:
- Allow "I don't know". Give the schema a way to say a value was not found, such as a
foundflag, so Claude is not pushed to invent one. - Ground in quotes. With long documents, have Claude pull out the supporting quotes first, then check in code that each quote appears in the source.
- Verify with citations. Ask for a source for each claim, and retract claims with no support.
- Compare runs. Send the same prompt several times; disagreement between outputs is a warning sign.
- Restrict to the provided material. Tell Claude to use only the documents given.
The docs are clear that these techniques reduce hallucinations but do not remove them. Check critical values against a system of record (does this customer ID exist, does this total add up) and put a person in the loop for high-stakes outputs. A confidence score that Claude writes about its own answer is more output, not a measurement. Do not route on it until you have checked it against labelled results.
Which check fits
| Situation | Choose | Why |
|---|---|---|
| Code parses the response | JSON outputs with a schema | Valid, schema-matching JSON every time the response completes |
| Claude must call a function with exact argument types | Strict tool use | Tool input always matches the schema |
| A field has a range or length limit | A check in code | The schema cannot enforce it |
| JSON is valid but a value is wrong | Business rule checks, then retry or review | Shape is guaranteed; correctness is not |
stop_reason is max_tokens | Discard, raise the limit, retry | Partial output is not a result |
| A streamed tool call or turn | Parse after content_block_stop, commit after message_stop | A stream that ended is not a message that completed |
| A high-stakes extraction | Quotes, verification against the source, human review | Confidence in tone proves nothing |
Rules that decide exam answers
- Schema-valid is not correct. Structured outputs guarantee shape; business rules check truth.
- Check
stop_reasonbefore you parse.max_tokensandrefusalmean the output may not match the schema. - Read content blocks by type. Code that reads
content[0].textbreaks when a thinking block comes first. - Use structured outputs or strict tools, not format instructions, prefill or a forced tool alone. Prefill is rejected on newer models, and forcing a tool does not validate its input.
- Enforce in code what the schema cannot. Ranges, lengths, cross-field totals and enum case.
- Confident is not verified. Check against the source or a system of record, and route high-stakes cases to a person.
Where it appears in the exam
Output Handling is 2.6% of the exam, inside Prompt and Context Engineering (11.0%). The guide's description points to questions about extraction and generation pipelines whose output is consumed by code: a parser that crashes, output that is well formed but wrong, or a team that trusts a confident answer it should have checked.
Two sample questions
These are original Timo practice questions. They are not official exam questions.
Build exercise
- Define the invoice schema above and extract five invoices with
output_config.format. Parse them withparse_invoice. - Set
max_tokensvery low and confirm the parser rejects the output throughstop_reason, not through a JSON error. - Edit one source invoice so its totals do not add up, and one so the invoice number is missing. Confirm both are caught.
- Add a single retry that sends the error back to Claude, then a review queue for second failures. Log the failure rate by type.
Practise this topic
- Claude Certified Developer practice exam: free, 20 questions, no sign-up
- CCDV-F study guide: all topics
- Worked example: Professional output validation
- Same topic in another exam: CCAR-F 4.3 Structured output with JSON schemas and CCAO-F Output Evaluation and Validation
- Previous topic: Prompt Engineering
- Next topic: AI Application Security
Sources
- Claude Certified Developer Foundations Exam Guide, version 1.0, effective July 2026 (Anthropic), topic: Output Handling (Domain 6, Prompt and Context Engineering)
- Anthropic documentation: Structured outputs
- Anthropic documentation: Strict tool use
- Anthropic documentation: Streaming messages
- Anthropic documentation: Reduce hallucinations
By Amotion AI