TimoBy Amotion AI

Output handling: CCDV-F study guide

CCDV-F · Prompt and Context Engineering, topic weight 2.6% of the exam

Output Handling sits in Prompt and Context Engineering, which is 11.0% of the CCDV-F exam, and carries 2.6% on its own. It tests what your code does with Claude's response before anything downstream uses it: get it in a predictable structure, check it, parse it without crashing, and refuse to trust an answer just because it sounds certain.

What the official guide covers

The Claude Certified Developer Foundations exam guide (version 1.0, effective July 2026) describes this topic as patterns for producing, validating and consuming Claude output.

What the guide listsWhat it means in practice
Structured output patternsJSON outputs with a schema, or tools with strict: true, instead of format instructions alone
Response validationCheck stop_reason, then the structure, then the values against your business rules
Defensive parsingRead content blocks by type, catch parse and validation errors, and never pass unchecked output on
Skepticism toward confident outputA fluent, certain answer can still be wrong; verify it against the source or a system of record

Getting structured output

PatternHow you ask for itWhat it guarantees
JSON outputsoutput_config={"format": {"type": "json_schema", "schema": ...}}The text response is valid JSON that matches the schema
Strict tool use"strict": True on a tool definitionThe tool's input matches its input_schema
Forced tool call without stricttool_choice={"type": "tool", "name": ...}That the tool is called, not that its input matches the schema; add strict: true for that. Not every model supports forced tool use
Format instructions only"Reply in JSON with these fields..."Nothing; parse with full error handling
Prefilling the replyA partial assistant messageRejected with a 400 error on newer models; use structured outputs

Structured outputs use constrained decoding, so the schema is followed. Some details still need your attention:

  • Objects in the schema need "additionalProperties": false.
  • JSON Schema support is limited. Constraints such as minimum, maximum, minLength and maxLength are not enforced; the SDK helpers move them into field descriptions. Check them in code.
  • Capitalisation of string enum and const values is not guaranteed. Normalise before you compare.
  • Citations cannot be combined with output_config.format; the request returns a 400 error.
  • The first request with a new schema has extra latency while its grammar compiles. A compiled grammar stays cached for 24 hours after its last use, so a stable schema pays this once, while schemas generated per request pay it every time.
  • The API adds a short system prompt describing the format, so input tokens rise slightly. Changing output_config.format also invalidates the prompt cache for that conversation.

The CCAR-F guide covers schema design in 4.3 Structured output with JSON schemas. This page is about what your code does next.

When schema-valid output still fails

FailureWhat you seeHandle it by
stop_reason is max_tokensTruncated JSON that does not parseTreat as incomplete; raise max_tokens and retry
stop_reason is refusalA refusal that may not match the schema, returned with HTTP 200Do not parse and do not retry blindly; your HTTP error handling will not catch it, so log it and follow your fallback
Thinking blocks come firstcontent[0] is not the text blockSelect blocks where type == "text"
Valid structure, wrong valueA total that does not add up, an invented IDBusiness rule checks against the source

The last row is the one the exam cares about most. A schema guarantees shape, not truth.

A defensive parser (Python)

import json
import anthropic
from pydantic import BaseModel, ValidationError

client = anthropic.Anthropic()

class Invoice(BaseModel):
    invoice_number: str
    currency: str
    line_total: float
    tax: float
    grand_total: float

INVOICE_SCHEMA = {
    "type": "object",
    "properties": {
        "invoice_number": {"type": "string"},
        "currency": {"type": "string"},
        "line_total": {"type": "number"},
        "tax": {"type": "number"},
        "grand_total": {"type": "number"},
    },
    "required": ["invoice_number", "currency", "line_total", "tax", "grand_total"],
    "additionalProperties": False,
}

class OutputError(Exception):
    pass

def parse_invoice(response, source_text: str) -> Invoice:
    # 1. Was the response complete?
    if response.stop_reason in ("max_tokens", "refusal"):
        raise OutputError(f"incomplete output: {response.stop_reason}")
    # 2. Read text blocks by type, never by position.
    text = "".join(b.text for b in response.content if b.type == "text")
    # 3. Parse and validate the structure.
    try:
        invoice = Invoice.model_validate(json.loads(text))
    except (json.JSONDecodeError, ValidationError) as err:
        raise OutputError(f"bad structure: {err}") from err
    # 4. Check the values, not just the shape.
    invoice.currency = invoice.currency.upper()
    if abs(invoice.line_total + invoice.tax - invoice.grand_total) > 0.01:
        raise OutputError("totals do not add up")
    if invoice.invoice_number not in source_text:
        raise OutputError("invoice number does not appear in the source document")
    return invoice

response = client.messages.create(
    model=MODEL,
    max_tokens=2048,
    messages=[{"role": "user", "content": f"Extract the invoice fields.\n<invoice>\n{source_text}\n</invoice>"}],
    output_config={"format": {"type": "json_schema", "schema": INVOICE_SCHEMA}},
)
invoice = parse_invoice(response, source_text)

When parse_invoice raises, do not pass a partial result on. Retry once with the error message added to the conversation (the CCAR-F guide covers this in 4.4 Validation and retry loops), and send repeated failures to a person. Log every failure type so you can see which one is growing.

Consuming a streamed response

Streaming changes when output is safe to use. The response arrives as events: message_start, then for each content block a content_block_start, one or more content_block_delta events and a content_block_stop, then message_delta (which carries the stop_reason) and finally message_stop.

  • Show text as it arrives, but act on nothing partial. A tool_use block's input arrives as input_json_delta fragments. They are not valid JSON until that block's content_block_stop, so parse and run the tool only then.
  • Commit the turn only after message_stop. If your loop appends the assistant turn whenever it stops reading, a dropped connection writes a half-built tool_use block into history, and the error appears on the next request, far from its cause.
  • On interruption, discard partial tool and thinking blocks. The streaming docs state they cannot be partially recovered. Retry from the last complete turn; only text can be continued.
  • Fine-grained tool streaming skips validation. With eager_input_streaming on a tool, input arrives unbuffered and can be invalid JSON even after the block closes, or cut mid-parameter by max_tokens. Guard the parse, and return unparseable input to Claude as a tool_result with is_error: true.

The SDK's stream helper can assemble the final message for you, which removes most of this bookkeeping.

Skepticism toward confident output

Claude states a wrong answer in the same fluent, confident tone as a right one. Tone is not evidence. Anthropic's guidance on reducing hallucinations gives techniques you can build into the application:

  • Allow "I don't know". Give the schema a way to say a value was not found, such as a found flag, so Claude is not pushed to invent one.
  • Ground in quotes. With long documents, have Claude pull out the supporting quotes first, then check in code that each quote appears in the source.
  • Verify with citations. Ask for a source for each claim, and retract claims with no support.
  • Compare runs. Send the same prompt several times; disagreement between outputs is a warning sign.
  • Restrict to the provided material. Tell Claude to use only the documents given.

The docs are clear that these techniques reduce hallucinations but do not remove them. Check critical values against a system of record (does this customer ID exist, does this total add up) and put a person in the loop for high-stakes outputs. A confidence score that Claude writes about its own answer is more output, not a measurement. Do not route on it until you have checked it against labelled results.

Which check fits

SituationChooseWhy
Code parses the responseJSON outputs with a schemaValid, schema-matching JSON every time the response completes
Claude must call a function with exact argument typesStrict tool useTool input always matches the schema
A field has a range or length limitA check in codeThe schema cannot enforce it
JSON is valid but a value is wrongBusiness rule checks, then retry or reviewShape is guaranteed; correctness is not
stop_reason is max_tokensDiscard, raise the limit, retryPartial output is not a result
A streamed tool call or turnParse after content_block_stop, commit after message_stopA stream that ended is not a message that completed
A high-stakes extractionQuotes, verification against the source, human reviewConfidence in tone proves nothing

Rules that decide exam answers

  • Schema-valid is not correct. Structured outputs guarantee shape; business rules check truth.
  • Check stop_reason before you parse. max_tokens and refusal mean the output may not match the schema.
  • Read content blocks by type. Code that reads content[0].text breaks when a thinking block comes first.
  • Use structured outputs or strict tools, not format instructions, prefill or a forced tool alone. Prefill is rejected on newer models, and forcing a tool does not validate its input.
  • Enforce in code what the schema cannot. Ranges, lengths, cross-field totals and enum case.
  • Confident is not verified. Check against the source or a system of record, and route high-stakes cases to a person.

Where it appears in the exam

Output Handling is 2.6% of the exam, inside Prompt and Context Engineering (11.0%). The guide's description points to questions about extraction and generation pipelines whose output is consumed by code: a parser that crashes, output that is well formed but wrong, or a team that trusts a confident answer it should have checked.

Two sample questions

These are original Timo practice questions. They are not official exam questions.

Question 1

An extraction service uses structured outputs with a JSON schema. Because the schema guarantees valid JSON, the team removed all parsing checks. One night a batch of unusually long documents produces truncated JSON that crashes the parser. Which check was missing?

Answer: C. Output cut off at max_tokens may not match the schema, so stop_reason must be checked before parsing. A affects schema validity, not truncation, B has nothing to do with the crash, and D would wrongly reject responses that start with a thinking block.

Question 2

An invoice extractor returns schema-valid JSON for every document. Finance finds that about one invoice in fifty has a grand total that does not equal the line total plus tax, although Claude's output looks clean and confident. What should the developer add?

Answer: B. A deterministic check in code catches every mismatch, and the retry or review path handles them. A still relies on Claude getting it right, C trusts an unmeasured self-reported score, and D changes the shape, which was never the problem.

Build exercise

  1. Define the invoice schema above and extract five invoices with output_config.format. Parse them with parse_invoice.
  2. Set max_tokens very low and confirm the parser rejects the output through stop_reason, not through a JSON error.
  3. Edit one source invoice so its totals do not add up, and one so the invoice number is missing. Confirm both are caught.
  4. Add a single retry that sends the error back to Claude, then a review queue for second failures. Log the failure rate by type.

Practise this topic

Sources