TimoBy Amotion AI

Claude API mechanics: CCDV-F study guide

CCDV-F · Applications and Integration, topic weight 6.8% of the exam

Claude API Mechanics is a 6.8% topic in Applications and Integration, which makes up 33.1% of the CCDV-F exam. It tests how the Messages API behaves across its whole surface: a single request, tools, streaming, images, thinking, caching, calling Claude through cloud platforms, and choosing realtime or batch.

What the official guide covers

The Claude Certified Developer Foundations exam guide (version 1.0, effective July 2026) describes this topic as Claude API behaviour and mechanics, from messages and tools to "tradeoffs between realtime and batch API selection".

What the guide listsWhat it means in practice
MessagesStateless requests that carry the full history; content blocks; stop_reason
ToolsTool definitions, tool_use and tool_result blocks, tool_choice, client vs server tools
StreamingServer-sent events, their order, and errors that arrive after a 200
Visionimage and document blocks, formats and placement
ThinkingAdaptive thinking with an effort level; returning thinking blocks in tool loops
Cachingcache_control breakpoints and the cache fields in usage
Third-party vendorsClaude on Amazon Bedrock and Google Cloud Vertex AI: clients, model IDs, feature gaps
Data access patternsInline base64, URLs, Files API IDs, token counting, the Models API, batch results
Batch API and realtime vs batchMessage Batches API for volume that can wait; Messages API when someone is waiting

One request, end to end

Every call is POST /v1/messages with three required fields: model, messages and max_tokens. The API is stateless: each request carries the whole conversation, with user and assistant turns alternating, and standing instructions go in the top-level system field.

The response holds content, a list of blocks (text, tool_use, thinking), plus stop_reason and usage. Read stop_reason before you use the answer: end_turn means finished, max_tokens means cut off, tool_use means Claude wants a tool run, refusal means Claude declined, stop_sequence means one of your stop_sequences was emitted, model_context_window_exceeded means generation reached the model's context limit, and pause_turn means a server-tool loop paused and you send the response back to continue.

Tools

A tool definition has three fields: name, description and input_schema. When Claude calls a client tool, the response has stop_reason: "tool_use" and a tool_use block with an id, a name and an input. Your code runs it and sends a tool_result with the matching tool_use_id in the next user message. Server tools such as web search run on Anthropic's infrastructure, with no handler to write. Anthropic-schema client tools, such as bash and the text editor, sit in between: you declare them by a versioned type, Claude knows their schema, and your code still executes them.

The result rules are strict, and breaking them fails validation:

  • Claude can return several tool_use blocks in one turn. Return one tool_result per call, all in the next user message, before any text in that message. To limit Claude to one call per turn, set tool_choice: {"type": "auto", "disable_parallel_tool_use": true}.
  • Append the whole assistant content list to history, including any text block that came with the tool call.
  • When a tool fails, return its tool_result with is_error: true and a message that says what went wrong and what to try next. An empty result reads as valid data.

tool_choice defaults to auto. any forces some tool, {"type": "tool", "name": "..."} forces one, and none blocks tools. Some newer models reject any and tool with a 400 error; use auto with strict: true on the tool instead. Writing the tools themselves is covered in Tool Implementation.

Streaming

Set stream: true, or use client.messages.stream() in the SDK. The response arrives as server-sent events in this order:

  1. message_start, with an empty content list
  2. For each block: content_block_start, one or more content_block_delta events, then content_block_stop
  3. message_delta, carrying stop_reason and the final usage
  4. message_stop

Delta types are text_delta, input_json_delta (partial tool input), thinking_delta and signature_delta. ping events can appear at any point. An error event, such as overloaded_error, can arrive after the HTTP status was already 200, so your client must handle it. For large max_tokens values the SDK requires streaming to avoid HTTP timeouts; stream.get_final_message() still gives you the complete message.

A stream that stops has not necessarily delivered a complete message. Tool input arrives as partial JSON strings in input_json_delta events, so parse it only after that block's content_block_stop, or use the SDK helpers. Add the assistant turn to history only after message_stop. If the connection drops first, do not save the partial turn as if it were complete: a half-built tool_use block makes the next request fail, and the error points at the retry rather than the stream. Retry from the last complete turn, or follow the documented recovery: keep the partial text and ask Claude to continue. By default the API buffers and validates each tool parameter before streaming it; setting eager_input_streaming: true on a tool streams it sooner, unvalidated, so your parser must handle invalid JSON.

Vision and documents

Send images as image blocks with a base64, url or file source. Supported formats are JPEG, PNG, GIF and WebP; only the first frame of an animated GIF is used. Claude reads an image in 28 by 28 pixel patches, one token each, so a 1,000 by 1,000 pixel image costs about 1,300 tokens. Images above the model's native resolution are downscaled first, so resizing them yourself saves upload size without losing detail Claude would see. Place images before the text that asks about them.

Pick the source by reuse: inline base64 for an image sent once, a Files API file_id for a reference image sent with every request, and the Message Batches API when thousands of inputs can wait.

Send PDFs as document blocks. Claude reads each page as an image plus its extracted text, so charts count. On Bedrock and Vertex AI, send PDFs as base64.

Thinking

On newer models, turn thinking on with thinking: {"type": "adaptive"} and set depth with output_config: {"effort": ...} (low, medium, high, xhigh or max). Older models use thinking: {"type": "enabled", "budget_tokens": N}, with the budget below max_tokens. The response includes thinking blocks. In a tool loop, send them back unchanged as part of the assistant turn, together with any redacted_thinking blocks; code that keeps only type == "thinking" silently drops those and breaks the loop. On newer models the thinking text is omitted by default; request thinking: {"type": "adaptive", "display": "summarized"} to see a summary. Thinking is billed as output tokens either way; usage.output_tokens_details.thinking_tokens shows how many. Leave it off for classification, extraction and format conversion; turn it on for multi-step reasoning and for planning across several tool calls.

Prompt caching

Add cache_control: {"type": "ephemeral"} to a block to mark a breakpoint (up to four), or put cache_control at the top level of the request for automatic caching. The cache is a prefix in the order tools, system, messages, so changing the tools invalidates everything after them. usage.cache_creation_input_tokens and usage.cache_read_input_tokens show writes and hits. Prompts below a model's minimum cacheable length are not cached, with no error. An entry lasts five minutes by default and each hit refreshes it; "ttl": "1h" inside cache_control keeps it for an hour. Cost modelling is covered in Cost and Token Management.

Example: the same request on two platforms

import base64
import os
from anthropic import Anthropic, AnthropicVertex

if os.environ.get("USE_VERTEX"):
    client = AnthropicVertex(project_id=os.environ["GCP_PROJECT"], region="global")
else:
    client = Anthropic()

MODEL = os.environ["CLAUDE_MODEL"]   # model IDs differ between platforms
image = base64.standard_b64encode(open("invoice.png", "rb").read()).decode()

with client.messages.stream(
    model=MODEL,
    max_tokens=1024,
    system=[{
        "type": "text",
        "text": open("extraction_rules.txt").read(),   # long and stable: cache it
        "cache_control": {"type": "ephemeral"},
    }],
    messages=[{
        "role": "user",
        "content": [
            {"type": "image", "source": {"type": "base64", "media_type": "image/png", "data": image}},
            {"type": "text", "text": "Extract the supplier, total and due date."},
        ],
    }],
) as stream:
    for text in stream.text_stream:
        print(text, end="", flush=True)
    final = stream.get_final_message()

print("\n", final.stop_reason, final.usage.cache_read_input_tokens)

On a cloud platform, authentication uses that cloud's credentials (Google Cloud application default credentials, AWS IAM), not an Anthropic API key, and model IDs follow the platform's format (Bedrock adds an anthropic. prefix). The Python clients are AnthropicVertex (anthropic[vertex]), AnthropicBedrockMantle for Bedrock's Messages API endpoint or AnthropicBedrock for the older bedrock-runtime path (anthropic[bedrock]), and AnthropicFoundry. Not every feature is offered there: the Message Batches API, Files API, URL sources and code execution are missing on both Bedrock and Vertex AI, and Bedrock also lacks structured outputs and web search.

Data access patterns

How data reaches ClaudeUse it whenWatch for
Inline base64 image or documentOne-off inputs, or on Bedrock and Vertex AICounts toward the request size limit
url sourceThe file is already reachable by URLNot available on Bedrock or Vertex AI
Files API file_idThe same file is used across many requestsFiles are kept until deleted, so not covered by Zero Data Retention
client.messages.count_tokensYou need the size before sendingCounts are estimates
Models APIYou need a model's limits and capabilitiesReturns max_input_tokens, max_tokens and capabilities

Realtime or batch

SituationChooseWhy
A person is waiting for the replyMessages API with streamingText appears as it is generated
Large volume that can wait for resultsMessage Batches APIAsynchronous, completes within 24 hours, lower cost
A request may run longer than 10 minutesStreaming or batchIdle connections can be dropped on long non-streaming calls
You need stream: true or fast modeMessages APIBatches do not support either
Results must be matched to inputsBatch with a unique custom_id per requestBatch results do not come back in input order

Rules that decide exam answers

  • The API keeps no conversation. Every request sends the full history. A missing earlier turn is your bug, not the API's.
  • Every tool_use gets a tool_result. Same tool_use_id, all in the next user message, with is_error: true when the tool failed.
  • A 200 does not end the story when streaming. Handle error events mid-stream, read stop_reason from message_delta, and add a turn to history only once message_stop arrives.
  • Thinking blocks go back unchanged. In a tool loop, return the whole assistant turn, thinking included.
  • Cache order is tools, system, messages. Put stable content first and the changing question last.
  • Batch is for volume that can wait. Match results by custom_id, never by position.

Where it appears in the exam

Claude API Mechanics is 6.8% of the exam, inside Applications and Integration (33.1%), so expect three or four of the 53 items. Questions describe a request or response that misbehaves, or ask which API feature fits a workload.

Two sample questions

These are original Timo practice questions. They are not official exam questions.

Question 1

A developer streams Claude's replies into a web chat. Now and then a reply stops halfway and the chat shows it as complete. Every request in the logs returned HTTP 200. What is the most likely gap in the client code?

Answer: C. In a stream, the status is sent before generation, so a later failure arrives as an error event that the client must detect. A would fail every request at once. B loses no text, since deltas carry it. D is wrong because a low max_tokens gives stop_reason: "max_tokens", not a 413.

Question 2

An agent uses adaptive thinking and two client tools. To save tokens, the developer removes the thinking block from Claude's assistant turn before sending the tool results back. Following Anthropic's documented rule, what should the next request contain?

Answer: A. Thinking blocks must be passed back unmodified during tool use. B is wrong because the API is stateless. C modifies the block. D breaks turn order: the results answer the assistant turn, so they come after it.

Build exercise

  1. Send one request with curl and the same request with the Python SDK. Compare the headers and the response JSON.
  2. Stream a long reply and log every event type in order. Confirm stop_reason arrives in message_delta.
  3. Run the example above twice within five minutes and compare cache_creation_input_tokens with cache_read_input_tokens.
  4. Submit 20 requests as a batch, poll until its status is ended, and match every result to its input by custom_id.

Practise this topic

Sources