Claude API mechanics: CCDV-F study guide
CCDV-F · Applications and Integration, topic weight 6.8% of the exam
Claude API Mechanics is a 6.8% topic in Applications and Integration, which makes up 33.1% of the CCDV-F exam. It tests how the Messages API behaves across its whole surface: a single request, tools, streaming, images, thinking, caching, calling Claude through cloud platforms, and choosing realtime or batch.
What the official guide covers
The Claude Certified Developer Foundations exam guide (version 1.0, effective July 2026) describes this topic as Claude API behaviour and mechanics, from messages and tools to "tradeoffs between realtime and batch API selection".
| What the guide lists | What it means in practice |
|---|---|
| Messages | Stateless requests that carry the full history; content blocks; stop_reason |
| Tools | Tool definitions, tool_use and tool_result blocks, tool_choice, client vs server tools |
| Streaming | Server-sent events, their order, and errors that arrive after a 200 |
| Vision | image and document blocks, formats and placement |
| Thinking | Adaptive thinking with an effort level; returning thinking blocks in tool loops |
| Caching | cache_control breakpoints and the cache fields in usage |
| Third-party vendors | Claude on Amazon Bedrock and Google Cloud Vertex AI: clients, model IDs, feature gaps |
| Data access patterns | Inline base64, URLs, Files API IDs, token counting, the Models API, batch results |
| Batch API and realtime vs batch | Message Batches API for volume that can wait; Messages API when someone is waiting |
One request, end to end
Every call is POST /v1/messages with three required fields: model, messages and max_tokens. The API is stateless: each request carries the whole conversation, with user and assistant turns alternating, and standing instructions go in the top-level system field.
The response holds content, a list of blocks (text, tool_use, thinking), plus stop_reason and usage. Read stop_reason before you use the answer: end_turn means finished, max_tokens means cut off, tool_use means Claude wants a tool run, refusal means Claude declined, stop_sequence means one of your stop_sequences was emitted, model_context_window_exceeded means generation reached the model's context limit, and pause_turn means a server-tool loop paused and you send the response back to continue.
Tools
A tool definition has three fields: name, description and input_schema. When Claude calls a client tool, the response has stop_reason: "tool_use" and a tool_use block with an id, a name and an input. Your code runs it and sends a tool_result with the matching tool_use_id in the next user message. Server tools such as web search run on Anthropic's infrastructure, with no handler to write. Anthropic-schema client tools, such as bash and the text editor, sit in between: you declare them by a versioned type, Claude knows their schema, and your code still executes them.
The result rules are strict, and breaking them fails validation:
- Claude can return several
tool_useblocks in one turn. Return onetool_resultper call, all in the next user message, before any text in that message. To limit Claude to one call per turn, settool_choice: {"type": "auto", "disable_parallel_tool_use": true}. - Append the whole assistant
contentlist to history, including any text block that came with the tool call. - When a tool fails, return its
tool_resultwithis_error: trueand a message that says what went wrong and what to try next. An empty result reads as valid data.
tool_choice defaults to auto. any forces some tool, {"type": "tool", "name": "..."} forces one, and none blocks tools. Some newer models reject any and tool with a 400 error; use auto with strict: true on the tool instead. Writing the tools themselves is covered in Tool Implementation.
Streaming
Set stream: true, or use client.messages.stream() in the SDK. The response arrives as server-sent events in this order:
message_start, with an emptycontentlist- For each block:
content_block_start, one or morecontent_block_deltaevents, thencontent_block_stop message_delta, carryingstop_reasonand the finalusagemessage_stop
Delta types are text_delta, input_json_delta (partial tool input), thinking_delta and signature_delta. ping events can appear at any point. An error event, such as overloaded_error, can arrive after the HTTP status was already 200, so your client must handle it. For large max_tokens values the SDK requires streaming to avoid HTTP timeouts; stream.get_final_message() still gives you the complete message.
A stream that stops has not necessarily delivered a complete message. Tool input arrives as partial JSON strings in input_json_delta events, so parse it only after that block's content_block_stop, or use the SDK helpers. Add the assistant turn to history only after message_stop. If the connection drops first, do not save the partial turn as if it were complete: a half-built tool_use block makes the next request fail, and the error points at the retry rather than the stream. Retry from the last complete turn, or follow the documented recovery: keep the partial text and ask Claude to continue. By default the API buffers and validates each tool parameter before streaming it; setting eager_input_streaming: true on a tool streams it sooner, unvalidated, so your parser must handle invalid JSON.
Vision and documents
Send images as image blocks with a base64, url or file source. Supported formats are JPEG, PNG, GIF and WebP; only the first frame of an animated GIF is used. Claude reads an image in 28 by 28 pixel patches, one token each, so a 1,000 by 1,000 pixel image costs about 1,300 tokens. Images above the model's native resolution are downscaled first, so resizing them yourself saves upload size without losing detail Claude would see. Place images before the text that asks about them.
Pick the source by reuse: inline base64 for an image sent once, a Files API file_id for a reference image sent with every request, and the Message Batches API when thousands of inputs can wait.
Send PDFs as document blocks. Claude reads each page as an image plus its extracted text, so charts count. On Bedrock and Vertex AI, send PDFs as base64.
Thinking
On newer models, turn thinking on with thinking: {"type": "adaptive"} and set depth with output_config: {"effort": ...} (low, medium, high, xhigh or max). Older models use thinking: {"type": "enabled", "budget_tokens": N}, with the budget below max_tokens. The response includes thinking blocks. In a tool loop, send them back unchanged as part of the assistant turn, together with any redacted_thinking blocks; code that keeps only type == "thinking" silently drops those and breaks the loop. On newer models the thinking text is omitted by default; request thinking: {"type": "adaptive", "display": "summarized"} to see a summary. Thinking is billed as output tokens either way; usage.output_tokens_details.thinking_tokens shows how many. Leave it off for classification, extraction and format conversion; turn it on for multi-step reasoning and for planning across several tool calls.
Prompt caching
Add cache_control: {"type": "ephemeral"} to a block to mark a breakpoint (up to four), or put cache_control at the top level of the request for automatic caching. The cache is a prefix in the order tools, system, messages, so changing the tools invalidates everything after them. usage.cache_creation_input_tokens and usage.cache_read_input_tokens show writes and hits. Prompts below a model's minimum cacheable length are not cached, with no error. An entry lasts five minutes by default and each hit refreshes it; "ttl": "1h" inside cache_control keeps it for an hour. Cost modelling is covered in Cost and Token Management.
Example: the same request on two platforms
import base64
import os
from anthropic import Anthropic, AnthropicVertex
if os.environ.get("USE_VERTEX"):
client = AnthropicVertex(project_id=os.environ["GCP_PROJECT"], region="global")
else:
client = Anthropic()
MODEL = os.environ["CLAUDE_MODEL"] # model IDs differ between platforms
image = base64.standard_b64encode(open("invoice.png", "rb").read()).decode()
with client.messages.stream(
model=MODEL,
max_tokens=1024,
system=[{
"type": "text",
"text": open("extraction_rules.txt").read(), # long and stable: cache it
"cache_control": {"type": "ephemeral"},
}],
messages=[{
"role": "user",
"content": [
{"type": "image", "source": {"type": "base64", "media_type": "image/png", "data": image}},
{"type": "text", "text": "Extract the supplier, total and due date."},
],
}],
) as stream:
for text in stream.text_stream:
print(text, end="", flush=True)
final = stream.get_final_message()
print("\n", final.stop_reason, final.usage.cache_read_input_tokens)
On a cloud platform, authentication uses that cloud's credentials (Google Cloud application default credentials, AWS IAM), not an Anthropic API key, and model IDs follow the platform's format (Bedrock adds an anthropic. prefix). The Python clients are AnthropicVertex (anthropic[vertex]), AnthropicBedrockMantle for Bedrock's Messages API endpoint or AnthropicBedrock for the older bedrock-runtime path (anthropic[bedrock]), and AnthropicFoundry. Not every feature is offered there: the Message Batches API, Files API, URL sources and code execution are missing on both Bedrock and Vertex AI, and Bedrock also lacks structured outputs and web search.
Data access patterns
| How data reaches Claude | Use it when | Watch for |
|---|---|---|
Inline base64 image or document | One-off inputs, or on Bedrock and Vertex AI | Counts toward the request size limit |
url source | The file is already reachable by URL | Not available on Bedrock or Vertex AI |
Files API file_id | The same file is used across many requests | Files are kept until deleted, so not covered by Zero Data Retention |
client.messages.count_tokens | You need the size before sending | Counts are estimates |
| Models API | You need a model's limits and capabilities | Returns max_input_tokens, max_tokens and capabilities |
Realtime or batch
| Situation | Choose | Why |
|---|---|---|
| A person is waiting for the reply | Messages API with streaming | Text appears as it is generated |
| Large volume that can wait for results | Message Batches API | Asynchronous, completes within 24 hours, lower cost |
| A request may run longer than 10 minutes | Streaming or batch | Idle connections can be dropped on long non-streaming calls |
You need stream: true or fast mode | Messages API | Batches do not support either |
| Results must be matched to inputs | Batch with a unique custom_id per request | Batch results do not come back in input order |
Rules that decide exam answers
- The API keeps no conversation. Every request sends the full history. A missing earlier turn is your bug, not the API's.
- Every
tool_usegets atool_result. Sametool_use_id, all in the next user message, withis_error: truewhen the tool failed. - A 200 does not end the story when streaming. Handle
errorevents mid-stream, readstop_reasonfrommessage_delta, and add a turn to history only oncemessage_stoparrives. - Thinking blocks go back unchanged. In a tool loop, return the whole assistant turn, thinking included.
- Cache order is tools, system, messages. Put stable content first and the changing question last.
- Batch is for volume that can wait. Match results by
custom_id, never by position.
Where it appears in the exam
Claude API Mechanics is 6.8% of the exam, inside Applications and Integration (33.1%), so expect three or four of the 53 items. Questions describe a request or response that misbehaves, or ask which API feature fits a workload.
Two sample questions
These are original Timo practice questions. They are not official exam questions.
Build exercise
- Send one request with
curland the same request with the Python SDK. Compare the headers and the response JSON. - Stream a long reply and log every event type in order. Confirm
stop_reasonarrives inmessage_delta. - Run the example above twice within five minutes and compare
cache_creation_input_tokenswithcache_read_input_tokens. - Submit 20 requests as a batch, poll until its status is
ended, and match every result to its input bycustom_id.
Practise this topic
- Claude Certified Developer practice exam: free, 20 questions, no sign-up
- CCDV-F study guide: all topics
- Worked example: Claude Batch API
- Same topic in another exam: 4.5 Message Batches API
- Previous topic: Systems Life Cycle
- Next topic: Software Engineering Foundations
Sources
- Claude Certified Developer Foundations Exam Guide, version 1.0, effective July 2026 (Anthropic), Domain 2 topic: Claude API Mechanics
- Anthropic documentation: Streaming messages
- Anthropic documentation: Prompt caching
- Anthropic documentation: Batch processing
- Anthropic documentation: Claude on Google Cloud
By Amotion AI