Message Batches API: CCAR-F task statement 4.5
CCAR-F · Prompt Engineering & Structured Output (20% of the exam)
Task statement 4.5 sits in Prompt Engineering & Structured Output, 20% of the CCAR-F exam. It tests one decision first: whether a workload can wait for the Message Batches API or needs the synchronous Messages API. It then tests how you run a batch well: matching results by custom_id, resubmitting only what failed, and timing submissions against a deadline.
What the official guide covers
The Claude Certified Architect Foundations exam guide (version 1.0, effective July 2026) lists this under task statement 4.5, "Design efficient batch processing strategies":
| Knowledge of | Skills in |
|---|---|
| The Message Batches API costs less than standard calls, can take up to 24 hours to process, and has no guaranteed latency | Using the synchronous API for blocking pre-merge checks and the batch API for overnight or weekly analysis |
| Batches suit work nobody waits on (overnight reports, weekly audits, nightly test generation) and do not suit blocking workflows such as pre-merge checks | Working out how often to submit batches so every item meets an SLA, given up to 24 hours of processing |
| A batch request cannot run your tools mid-request and send the results back to Claude | Resubmitting only the failed documents, found by custom_id, with changes such as splitting documents that were too long |
custom_id fields link each request to its result | Refining the prompt on a sample before processing large volumes, so more requests succeed on the first pass |
How a batch works
- Create. Send a list of requests to
client.messages.batches.create(). Each request has acustom_idandparams, which are the same parameters as a normal Messages API call. - Wait. The batch starts with
processing_statusset toin_progress. Poll it withclient.messages.batches.retrieve()until the status isended. Anthropic's docs say most batches finish within an hour, but any request not processed within 24 hours expires. - Read the results.
client.messages.batches.results()streams a JSONL file, one line per request. Each line has thecustom_idand a result of typesucceeded,errored,canceledorexpired. - Match by
custom_id. Results do not come back in request order.
Other behaviour the exam can draw on:
- A batch can hold 256 MB or 100,000 requests, and the first limit reached applies.
- A
custom_idmust be unique within the batch and use 1 to 64 letters, digits, hyphens or underscores. - One failed request does not affect the others. Only
succeededrequests are billed. - Request parameters are validated when the batch is processed, not when you submit it. Anthropic's docs recommend testing the request shape with the Messages API first.
- Results stay available for 29 days after the batch is created.
What a batch request cannot do
Each batch request returns one response, and nothing on your side is connected while it runs. If Claude calls one of your own tools, the result ends with stop_reason "tool_use" and no code runs that tool. To continue you would send a new request containing the tool_result, which is an agent loop, and an agent loop belongs on the synchronous API.
Anthropic's server tools, such as web search and code execution, are different: they run inside the batch. If such a result ends with pause_turn, the turn is unfinished and you continue it in a follow-up request. Batches also do not stream.
Synchronous or batch?
| Workload | Choose | Why |
|---|---|---|
| Pre-merge check that developers wait on | Messages API | A batch has no latency guarantee |
| Overnight technical debt report or weekly audit | Message Batches API | Nobody waits; lower cost |
| Nightly test generation for the whole repository | Message Batches API | Results are needed the next morning, not now |
| Agent that calls your tools across several turns | Messages API | A batch request cannot return tool results to Claude |
| 50,000 documents to classify by the end of the week | Message Batches API, after a sample run | Large, latency-tolerant volume |
| Re-running a 2,000-case eval on a settled prompt overnight | Message Batches API | Nobody waits for the scores; same request parameters |
A loop of synchronous calls is not a batch
A common mistake: a nightly job hits rate limits, so the team splits its 5,000 items into chunks of 100 and calls the Messages API for each item in each chunk. The API still sees 5,000 separate synchronous requests, so the job hits the same limits. Chunk size was never the problem.
The Message Batches API is a different way of submitting work. One call creates the batch, Anthropic processes it asynchronously, and the batch has its own rate limits, which apply to the batch HTTP requests and to the number of requests waiting to be processed. Use it whenever the volume is high and nobody is waiting.
Batches that share a long prefix: add prompt caching
When every request in a batch carries the same long system prompt or reference document, mark that block with an identical cache_control in every request. The batch and caching discounts stack. Cache hits in a batch are best effort, because requests run concurrently, and a batch can run longer than the default five-minute cache lifetime. Anthropic's batch docs suggest the one-hour cache duration ("ttl": "1h") for batches with shared context, as in the example below.
Working out how often to submit
The worst case for one item is: it arrives just after a batch was submitted, waits for the next one, then takes the full 24 hours to process. Add the time your own code needs afterwards.
Worst case = submission interval + 24 hours + your post-processing time
Example: the team promises results within 36 hours of upload and needs 2 hours to load results into its database. That leaves 36 minus 24 minus 2, so 10 hours at most between submissions. Submitting every 8 hours, at fixed times, leaves a 2-hour margin.
The guide's own skill line works the same way: short submission windows are what make a deadline safe when processing alone can take 24 hours.
Submit, collect and resubmit failures (Python)
import time
import anthropic
client = anthropic.Anthropic()
MODEL = "your-model-id"
def build_request(doc_id, text):
return {
"custom_id": f"doc-{doc_id}",
"params": {
"model": MODEL,
"max_tokens": 1024,
"system": [{
"type": "text",
"text": SYSTEM_PROMPT, # identical in every request
"cache_control": {"type": "ephemeral", "ttl": "1h"},
}],
"messages": [{"role": "user", "content": text}],
},
}
batch = client.messages.batches.create(
requests=[build_request(doc_id, text) for doc_id, text in documents.items()]
)
while True:
batch = client.messages.batches.retrieve(batch.id)
if batch.processing_status == "ended":
break
time.sleep(60)
succeeded, to_fix, to_resend = {}, [], []
for result in client.messages.batches.results(batch.id):
outcome = result.result
if outcome.type == "succeeded":
succeeded[result.custom_id] = outcome.message
elif outcome.type == "errored" and outcome.error.error.type == "invalid_request_error":
to_fix.append(result.custom_id) # change the request first, e.g. split a long document
else:
to_resend.append(result.custom_id) # server error or expired: send again as it is
The next batch contains only the IDs in to_fix (after the change) and to_resend. The documents that succeeded are never sent twice. Before you report a run as finished, check that every custom_id you submitted appears in one of the three groups. A summary that counts only the results it received can read as complete when it is not.
Refine on a sample first
A prompt error repeated across 50,000 requests is 50,000 wasted requests. Run 20 to 50 documents through the synchronous API, check the outputs, fix the prompt, and only then submit the full batch. This also catches parameter errors that a batch would only report once processing ends.
Rules that decide exam answers
- If someone waits for the result, do not batch it. Cost savings never outweigh a blocking workflow with no latency guarantee. And a loop of synchronous calls is not a batch, however you chunk it.
- Match results by
custom_id, never by position. Results come back in any order. - Resubmit only the failures. Use
custom_idto find them, and fix the request first when the error was in the request. - No tool round trips inside a batch request. Multi-turn work with your own tools needs the synchronous API.
- Test on a sample before the full run. It raises first-pass success and avoids paying for a broken prompt at scale.
- Plan for the worst case. Submission interval plus 24 hours plus your own processing must fit inside the SLA.
Where it appears in the exam
Prompt Engineering & Structured Output is a primary domain in two of the six exam scenarios: Claude Code for Continuous Integration and Structured Data Extraction. Choosing between batch and synchronous calls fits the CI scenario (pre-merge checks against nightly jobs). Large document runs, failures by custom_id and SLA timing fit the extraction scenario, and the guide's extraction pipeline exercise includes a batch of 100 documents.
Two sample questions
These are original Timo practice questions. They are not official exam questions.
Build exercise
- Run 10 documents through the synchronous Messages API with your prompt and fix it until every output is right.
- Submit 100 documents as a batch with meaningful
custom_idvalues. Poll until it ends and note how long it took. - Include two bad requests: one with an invalid parameter and one document too long for the context window. Confirm they come back
erroredwhile the rest succeed, then resubmit only those two after fixing them (split the long one). - Write down your team's own SLA calculation: submission interval, 24 hours, post-processing time.
Practise this topic
- Claude Certified Architect practice exam: free, 20 questions, no sign-up
- Claude Certified Architect hub
- CCAR-F study guide: all topics
- Worked example: Claude Batch API
- Same topic in another exam: Claude API Mechanics (CCDV-F)
- Previous topic: 4.4 Validation and retry loops
- Next topic: 4.6 Multi-pass review
Sources
- Claude Certified Architect Foundations Exam Guide, version 1.0, effective July 2026 (Anthropic), task statement 4.5
- Anthropic documentation: Batch processing
- Anthropic API reference: Retrieve Message Batch results
- Anthropic documentation: Prompt caching
By Amotion AI