Cost and token management: CCDV-F study guide
CCDV-F · Model Selection and Optimization, topic weight 2.8% of the exam
Cost and Token Management sits in Model Selection and Optimization, which is 16.8% of the CCDV-F exam, and carries 2.8% on its own. It tests whether you can measure what a Claude application spends in tokens and cut that spend, above all with prompt caching placed at the right point in the prompt.
What the official guide covers
The Claude Certified Developer Foundations exam guide (version 1.0, effective July 2026) describes this topic as token budgeting and cost management, including usage tracking, cost modelling and caching techniques such as prompt caching and "cache check-pointing".
| What the guide lists | What it means in practice |
|---|---|
| Token budgeting | Know what each request carries, cap output with max_tokens and effort, and count tokens before you send |
| Token usage tracking | Read the usage object on every response and aggregate it per feature, model and customer |
| Cost modelling | Multiply each kind of token by its rate and by expected volume |
| Prompt caching | Reuse an identical prompt prefix across requests instead of processing it again |
| Cache check-pointing | Place cache_control breakpoints so the stable part of the prompt is cached |
Where token usage is reported
Every Messages API response has a usage object. With caching on, input is split three ways:
| Field | What it counts |
|---|---|
input_tokens | Input tokens after the last cache breakpoint, which were neither read from nor written to the cache |
cache_creation_input_tokens | Tokens written to the cache by this request |
cache_read_input_tokens | Tokens read from the cache by this request |
output_tokens | Tokens generated, including thinking |
Total input is the sum of the first three. A dashboard that logs only input_tokens will show input "dropping" when caching starts, even though the prompt did not change. Thinking tokens are billed as output, and usage.output_tokens_details.thinking_tokens reports how many of the output tokens were thinking.
Three more places to measure:
- Before sending:
client.messages.count_tokens(...)takes the samemodel,system,toolsandmessagesas a real request and returnsinput_tokens. The docs call the result an estimate. - Across the organisation: the Usage and Cost Admin API reports historical usage and cost, grouped by model, workspace, API key and more, with uncached input, cache writes, cache reads and output listed separately. It needs an Admin API key, not a normal workspace key.
- In Claude Code:
/usage(alias/cost) shows session cost and usage.
A cost model without guesswork
Cost per request is a sum of four products:
uncached input tokens × input rate
+ cache write tokens × cache write rate
+ cache read tokens × cache read rate
+ output tokens × output rate
Rates differ by model. Cache writes cost more than normal input and cache reads cost much less, so caching pays off when a prefix is read several times within its lifetime. A write that is never read costs more than not caching. Multiply by requests per day to get a budget, and take current rates from Anthropic's pricing page rather than hard-coding them.
Other levers that change the sum: a cheaper tier, lower effort (fewer output and thinking tokens), a lower max_tokens cap, shorter tool results, fewer tool calls per task, and the Message Batches API for work that can wait (see CCAR-F 4.5 Message Batches API). Batch and caching discounts stack, so a scheduled job that sends the same long system prompt with thousands of requests benefits from both. Note that a newer model's tokenizer can produce more tokens for the same text, so re-count after an upgrade.
Two patterns multiply the sum quietly. Few-shot examples are resent and billed on every call, so a long, stable example block belongs in the cached prefix. And when an orchestrator fans a task out to parallel subagents, each subagent spends its own tokens in its own context, so the bill scales with the number of workers. That only pays off when the task really splits into independent parts.
Find the expensive step before you optimise
A monthly invoice tells you the total, not the cause. Log three numbers on every call, tagged with the step that made it: tokens (from usage), latency and errors. Then sort by step. Often one step, such as a tool that returns a whole file, accounts for most of the spend, and that is where to work.
import time
import anthropic
client = anthropic.Anthropic()
def tracked_call(step: str, **request):
start = time.perf_counter()
try:
response = client.messages.create(**request)
except anthropic.APIError as err:
log_metric(step=step, error=type(err).__name__)
raise
u = response.usage
log_metric(
step=step,
latency_ms=round((time.perf_counter() - start) * 1000),
input=u.input_tokens,
cache_write=u.cache_creation_input_tokens or 0,
cache_read=u.cache_read_input_tokens or 0,
output=u.output_tokens,
stop_reason=response.stop_reason,
)
return response
Set the reliability floor before you cut cost: a latency target, a retry budget and a minimum score on your test set. A saving that pushes any of those below the floor swaps a visible cost for silent failures, so it does not ship.
How prompt caching works
- The prompt is cached as a prefix, in a fixed order:
tools, thensystem, thenmessages. A cache hit needs every block up to and including the breakpoint to be identical. - You mark where the cached prefix ends. Either put one
cache_controlfield at the top level of the request (automatic caching, where the breakpoint moves to the last cacheable block as a conversation grows), or putcache_controlon individual blocks. You can set up to 4 explicit breakpoints. - The cache lives 5 minutes by default and is refreshed each time it is read.
{"type": "ephemeral", "ttl": "1h"}gives a one-hour lifetime for prompts reused less often than every 5 minutes. Longer-lived entries must come before shorter ones. - There is a minimum size. Prompts below the model's minimum cacheable length are processed without caching, and no error is returned. The minimum varies by model.
- Each breakpoint looks back at most 20 blocks for an earlier cache entry, and only finds entries that earlier requests wrote at their own breakpoints.
Changes invalidate the cache from the level where they happen downwards. Changing a tool definition invalidates everything. Changing tool_choice, adding or removing images, changing the thinking configuration or the effort level, switching fast mode on or off, or changing output_config.format also breaks cache hits. Caching never changes the response: output is identical with or without it. Caches are isolated between organisations, and on the Claude API also between workspaces.
Caching a long system prompt (Python)
import anthropic
client = anthropic.Anthropic()
POLICY_MANUAL = open("policy_manual.md").read() # large and identical on every call
def answer(question: str):
response = client.messages.create(
model=MODEL,
max_tokens=1024,
system=[
{"type": "text", "text": "You answer staff questions using the policy manual."},
{
"type": "text",
"text": POLICY_MANUAL,
"cache_control": {"type": "ephemeral"}, # breakpoint: cache up to here
},
],
messages=[{"role": "user", "content": question}], # changes on every call
)
u = response.usage
print(f"written={u.cache_creation_input_tokens or 0} "
f"read={u.cache_read_input_tokens or 0} "
f"uncached={u.input_tokens} output={u.output_tokens}")
return response
# Check the prefix is above the model's minimum cacheable length.
count = client.messages.count_tokens(
model=MODEL,
system=POLICY_MANUAL,
messages=[{"role": "user", "content": "test"}],
)
print(count.input_tokens)
On the first call, written is large and read is 0. On a second call within 5 minutes, read is large and written is near 0. If read stays at 0, something before the breakpoint is changing. The classic cause is a timestamp or request ID at the start of the system prompt. Setting max_tokens=0 writes the cache without generating output, which lets you warm it before traffic arrives. This is rejected inside a batch request.
Which technique fits
| Situation | Choose | Why |
|---|---|---|
| The same long system prompt or document on every request | Explicit breakpoint on the last static block | Static content is cached; the changing question sits after it |
| A multi-turn chat that grows each turn | Top-level cache_control (automatic caching) | The breakpoint moves forward with the conversation |
| Users come back after 10 to 30 minutes | 1-hour TTL | The 5-minute entry would expire between turns |
| A prompt below the model's minimum cacheable length | Do not rely on caching | It is processed without caching, silently |
| 50,000 documents to process by tomorrow | Message Batches API | Lower cost for work that is not urgent |
| A simple, high-volume step is too expensive | Cheaper tier or lower effort | Fewer or cheaper tokens per request |
| The bill rose and nobody knows which step caused it | Per-call logging tagged by step | You cannot cut a cost you have not located |
| A nightly job reuses one long prompt across thousands of requests | Batches plus caching | The two discounts stack |
| A user-facing reply must feel instant | Streaming | It shortens the wait before the first words appear; it does not lower cost |
| A single-fact lookup against a stable reference set | One call on the smallest tier that passes | Fan-out or extra agent turns pay for work the task never needed |
| A task whose steps each depend on the last | One agent with good context, not parallel subagents | Workers would wait on each other while each spends its own tokens |
Rules that decide exam answers
input_tokensis not total input. With caching, addcache_creation_input_tokensandcache_read_input_tokens.- Put the breakpoint on the last block that never changes. Dynamic content (timestamps, user input) goes after it, never before.
- Anything that changes before the breakpoint breaks the cache. That includes tool definitions,
tool_choice, images, thinking settings and effort. - Caching changes cost and latency, never the answer. An option claiming caching makes output more consistent is wrong.
- Match the lever to the constraint. Streaming cuts perceived latency, not cost; batching cuts cost, not latency; fan-out pays only when the work splits into independent parts.
- Below the minimum length, nothing is cached. No error tells you, so check
cache_creation_input_tokensafter the first call.
Where it appears in the exam
Cost and Token Management is 2.8% of the exam, inside Model Selection and Optimization (16.8%). The guide's description points to questions about tracking token usage, estimating what a workload will cost, and placing cache checkpoints, usually as an application whose bill is too high or whose cache never hits.
Two sample questions
These are original Timo practice questions. They are not official exam questions.
Build exercise
- Put a document of several thousand tokens in the system prompt with a breakpoint, call twice within 5 minutes, and log all four
usagefields. - Add a timestamp before the breakpoint and watch cache reads drop to 0. Move it after the breakpoint and confirm reads return.
- Write a cost function that takes
usageand a dictionary of rates from Anthropic's pricing page. Compare 100 requests with and without caching. - Compare
count_tokensfor your prompt with theusagefigures from the real call.
Practise this topic
- Claude Certified Developer practice exam: free, 20 questions, no sign-up
- CCDV-F study guide: all topics
- Same topic in another exam: CCAR-P Models, Prompting and Context
- Previous topic: Model Selection and Tradeoffs
- Next topic: Context Engineering
Sources
- Claude Certified Developer Foundations Exam Guide, version 1.0, effective July 2026 (Anthropic), topic: Cost and Token Management (Domain 5, Model Selection and Optimization)
- Anthropic documentation: Prompt caching
- Anthropic documentation: Token counting
- Anthropic documentation: Usage and Cost API
By Amotion AI