TimoBy Amotion AI

Cost and token management: CCDV-F study guide

CCDV-F · Model Selection and Optimization, topic weight 2.8% of the exam

Cost and Token Management sits in Model Selection and Optimization, which is 16.8% of the CCDV-F exam, and carries 2.8% on its own. It tests whether you can measure what a Claude application spends in tokens and cut that spend, above all with prompt caching placed at the right point in the prompt.

What the official guide covers

The Claude Certified Developer Foundations exam guide (version 1.0, effective July 2026) describes this topic as token budgeting and cost management, including usage tracking, cost modelling and caching techniques such as prompt caching and "cache check-pointing".

What the guide listsWhat it means in practice
Token budgetingKnow what each request carries, cap output with max_tokens and effort, and count tokens before you send
Token usage trackingRead the usage object on every response and aggregate it per feature, model and customer
Cost modellingMultiply each kind of token by its rate and by expected volume
Prompt cachingReuse an identical prompt prefix across requests instead of processing it again
Cache check-pointingPlace cache_control breakpoints so the stable part of the prompt is cached

Where token usage is reported

Every Messages API response has a usage object. With caching on, input is split three ways:

FieldWhat it counts
input_tokensInput tokens after the last cache breakpoint, which were neither read from nor written to the cache
cache_creation_input_tokensTokens written to the cache by this request
cache_read_input_tokensTokens read from the cache by this request
output_tokensTokens generated, including thinking

Total input is the sum of the first three. A dashboard that logs only input_tokens will show input "dropping" when caching starts, even though the prompt did not change. Thinking tokens are billed as output, and usage.output_tokens_details.thinking_tokens reports how many of the output tokens were thinking.

Three more places to measure:

  • Before sending: client.messages.count_tokens(...) takes the same model, system, tools and messages as a real request and returns input_tokens. The docs call the result an estimate.
  • Across the organisation: the Usage and Cost Admin API reports historical usage and cost, grouped by model, workspace, API key and more, with uncached input, cache writes, cache reads and output listed separately. It needs an Admin API key, not a normal workspace key.
  • In Claude Code: /usage (alias /cost) shows session cost and usage.

A cost model without guesswork

Cost per request is a sum of four products:

uncached input tokens × input rate
+ cache write tokens   × cache write rate
+ cache read tokens    × cache read rate
+ output tokens        × output rate

Rates differ by model. Cache writes cost more than normal input and cache reads cost much less, so caching pays off when a prefix is read several times within its lifetime. A write that is never read costs more than not caching. Multiply by requests per day to get a budget, and take current rates from Anthropic's pricing page rather than hard-coding them.

Other levers that change the sum: a cheaper tier, lower effort (fewer output and thinking tokens), a lower max_tokens cap, shorter tool results, fewer tool calls per task, and the Message Batches API for work that can wait (see CCAR-F 4.5 Message Batches API). Batch and caching discounts stack, so a scheduled job that sends the same long system prompt with thousands of requests benefits from both. Note that a newer model's tokenizer can produce more tokens for the same text, so re-count after an upgrade.

Two patterns multiply the sum quietly. Few-shot examples are resent and billed on every call, so a long, stable example block belongs in the cached prefix. And when an orchestrator fans a task out to parallel subagents, each subagent spends its own tokens in its own context, so the bill scales with the number of workers. That only pays off when the task really splits into independent parts.

Find the expensive step before you optimise

A monthly invoice tells you the total, not the cause. Log three numbers on every call, tagged with the step that made it: tokens (from usage), latency and errors. Then sort by step. Often one step, such as a tool that returns a whole file, accounts for most of the spend, and that is where to work.

import time
import anthropic

client = anthropic.Anthropic()

def tracked_call(step: str, **request):
    start = time.perf_counter()
    try:
        response = client.messages.create(**request)
    except anthropic.APIError as err:
        log_metric(step=step, error=type(err).__name__)
        raise
    u = response.usage
    log_metric(
        step=step,
        latency_ms=round((time.perf_counter() - start) * 1000),
        input=u.input_tokens,
        cache_write=u.cache_creation_input_tokens or 0,
        cache_read=u.cache_read_input_tokens or 0,
        output=u.output_tokens,
        stop_reason=response.stop_reason,
    )
    return response

Set the reliability floor before you cut cost: a latency target, a retry budget and a minimum score on your test set. A saving that pushes any of those below the floor swaps a visible cost for silent failures, so it does not ship.

How prompt caching works

  1. The prompt is cached as a prefix, in a fixed order: tools, then system, then messages. A cache hit needs every block up to and including the breakpoint to be identical.
  2. You mark where the cached prefix ends. Either put one cache_control field at the top level of the request (automatic caching, where the breakpoint moves to the last cacheable block as a conversation grows), or put cache_control on individual blocks. You can set up to 4 explicit breakpoints.
  3. The cache lives 5 minutes by default and is refreshed each time it is read. {"type": "ephemeral", "ttl": "1h"} gives a one-hour lifetime for prompts reused less often than every 5 minutes. Longer-lived entries must come before shorter ones.
  4. There is a minimum size. Prompts below the model's minimum cacheable length are processed without caching, and no error is returned. The minimum varies by model.
  5. Each breakpoint looks back at most 20 blocks for an earlier cache entry, and only finds entries that earlier requests wrote at their own breakpoints.

Changes invalidate the cache from the level where they happen downwards. Changing a tool definition invalidates everything. Changing tool_choice, adding or removing images, changing the thinking configuration or the effort level, switching fast mode on or off, or changing output_config.format also breaks cache hits. Caching never changes the response: output is identical with or without it. Caches are isolated between organisations, and on the Claude API also between workspaces.

Caching a long system prompt (Python)

import anthropic

client = anthropic.Anthropic()
POLICY_MANUAL = open("policy_manual.md").read()   # large and identical on every call

def answer(question: str):
    response = client.messages.create(
        model=MODEL,
        max_tokens=1024,
        system=[
            {"type": "text", "text": "You answer staff questions using the policy manual."},
            {
                "type": "text",
                "text": POLICY_MANUAL,
                "cache_control": {"type": "ephemeral"},   # breakpoint: cache up to here
            },
        ],
        messages=[{"role": "user", "content": question}],  # changes on every call
    )
    u = response.usage
    print(f"written={u.cache_creation_input_tokens or 0} "
          f"read={u.cache_read_input_tokens or 0} "
          f"uncached={u.input_tokens} output={u.output_tokens}")
    return response

# Check the prefix is above the model's minimum cacheable length.
count = client.messages.count_tokens(
    model=MODEL,
    system=POLICY_MANUAL,
    messages=[{"role": "user", "content": "test"}],
)
print(count.input_tokens)

On the first call, written is large and read is 0. On a second call within 5 minutes, read is large and written is near 0. If read stays at 0, something before the breakpoint is changing. The classic cause is a timestamp or request ID at the start of the system prompt. Setting max_tokens=0 writes the cache without generating output, which lets you warm it before traffic arrives. This is rejected inside a batch request.

Which technique fits

SituationChooseWhy
The same long system prompt or document on every requestExplicit breakpoint on the last static blockStatic content is cached; the changing question sits after it
A multi-turn chat that grows each turnTop-level cache_control (automatic caching)The breakpoint moves forward with the conversation
Users come back after 10 to 30 minutes1-hour TTLThe 5-minute entry would expire between turns
A prompt below the model's minimum cacheable lengthDo not rely on cachingIt is processed without caching, silently
50,000 documents to process by tomorrowMessage Batches APILower cost for work that is not urgent
A simple, high-volume step is too expensiveCheaper tier or lower effortFewer or cheaper tokens per request
The bill rose and nobody knows which step caused itPer-call logging tagged by stepYou cannot cut a cost you have not located
A nightly job reuses one long prompt across thousands of requestsBatches plus cachingThe two discounts stack
A user-facing reply must feel instantStreamingIt shortens the wait before the first words appear; it does not lower cost
A single-fact lookup against a stable reference setOne call on the smallest tier that passesFan-out or extra agent turns pay for work the task never needed
A task whose steps each depend on the lastOne agent with good context, not parallel subagentsWorkers would wait on each other while each spends its own tokens

Rules that decide exam answers

  • input_tokens is not total input. With caching, add cache_creation_input_tokens and cache_read_input_tokens.
  • Put the breakpoint on the last block that never changes. Dynamic content (timestamps, user input) goes after it, never before.
  • Anything that changes before the breakpoint breaks the cache. That includes tool definitions, tool_choice, images, thinking settings and effort.
  • Caching changes cost and latency, never the answer. An option claiming caching makes output more consistent is wrong.
  • Match the lever to the constraint. Streaming cuts perceived latency, not cost; batching cuts cost, not latency; fan-out pays only when the work splits into independent parts.
  • Below the minimum length, nothing is cached. No error tells you, so check cache_creation_input_tokens after the first call.

Where it appears in the exam

Cost and Token Management is 2.8% of the exam, inside Model Selection and Optimization (16.8%). The guide's description points to questions about tracking token usage, estimating what a workload will cost, and placing cache checkpoints, usually as an application whose bill is too high or whose cache never hits.

Two sample questions

These are original Timo practice questions. They are not official exam questions.

Question 1

An internal helpdesk app sends a 30,000-token policy manual in its system prompt on every request. The team added cache_control to the system block, but cache_read_input_tokens is always 0 and cache_creation_input_tokens is high on every call. The first line of the system prompt is "Current time: {now}". What should they fix?

Answer: C. A cache hit needs an identical prefix up to the breakpoint, and the timestamp changes it on every call. A confuses output length with caching, B keeps writing a new entry every call, and D still includes the changing timestamp in every prefix.

Question 2

A finance team builds a dashboard of input tokens per request from response.usage.input_tokens. After the team turns on prompt caching, logged input drops sharply although the prompts have not changed. What should the dashboard record as total input?

Answer: A. With caching, input_tokens covers only tokens after the last breakpoint, so total input is the sum of all three fields. B drops uncached and newly written tokens, C reverses which number is the estimate, and D is false because input is still billed.

Build exercise

  1. Put a document of several thousand tokens in the system prompt with a breakpoint, call twice within 5 minutes, and log all four usage fields.
  2. Add a timestamp before the breakpoint and watch cache reads drop to 0. Move it after the breakpoint and confirm reads return.
  3. Write a cost function that takes usage and a dictionary of rates from Anthropic's pricing page. Compare 100 requests with and without caching.
  4. Compare count_tokens for your prompt with the usage figures from the real call.

Practise this topic

Sources

  • Claude Certified Developer Foundations Exam Guide, version 1.0, effective July 2026 (Anthropic), topic: Cost and Token Management (Domain 5, Model Selection and Optimization)
  • Anthropic documentation: Prompt caching
  • Anthropic documentation: Token counting
  • Anthropic documentation: Usage and Cost API