TimoBy Amotion AI

Model selection and tradeoffs: CCDV-F study guide

CCDV-F · Model Selection and Optimization, topic weight 2.7% of the exam

Model Selection and Tradeoffs sits in Model Selection and Optimization, which is 16.8% of the CCDV-F exam, and carries 2.7% on its own. It tests one decision: which Claude tier to use for a task once you weigh quality, latency and cost, and what to check before you move that task to a newer model.

What the official guide covers

The Claude Certified Developer Foundations exam guide (version 1.0, effective July 2026) describes this topic as Claude model capabilities, the tradeoffs across quality, latency and cost, and behaviour changes across model releases.

What the guide listsWhat it means in practice
Opus vs Sonnet vs Haiku use casesKnow what each tier is for and which kinds of work sit in each
Adaptive thinking supportThinking modes and the effort parameter differ by model; check support before you design around them
Tradeoffs across quality, latency and costPick the cheapest, fastest option that still passes your quality bar, measured on your own tests
Breaking behaviour changes across model releasesA newer model can reject parameters your code sends or return responses in a different shape

What each tier is for

TierWhere it sitsTypical work
OpusThe most capable and the most expensive of the threeLong-running coding agents, large refactors, complex analysis and planning
SonnetBalance of speed and intelligenceCode generation, data analysis, content creation, agentic tool use
HaikuThe fastest and cheapestReal-time features, high-volume classification and extraction, subagents, cost-sensitive work

The exam guide names these three tiers. Anthropic's current models overview also lists a Fable tier above Opus, for the most demanding reasoning and long-horizon agentic work, and suggests it when your tests on Opus at higher effort still fall short. The selection logic on this page is the same for it.

Feature support also differs by model. Today the current Opus and Sonnet models use adaptive thinking and accept the effort parameter, and some of them think by default. The current Haiku model uses extended thinking with a token budget and does not support effort. Check the models overview before you rely on a feature.

Two ways to choose

Anthropic's "Choosing the right model" page describes two starting points.

  • Efficiency first. Start with Haiku, test it on your task, and move up a tier only where it falls short. This suits prototypes, tight latency targets and cost-sensitive work.
  • Capability first. Start with the most capable model to prove the task can be done, then optimise: improve the prompt, lower effort, or move steps to a cheaper tier. This suits complex reasoning and work where accuracy matters more than cost.

Both need a set of test cases with known good answers. Without one, "Sonnet is good enough" is a guess.

You do not need one model for a whole system. Route each step to the tier it needs, such as Haiku for triage, Sonnet for generation and Opus for planning. Effort is a second lever: a lower level on the same model cuts tokens and latency, and the docs note that tuning effort is often a better lever than switching models.

Compare cost per finished task, not price per token

The rate card is only half the sum. What a workload really costs is the tokens each completed task uses, times the rate, plus the cost of the wrong answers that get through.

  • Tokens per task. A stronger model can sometimes finish in fewer turns, retries or tool calls, so measure total usage per completed task on your test set rather than comparing rates.
  • Cost of a mistake. A cheaper tier that saves a little per request is a bad trade if its extra errors cause refunds, rework or manual review downstream. Put that cost in the comparison.
  • Prompt and tier together. Adding a few examples can let a cheaper tier pass; a stronger tier may pass with a shorter prompt. Test the combinations, not the tier alone.

Anthropic's model guide describes two multi-model patterns that bill most tokens at the lower rate: an executor on a cheaper model that escalates hard decisions to a frontier model, and an orchestrator that hands bulk work to cheaper workers. A simpler version is a default model with an override: a cheap signal read from the request (task type, input length, a quick classification) sends the minority of hard requests to a larger tier. If every request has the same shape, drop the router and pin one model.

How quality, latency and cost trade off

For any workload, name the one constraint that would change the answer if you picked a different tier: cost at high volume, quality on hard reasoning where a wrong answer is expensive, a response-time target, or traffic that mixes easy and hard requests. Then let your test results confirm the tier. Defaulting to the most capable model without that evidence is a common and costly mistake.

SituationChooseWhy
High-volume classification with a strict response-time targetHaikuFastest and cheapest; check accuracy on your tests first
Customer-facing chat that needs good writing and tool useSonnetBalance of quality and speed
A coding agent that runs for a long time on a large codebaseOpusStrongest at long, complex agentic work
Same model passes but is too slow or costlyLower effort before changing tierFewer tokens without changing the model
Interactive work on Opus where output speed mattersFast mode, if your account has accessSame model, faster output, higher price
One step needs deep reasoning, the rest are simpleMixed tiers, routed per stepPay for capability only where it is used
An orchestrator fans work out to many parallel workersA capable lead and cheaper workersEvery worker spends its own tokens, so the worker tier sets most of the bill
A cheaper tier passes most tests but its misses are costly to fixKeep the stronger tier, or route only risky cases to itError cost belongs in the comparison

Breaking changes across model releases

Model IDs are fixed. Anthropic does not change the weights or configuration behind an existing model ID; a new version ships under a new ID. Newer models use dateless IDs that each map to one fixed snapshot. Some older models also have aliases that point to the latest snapshot of that version, so an alias can change under you. Pin an exact ID in production configuration and change it on purpose. The docs also note that serving infrastructure can cause small differences in behaviour even when the ID stays the same.

A new model is a new dependency, and recent migration guides list changes that break working code:

Change in newer modelsWhat breaksFix
Non-default temperature, top_p or top_k rejectedRequests fail with a 400 errorRemove the parameters; steer with the prompt
Extended thinking with budget_tokens rejectedThinking requests fail with a 400 errorUse adaptive thinking and set effort
Prefilling the assistant turn rejectedPrefill-based JSON tricks failUse structured outputs or system prompt instructions
Forced tool_choice (any or tool) rejected on some modelsForced tool calls failUse auto with strict tools and say in the prompt when the tool applies
New tokenizerThe same text uses more tokens, so cost, max_tokens headroom and compaction triggers shiftCount again and add headroom
Responses can start with thinking blocksCode that reads content[0].text breaksSelect blocks by their type
Default effort changedCost and quality shift with no code changeSet effort explicitly and rerun your evals
New stop reasons such as refusalUnhandled responses look like successHandle every stop_reason

A model config that survives upgrades (Python)

import os
import anthropic

client = anthropic.Anthropic()

# Exact model IDs live in configuration, one per job, and change only through review.
MODELS = {
    "triage": os.environ["TRIAGE_MODEL_ID"],      # a Haiku model
    "drafting": os.environ["DRAFTING_MODEL_ID"],  # a Sonnet model
    "planning": os.environ["PLANNING_MODEL_ID"],  # an Opus model
}

INCOMPLETE = {"max_tokens", "refusal", "model_context_window_exceeded"}

def ask(job: str, prompt: str, effort: str | None = None) -> str:
    extra = {"output_config": {"effort": effort}} if effort else {}  # not every model supports effort
    response = client.messages.create(
        model=MODELS[job],
        max_tokens=8000,
        messages=[{"role": "user", "content": prompt}],
        **extra,
    )
    if response.stop_reason in INCOMPLETE:
        raise RuntimeError(f"{job}: incomplete response ({response.stop_reason})")
    # Newer models can return thinking blocks first, so never read content[0].
    return "".join(b.text for b in response.content if b.type == "text")

Before you change any entry in MODELS, run your test set on the old and new IDs side by side and compare quality, latency, token counts and errors.

Rules that decide exam answers

  • Choose the cheapest tier that passes your tests. "Always use the most capable model" ignores latency and cost; "always use the cheapest" ignores quality.
  • Decide from measurements, not reputation. The right answer runs the task on candidate tiers against an eval set.
  • Pin exact model IDs in production. A model upgrade is a deliberate change, tested like any other dependency upgrade.
  • Read the migration guide before you switch. Rejected parameters, new response shapes and tokenizer changes break code, not only answers.
  • Effort and tier are separate levers. Lowering effort can fix cost or latency without a weaker model.
  • Route only when traffic is mixed. A default tier with an override for hard requests fits mixed traffic; uniform traffic gets one pinned model and no router.

Where it appears in the exam

Model Selection and Tradeoffs is 2.7% of the exam, inside Model Selection and Optimization (16.8%). The guide's description points to questions that give you a workload with constraints, such as volume, a response-time target, an accuracy need or a budget, and ask which tier fits, and to questions about an application that misbehaves after a move to a newer model.

Two sample questions

These are original Timo practice questions. They are not official exam questions.

Question 1

A utility company classifies 200,000 incoming emails a day into eight categories, and each label must be ready within a second. On the team's test set of 500 labelled emails, Haiku meets the accuracy target and Opus scores slightly higher. Which choice fits the requirements?

Answer: A. Haiku passes the measured quality bar and best fits the volume and latency constraints. B ignores cost and latency, C is wrong because fast mode costs more, not less, and D is wrong because higher effort means more work per response, not faster responses.

Question 2

A team moves a working application to a newer Sonnet model by changing only the model ID. Most requests now fail with 400 errors. The code sends temperature=0.2, and in the responses that succeed the first content block is sometimes a thinking block, which the app shows to users as a blank answer. What should the team change?

Answer: C. Newer models reject non-default sampling values with a 400 error, and responses can begin with thinking blocks, so code must select blocks by type. A avoids the fix and leaves the app on an older model, B treats a parsing bug as truncation, and D retries a request that will fail every time.

Build exercise

  1. Write 30 test inputs with known good outputs for one task, such as ticket triage.
  2. Run the set on Haiku, Sonnet and Opus. Record accuracy, time per request and usage tokens, then pick the cheapest tier that passes.
  3. On a tier that supports effort, run the set again at two effort levels and compare.
  4. Read the migration guide for the newest model in your chosen tier. List every change that would break your code, and fix any place that reads response content by position.

Practise this topic

Sources