TimoBy Amotion AI

LLM fundamentals: CCDV-F study guide

CCDV-F · Model Selection and Optimization, topic weight 5.2% of the exam

LLM Fundamentals sits in Model Selection and Optimization, which is 16.8% of the CCDV-F exam, and carries 5.2% on its own. It tests whether you can predict how Claude will behave from how it works: what a token is, what fills the context window, why the same prompt gives different answers, and which thinking and effort settings fit a task.

What the official guide covers

The Claude Certified Developer Foundations exam guide (version 1.0, effective July 2026) describes this topic as a basic understanding of LLMs, the model options you can set, and fundamental prompting techniques.

What the guide listsWhat it means in practice
Tokens and next-token generationClaude reads and writes tokens, generating one at a time
Context windowsThe prompt, tools and response must fit in one window; more context is not always better
Sampling and non-determinismEach next token is picked from a probability distribution, so one prompt can give different outputs
Fast modeThe same model with faster output, at a higher price, on some Opus models
Extended thinking, adaptive thinking, effort levelsControls for whether Claude reasons before it answers and how much work it puts in
Zero-shot, single-shot, multi-shot promptingGiving no example, one example or several examples of the task

How Claude produces text

Claude is an autoregressive model: it predicts the next token from all the tokens before it, adds it, and repeats. Three things follow.

  1. Tokens are the unit of everything. Context limits, max_tokens, billing and the usage object are counted in tokens, usually pieces of words. Tokenizers differ between models, so count again when you change models.
  2. Each token depends on what came before. An early mistake in the output carries forward. This is why reasoning before the answer, or thinking, helps multi-step problems.
  3. Generation always stops for a stated reason. The response's stop_reason tells you which: end_turn (finished), max_tokens (hit your limit), stop_sequence, tool_use, refusal, or, on newer models, model_context_window_exceeded.

What fills the context window

The context window is everything Claude can reference while it generates a response, including the response itself. The system prompt, every message (tool results, images and documents included), the tool definitions and the output, thinking included, all count.

Three behaviours to know:

  • Thinking from earlier turns. On newer Opus and Sonnet models the API keeps previous thinking blocks by default, and they count as input. On older models and on Haiku models the API strips them automatically.
  • Context rot. Anthropic's docs state that accuracy and recall degrade as the token count grows. A large window is room, not a reason to fill it. The Context Engineering topic covers what to remove.
  • Overflow has two edges. If the input alone is larger than the window, every model rejects the request with a 400 error ("prompt is too long") before any generation. If the input fits but input plus max_tokens exceeds the window, newer models accept the request and stop with model_context_window_exceeded if generation reaches the limit; older models return a validation error instead. Either way, the API does not trim old turns for you, and raising max_tokens does not help: it caps only what Claude writes, not what it can read. Keeping a long session alive is your application's job, and the token counting endpoint lets you check the size before you send.

Sampling and non-determinism

At each step the model has a probability for every possible next token, and sampling picks one. So two calls with the same prompt can return different wording or, now and then, a different answer.

ControlWhat it doesStatus today
temperatureHigher gives more varied output; lower gives more conservative outputOlder models accept 0.0 to 1.0, default 1.0. Newer models reject any non-default value with a 400 error
top_p, top_kCut off low-probability tokensFor advanced use only on older models. Newer models reject non-default values

The API reference says that even at temperature 0.0, results are not fully deterministic, and on newer models Anthropic recommends prompting instead of sampling settings. Design for variation: test prompts over many runs, use structured outputs when code parses the result, and validate in code.

Testing a feature whose output varies

A test that compares Claude's reply to a fixed string will fail on a correct answer worded differently. Test the property that must hold instead, and pick the grader by the shape of the output:

Output shapeGrade withExample check
One correct formExact matchThe label equals BILLING
Structure or a ruleCodeThe JSON parses, a required field is present, a number is in range
Meaning or qualityA second model with a rubric"Does the summary name the issue and its current status?"

Code-based grading is the fastest and most reliable, so use it wherever a rule can decide. For model-based grading, write a detailed rubric and have the grader reason before it scores. Check a model grader against cases a person has labelled before you trust its numbers. Because one run proves little, run each case several times and track the pass rate.

Thinking, effort and fast mode

OptionHow you set itWhat it does
Extended thinkingthinking={"type": "enabled", "budget_tokens": N}Claude reasons in thinking blocks up to a token budget, which must be less than max_tokens. Deprecated on some models and rejected by the newest ones
Adaptive thinkingthinking={"type": "adaptive"}Claude decides per request whether to think and how much. Some newer models use it by default
Effortoutput_config={"effort": "high"}Soft guidance on how much work Claude puts into the whole response: text, tool calls and thinking. Levels are low, medium, high, xhigh and max; the default depends on the model
Fast modespeed="fast" with a beta headerResearch preview on some Opus models. Same model and capabilities, more output tokens per second, premium pricing

What the exam can test:

  • You pay for thinking tokens at the output rate, and they count toward max_tokens, even when the thinking text is left out of the response. Omitting it cuts latency, not cost.
  • In a tool-use loop, pass thinking blocks back complete and unmodified. Each one carries a signature.
  • Where adaptive thinking is available, effort is the recommended control for thinking depth. Anthropic suggests low for simple, speed-sensitive work such as subagents.
  • Changing the thinking configuration or effort between requests invalidates prompt cache breakpoints.
  • Fast mode raises output speed, not time to first token, and is not available on the Batch API.
import anthropic

client = anthropic.Anthropic()

response = client.messages.create(
    model=MODEL,                         # a model that supports adaptive thinking
    max_tokens=16000,                    # hard cap on thinking plus answer
    thinking={"type": "adaptive"},
    output_config={"effort": "high"},
    messages=[{"role": "user", "content": "Find the bug in this function: ..."}],
)

for block in response.content:          # read blocks by type, not by position
    if block.type == "text":
        print(block.text)

print(response.stop_reason, response.usage.output_tokens)

Zero-shot, single-shot and multi-shot prompts

ApproachWhat you give ClaudeUse it when
Zero-shotInstructions onlyThe task is common and the format is simple
Single-shotInstructions and one exampleYou need one exact format and the inputs look alike
Multi-shot (few-shot)Several varied examples in <example> tagsOutput must follow a pattern across varied inputs, or edge cases matter

A single example can lead Claude to copy its details too closely, so use several varied ones when inputs vary.

Examples are not training. They sit in the prompt, so every example is sent and billed as input on every call. Start zero-shot, and add examples only when a description keeps missing the format, casing or an edge case. A long, stable example block is a good candidate for prompt caching.

Three separate levers

Model tier, thinking settings and the number of examples are set independently, and they trade against each other. A more capable model may handle a task zero-shot where a smaller one needs three examples, and a few examples can let a cheaper tier pass your tests. Thinking can be on or off on either. Try the smallest combination that passes your test set, and add capability, effort or examples only where the results say you need them.

Which setting fits the problem

SituationChooseWhy
Hard multi-step reasoning where accuracy matters mostAdaptive thinking with higher effortClaude reasons before it answers and between tool calls
High-volume simple classificationLower effort and a short promptFewer output tokens; extra reasoning adds cost and latency for little gain
An agent plans a chain of dependent tool callsThinking on, with thinking blocks returned unchanged in the loopReasoning before the plan cuts wrong early steps; edited or dropped blocks within a tool-use turn are rejected
A tuned prompt still misses your accuracy target on testsTurn thinking on or raise effort, then measure againFix the prompt first; thinking is the next lever, not the first
The same prompt gives different answers across runsClearer criteria, examples, structured output and validationSampling cannot be made fully deterministic
Interactive coding where output speed mattersFast mode, if your account has accessFaster output from the same model, at a higher price

Rules that decide exam answers

  • Temperature 0 does not mean deterministic. Newer models also reject non-default sampling values. An option that "sets temperature to 0 to guarantee identical output" is wrong.
  • Thinking costs output tokens. Hiding the thinking text reduces latency, not cost.
  • Tune adaptive thinking with effort. budget_tokens is the older control, and the newest models reject it.
  • Read stop_reason, not the text. max_tokens means the output was cut off; treat it as incomplete.
  • Fast mode is the same model. It changes speed and price, not intelligence. To change capability, change tier.
  • Test properties, not wording. An exact-text assertion on open-ended output is a flaky test, not a quality check.

Where it appears in the exam

LLM Fundamentals is 5.2% of the exam, inside Model Selection and Optimization (16.8%). Expect questions that start from a symptom, such as answers that vary between runs, cut-off responses or slow output, and ask which model option or prompting technique explains or fixes it.

Two sample questions

These are original Timo practice questions. They are not official exam questions.

Question 1

A team's essay-grading service sends the same essay to Claude twice and gets two slightly different scores. A developer proposes changing a sampling setting on their newer model so every run gives an identical score. What should the team do?

Answer: D. Sampling cannot be made fully deterministic: the API reference says results vary even at temperature 0.0, and newer models reject non-default temperature values. Clear criteria and examples reduce the variation that matters. A is false on both counts, B changes how much work Claude does rather than sampling, and C changes speed, not output.

Question 2

A developer enables adaptive thinking for a contract-review feature. To keep costs flat, they set the thinking display to omitted so no thinking text comes back in responses. The bill still rises. Why?

Answer: B. You pay for the full thinking tokens; omitting the text reduces latency, not cost. A names the wrong rate, C has no basis in the docs, and D is wrong because max_tokens is a cap, not a target.

Build exercise

  1. Send the same open-ended prompt 10 times and note where the outputs differ. Then send temperature=0.5 to a newer model and record the 400 error.
  2. Call client.messages.count_tokens on the same two-page document with two different models, if you have access, and compare the counts.
  3. Run a multi-step reasoning task with adaptive thinking at effort low, then high. Record usage.output_tokens, time taken and correctness.
  4. Set max_tokens deliberately low, confirm stop_reason is max_tokens, and write code that treats that response as incomplete.

Practise this topic

Sources

  • Claude Certified Developer Foundations Exam Guide, version 1.0, effective July 2026 (Anthropic), topic: LLM Fundamentals (Domain 5, Model Selection and Optimization)
  • Anthropic documentation: Context windows
  • Anthropic documentation: Thinking
  • Anthropic documentation: Effort
  • Anthropic API reference: Create a Message