LLM fundamentals: CCDV-F study guide
CCDV-F · Model Selection and Optimization, topic weight 5.2% of the exam
LLM Fundamentals sits in Model Selection and Optimization, which is 16.8% of the CCDV-F exam, and carries 5.2% on its own. It tests whether you can predict how Claude will behave from how it works: what a token is, what fills the context window, why the same prompt gives different answers, and which thinking and effort settings fit a task.
What the official guide covers
The Claude Certified Developer Foundations exam guide (version 1.0, effective July 2026) describes this topic as a basic understanding of LLMs, the model options you can set, and fundamental prompting techniques.
| What the guide lists | What it means in practice |
|---|---|
| Tokens and next-token generation | Claude reads and writes tokens, generating one at a time |
| Context windows | The prompt, tools and response must fit in one window; more context is not always better |
| Sampling and non-determinism | Each next token is picked from a probability distribution, so one prompt can give different outputs |
| Fast mode | The same model with faster output, at a higher price, on some Opus models |
| Extended thinking, adaptive thinking, effort levels | Controls for whether Claude reasons before it answers and how much work it puts in |
| Zero-shot, single-shot, multi-shot prompting | Giving no example, one example or several examples of the task |
How Claude produces text
Claude is an autoregressive model: it predicts the next token from all the tokens before it, adds it, and repeats. Three things follow.
- Tokens are the unit of everything. Context limits,
max_tokens, billing and theusageobject are counted in tokens, usually pieces of words. Tokenizers differ between models, so count again when you change models. - Each token depends on what came before. An early mistake in the output carries forward. This is why reasoning before the answer, or thinking, helps multi-step problems.
- Generation always stops for a stated reason. The response's
stop_reasontells you which:end_turn(finished),max_tokens(hit your limit),stop_sequence,tool_use,refusal, or, on newer models,model_context_window_exceeded.
What fills the context window
The context window is everything Claude can reference while it generates a response, including the response itself. The system prompt, every message (tool results, images and documents included), the tool definitions and the output, thinking included, all count.
Three behaviours to know:
- Thinking from earlier turns. On newer Opus and Sonnet models the API keeps previous thinking blocks by default, and they count as input. On older models and on Haiku models the API strips them automatically.
- Context rot. Anthropic's docs state that accuracy and recall degrade as the token count grows. A large window is room, not a reason to fill it. The Context Engineering topic covers what to remove.
- Overflow has two edges. If the input alone is larger than the window, every model rejects the request with a 400 error ("prompt is too long") before any generation. If the input fits but input plus
max_tokensexceeds the window, newer models accept the request and stop withmodel_context_window_exceededif generation reaches the limit; older models return a validation error instead. Either way, the API does not trim old turns for you, and raisingmax_tokensdoes not help: it caps only what Claude writes, not what it can read. Keeping a long session alive is your application's job, and the token counting endpoint lets you check the size before you send.
Sampling and non-determinism
At each step the model has a probability for every possible next token, and sampling picks one. So two calls with the same prompt can return different wording or, now and then, a different answer.
| Control | What it does | Status today |
|---|---|---|
temperature | Higher gives more varied output; lower gives more conservative output | Older models accept 0.0 to 1.0, default 1.0. Newer models reject any non-default value with a 400 error |
top_p, top_k | Cut off low-probability tokens | For advanced use only on older models. Newer models reject non-default values |
The API reference says that even at temperature 0.0, results are not fully deterministic, and on newer models Anthropic recommends prompting instead of sampling settings. Design for variation: test prompts over many runs, use structured outputs when code parses the result, and validate in code.
Testing a feature whose output varies
A test that compares Claude's reply to a fixed string will fail on a correct answer worded differently. Test the property that must hold instead, and pick the grader by the shape of the output:
| Output shape | Grade with | Example check |
|---|---|---|
| One correct form | Exact match | The label equals BILLING |
| Structure or a rule | Code | The JSON parses, a required field is present, a number is in range |
| Meaning or quality | A second model with a rubric | "Does the summary name the issue and its current status?" |
Code-based grading is the fastest and most reliable, so use it wherever a rule can decide. For model-based grading, write a detailed rubric and have the grader reason before it scores. Check a model grader against cases a person has labelled before you trust its numbers. Because one run proves little, run each case several times and track the pass rate.
Thinking, effort and fast mode
| Option | How you set it | What it does |
|---|---|---|
| Extended thinking | thinking={"type": "enabled", "budget_tokens": N} | Claude reasons in thinking blocks up to a token budget, which must be less than max_tokens. Deprecated on some models and rejected by the newest ones |
| Adaptive thinking | thinking={"type": "adaptive"} | Claude decides per request whether to think and how much. Some newer models use it by default |
| Effort | output_config={"effort": "high"} | Soft guidance on how much work Claude puts into the whole response: text, tool calls and thinking. Levels are low, medium, high, xhigh and max; the default depends on the model |
| Fast mode | speed="fast" with a beta header | Research preview on some Opus models. Same model and capabilities, more output tokens per second, premium pricing |
What the exam can test:
- You pay for thinking tokens at the output rate, and they count toward
max_tokens, even when the thinking text is left out of the response. Omitting it cuts latency, not cost. - In a tool-use loop, pass thinking blocks back complete and unmodified. Each one carries a
signature. - Where adaptive thinking is available, effort is the recommended control for thinking depth. Anthropic suggests
lowfor simple, speed-sensitive work such as subagents. - Changing the thinking configuration or effort between requests invalidates prompt cache breakpoints.
- Fast mode raises output speed, not time to first token, and is not available on the Batch API.
import anthropic
client = anthropic.Anthropic()
response = client.messages.create(
model=MODEL, # a model that supports adaptive thinking
max_tokens=16000, # hard cap on thinking plus answer
thinking={"type": "adaptive"},
output_config={"effort": "high"},
messages=[{"role": "user", "content": "Find the bug in this function: ..."}],
)
for block in response.content: # read blocks by type, not by position
if block.type == "text":
print(block.text)
print(response.stop_reason, response.usage.output_tokens)
Zero-shot, single-shot and multi-shot prompts
| Approach | What you give Claude | Use it when |
|---|---|---|
| Zero-shot | Instructions only | The task is common and the format is simple |
| Single-shot | Instructions and one example | You need one exact format and the inputs look alike |
| Multi-shot (few-shot) | Several varied examples in <example> tags | Output must follow a pattern across varied inputs, or edge cases matter |
A single example can lead Claude to copy its details too closely, so use several varied ones when inputs vary.
Examples are not training. They sit in the prompt, so every example is sent and billed as input on every call. Start zero-shot, and add examples only when a description keeps missing the format, casing or an edge case. A long, stable example block is a good candidate for prompt caching.
Three separate levers
Model tier, thinking settings and the number of examples are set independently, and they trade against each other. A more capable model may handle a task zero-shot where a smaller one needs three examples, and a few examples can let a cheaper tier pass your tests. Thinking can be on or off on either. Try the smallest combination that passes your test set, and add capability, effort or examples only where the results say you need them.
Which setting fits the problem
| Situation | Choose | Why |
|---|---|---|
| Hard multi-step reasoning where accuracy matters most | Adaptive thinking with higher effort | Claude reasons before it answers and between tool calls |
| High-volume simple classification | Lower effort and a short prompt | Fewer output tokens; extra reasoning adds cost and latency for little gain |
| An agent plans a chain of dependent tool calls | Thinking on, with thinking blocks returned unchanged in the loop | Reasoning before the plan cuts wrong early steps; edited or dropped blocks within a tool-use turn are rejected |
| A tuned prompt still misses your accuracy target on tests | Turn thinking on or raise effort, then measure again | Fix the prompt first; thinking is the next lever, not the first |
| The same prompt gives different answers across runs | Clearer criteria, examples, structured output and validation | Sampling cannot be made fully deterministic |
| Interactive coding where output speed matters | Fast mode, if your account has access | Faster output from the same model, at a higher price |
Rules that decide exam answers
- Temperature 0 does not mean deterministic. Newer models also reject non-default sampling values. An option that "sets temperature to 0 to guarantee identical output" is wrong.
- Thinking costs output tokens. Hiding the thinking text reduces latency, not cost.
- Tune adaptive thinking with effort.
budget_tokensis the older control, and the newest models reject it. - Read
stop_reason, not the text.max_tokensmeans the output was cut off; treat it as incomplete. - Fast mode is the same model. It changes speed and price, not intelligence. To change capability, change tier.
- Test properties, not wording. An exact-text assertion on open-ended output is a flaky test, not a quality check.
Where it appears in the exam
LLM Fundamentals is 5.2% of the exam, inside Model Selection and Optimization (16.8%). Expect questions that start from a symptom, such as answers that vary between runs, cut-off responses or slow output, and ask which model option or prompting technique explains or fixes it.
Two sample questions
These are original Timo practice questions. They are not official exam questions.
Build exercise
- Send the same open-ended prompt 10 times and note where the outputs differ. Then send
temperature=0.5to a newer model and record the 400 error. - Call
client.messages.count_tokenson the same two-page document with two different models, if you have access, and compare the counts. - Run a multi-step reasoning task with adaptive thinking at effort
low, thenhigh. Recordusage.output_tokens, time taken and correctness. - Set
max_tokensdeliberately low, confirmstop_reasonismax_tokens, and write code that treats that response as incomplete.
Practise this topic
- Claude Certified Developer practice exam: free, 20 questions, no sign-up
- CCDV-F study guide: all topics
- Previous topic: Debugging and Error Handling
- Next topic: Technical Fundamentals
Sources
- Claude Certified Developer Foundations Exam Guide, version 1.0, effective July 2026 (Anthropic), topic: LLM Fundamentals (Domain 5, Model Selection and Optimization)
- Anthropic documentation: Context windows
- Anthropic documentation: Thinking
- Anthropic documentation: Effort
- Anthropic API reference: Create a Message
By Amotion AI