TimoBy Amotion AI

Models, Prompting and Context: CCAR-P domain 2 study guide

CCAR-P · Claude Models, Prompting & Context Engineering (13% of the exam)

Domain 2 is Claude Models, Prompting & Context Engineering, 13% of the CCAR-P exam. It tests design trade-offs: which model and effort level each step needs, how prompts are structured and reused, and how to keep the context window holding only what the task needs.

What the official guide covers

The Claude Certified Architect Professional exam guide (version 1.0, effective July 2026) lists five tasks under Domain 2, "Claude Models, Prompting & Context Engineering":

What the guide listsWhat it means in practice
Select Claude models based on trade-offsMatch each step to a tier and effort level using capability, speed and cost, then confirm with tests
Design system prompts, templates and guardrailsSeparate fixed instructions from variable input, and know which rules a prompt cannot guarantee
Apply zero-shot, few-shot and chain-of-thought techniquesChoose between instructions alone, worked examples and extra reasoning for each task
Optimise context windows and manage token usageDecide what goes into each request and remove what does not earn its place
Implement prompt reuse (caching, modular prompts, Skills)Build prompts from shared, versioned parts and load specialist instructions only when needed

Choosing a model and an effort level

Anthropic's model selection guidance weighs capability, speed and cost. Either start with the fastest model and move up where tests show a gap, or start with the most capable and move work down once proven. In general terms, Opus suits complex reasoning and agentic work, Sonnet balances capability and speed, and Haiku gives the lowest latency for high volumes.

A second lever sits inside one model. The effort setting (output_config.effort: low, medium, high, xhigh or max) controls how many tokens Claude spends on text, tool calls and thinking. Anthropic notes that tuning effort is often a better lever than switching models. Not every model supports it.

Model IDs are fixed: an updated model ships under a new ID, so an upgrade is a release you test.

Step in the systemStart withWhy
Classify or route high volumes of short requestsHaiku tierLowest latency; simple decision
Draft replies, extract fields, everyday agent workSonnet tier, effort tuned per stepBalance of quality and speed
Hard reasoning, complex agent tasks, high-stakes reviewOpus tier, higher effortErrors cost more than tokens
Same step, quality slightly shortRaise effort before changing tierKeeps tested behaviour

Leaving the tier undecided is itself a decision: every step then runs on whatever the team used in the demo, often the most expensive tier, including steps such as a router that need no reasoning. Choose per step, and treat any change of tier as a release:

  1. Build a test set that mirrors real traffic, grouped by request type so each group has enough cases to score on its own.
  2. Run old and new configurations through the same grader.
  3. Write the pass rule before you run it, per group as well as overall (for example, "no group drops more than two points").

The per-group rule matters. A cheaper tier can pass on average while failing two document types; the answer is then a partial move, routing those types to the stronger tier, which an average score would hide.

System prompts, templates and guardrails

Build a system prompt from parts that change at different rates:

  1. Role and task, and what a good result is.
  2. Rules, including how to refuse.
  3. Reference material: policies, product data, examples.
  4. Output format.

Variable input goes in the user turn, wrapped in tags such as <document> or <question> so Claude can tell instructions from data. A template is this structure with named slots, kept in version control.

Design the template so the fixed scaffolding carries every guardrail; filling a slot should never be able to remove a constraint. Then read the prompt for what it leaves unsaid. For each requirement, check that the prompt states the scope, the exact output format and the constraints, rather than hoping for them: wherever the prompt is silent, Claude fills the gap with its own assumption, and not the same one each time. Put a key guardrail into the output contract as well as the rules, for example a required field that lists any commitment made, which code can check. Write neutrally and balance few-shot examples across the cases you expect, since leading phrasing and one-sided examples skew results in ways no single output reveals.

Prompt guardrails steer behaviour; they do not guarantee it. A rule that must hold every time belongs in code: validation, a permission or a hook (see 1.5 Agent SDK hooks and Governance, Safety and Risk).

Zero-shot, few-shot or more thinking?

SituationChooseWhy
Task is clear and the format simple: classify into named categories, summariseZero-shot with explicit criteriaFewest tokens; step-by-step reasoning adds cost here, not quality
Pattern hard to describe, or edge cases keep failingFew-shot: 3 to 5 varied examplesExamples show the boundary better than rules
Multi-step reasoning, calculations, comparing sourcesThinking enabled, effort set to matchClaude reasons before answering
Thinking helps but latency mattersLower effort, or thinking only on the hard routeReasoning spent where it pays

On current models, thinking is set with thinking={"type": "adaptive"} and its depth with the effort setting; older models take {"type": "enabled", "budget_tokens": ...}. Thinking tokens are billed as output. Measure without thinking first, fix the prompt, and enable thinking only where the test set still shows a gap. For the craft, see 4.1 Prompts with explicit criteria and 4.2 Few-shot prompting.

Managing the context window

Everything counts toward the context window: system prompt, tool definitions, messages, tool results and output including thinking. Two different failures happen at the edge. If the input alone exceeds the window, the API rejects the request with an error whatever max_tokens is. If the input fits but generation runs into the limit, the response succeeds with stop_reason set to model_context_window_exceeded and the output cut short. Size the budget for the longest realistic conversation, the retrieved material and a margin, not the full window. The levers:

  • Measure first. The token counting API tells you what a request will cost before you send it.
  • Send less. Retrieve the relevant passages instead of whole documents; return concise tool results.
  • Load tools on demand. The tool search tool keeps rarely used definitions out of context (see Integration).
  • Clear old material. Context editing can clear old tool results in long agent runs, and server-side compaction (in beta) summarises earlier turns.
  • Split the work. A subagent explores in its own context and returns a summary.

Real systems combine strategies. Ask four separate questions: what the model needs at the start (a small, stable, cacheable prefix), what it needs from recent steps (carry those forward), what it might fetch on demand (retrieval or tools), and what older material can be summarised without losing identifiers, figures or decisions (compaction). For conversation-level detail, see 5.1 Context across long conversations.

Reusing prompts: caching, modules and Skills

Prompt caching reuses a repeated prefix, built in the order tools, then system, then messages. Stable content must come first and variable content last. Mark the end of the cacheable part with cache_control, or set it once at the top level. The default lifetime is five minutes, with a one-hour option. Changing tool definitions invalidates the whole cache, so a design that edits tools per request gains nothing. usage reports cache_creation_input_tokens and cache_read_input_tokens, so you can track the hit rate. Caching is not free: a cache write costs more than ordinary input, prompts below a model-specific minimum length cannot be cached, and a prompt called less often than the cache lifetime pays for writes it never reads. For cost mechanics, see Cost and Token Management.

Modular prompts keep shared parts (tone, refusals, formats) in one versioned place, so one fix reaches every assistant.

Skills package instructions, scripts and files in a folder with a SKILL.md whose frontmatter needs a name and description. Only those two load at startup; the body loads when a request matches the description, and bundled files only when referenced. On the API, Skills run in the code execution tool's container; in Claude Code, Claude finds them on its own. See 3.2 Slash commands and skills.

Reuse needPrompt librarySkill
How it is usedAssembled and adjusted per useSame procedure run the same way each time
Who shares itOne codebase or teamSeveral teams or products
ControlEngineers own the fragmentsVersioned, reviewed, can be rolled back

Example: a reusable, cache-friendly request

import anthropic

client = anthropic.Anthropic()

# Versioned prompt modules, loaded from the repository
ROLE_AND_RULES = load_module("support/role_and_rules@v12")
POLICY_HANDBOOK = load_module("support/policy_handbook@v31")   # large, changes monthly

SYSTEM = [
    {"type": "text", "text": ROLE_AND_RULES},
    {"type": "text", "text": POLICY_HANDBOOK,
     "cache_control": {"type": "ephemeral"}},     # cache everything up to here
]

def answer(question: str, effort: str = "low"):
    return client.messages.create(
        model=SONNET_MODEL_ID,                     # pinned ID from config
        max_tokens=2048,
        system=SYSTEM,
        messages=[{"role": "user",
                   "content": f"<question>{question}</question>"}],
        output_config={"effort": effort},
    )

resp = answer("Can I carry unused leave into next year?")
print(resp.usage.cache_read_input_tokens, resp.usage.cache_creation_input_tokens)

Fixed modules come first and the question last, so later requests read the handbook from cache. Changing effort invalidates only the message cache, not the cached system prompt.

Rules that decide exam answers

  • Test before you switch models. A tier change without an evaluation set is a guess.
  • Try effort before a new tier. When quality is slightly short on one step, raising effort keeps the tested prompt and behaviour.
  • Stable content first, variable content last. Any design that puts the user's message or a timestamp before the static prompt defeats caching.
  • Prompts steer; code enforces. A rule that must never break needs a check in code, not stronger wording.
  • Examples beat longer rules for format and edge cases. Add varied examples rather than more instructions.
  • Cut context before buying a bigger window. Retrieval, concise tool results, tool search and clearing old results fix the cause.

Where it appears in the exam

Claude Models, Prompting & Context Engineering carries 13% of the CCAR-P exam. The guide's own sample item for this domain asks how to cut latency and cost for a large, repeated system prompt. Expect cost or latency targets, a quality gap on one request type, growing prompts, or context running out in long sessions.

Two sample questions

These are original Timo practice questions. They are not official exam questions.

Question 1

A contract assistant runs every request on the most capable model tier. About 70% of requests are simple clause lookups; the rest are multi-clause risk analyses. Users complain about slow answers to lookups. Quality on risk analyses must not drop. What should the architect do?

Answer: A. Routing puts each request type on the model it needs, and testing both routes protects analysis quality. B risks the analyses, C does not speed up token generation, and D makes every lookup longer without proof the smaller tier can do analyses.

Question 2

A support agent loads 60 tool definitions and keeps every tool result in the conversation. In long sessions requests fail with a "prompt is too long" error. Which change fixes the cause while keeping the information the agent needs?

Answer: C. The input has grown too large, so removing stale results and unused definitions fixes the cause. A fails because the input alone exceeds the window, B removes the agent's instructions, and D discards facts the user already gave.

Build exercise

  1. List every part of one assistant's prompt and mark each as fixed, monthly or per request.
  2. Put fixed parts first, add a cache breakpoint after them, run ten requests and record cache_read_input_tokens.
  3. Run 20 test requests at two effort levels on two tiers, and write a decision matrix choosing tier and effort per request type.
  4. Move one set of specialist instructions into a Skill with a clear description, and check that it loads only for matching requests.

Practise this topic

Sources