Systems life cycle: CCDV-F study guide
CCDV-F · Applications and Integration, topic weight 2.8% of the exam
Systems Life Cycle is a 2.8% topic in Applications and Integration, which makes up 33.1% of the CCDV-F exam. It tests whether you can take a Claude system through the usual stages of any IT system, and handle the parts that are specific to Claude: pinned model versions, evaluation gates before release, monitoring in production and model deprecation.
What the official guide covers
The Claude Certified Developer Foundations exam guide (version 1.0, effective July 2026) describes this topic as life cycle concepts and frameworks used to "develop, implement, operate, and maintain IT systems".
| What the guide lists | What it means in practice |
|---|---|
| Develop | Build the prompt, tool definitions and an evaluation set together, all in version control |
| Implement | Release behind an eval gate, with a pinned model ID and a staged rollout |
| Operate | Watch stop reasons, error types, token use and sampled outputs; keep request IDs |
| Maintain | Handle model deprecations, upgrades and prompt changes, re-running evals each time |
| Life cycle frameworks | The usual stages still apply; the model is a versioned dependency, like a library |
The stages of a Claude system
| Stage | What is specific to Claude | What you produce |
|---|---|---|
| Plan | Success criteria that are specific and measurable, and the real inputs, including edge cases, that will test them | Requirements and success criteria |
| Design | Platform, model tier and trust boundaries chosen against the requirements, such as residency and identity | A design record |
| Build | The system prompt, tool schemas and model ID live in config files under version control | A versioned prompt and config |
| Test | Code-graded checks where possible, LLM grading with a clear rubric where needed, people for a sample | An eval set and a report with a pass rate |
| Release | Same eval gate for model, prompt and tool changes; roll out to a small share of traffic first | A gated release |
| Operate | Track stop_reason (a rise in max_tokens or refusal is a signal), error types such as 429 and 529, and usage | Dashboards and alerts |
| Maintain and iterate | Deprecation notices, migrations, prompt fixes; production failures become new eval cases and, where needed, new requirements | A new release through the same gate |
Exam scenarios often ask where a piece of work belongs. A rule that data must be processed in a region is a requirement; choosing the platform that meets it is design. Writing the eval cases and rubric is test work; refusing to promote a version until it meets the pinned baseline is a release decision, as is pinning the model ID and keeping the previous one. Tracking what each production call costs in tokens and how long it takes is operate.
Log the request-id response header (in the Python SDK, response._request_id) with every call. It is the reference Anthropic support needs when you report a problem.
Gates between stages
A gate is the decision to move from one stage to the next. It is where a team, or a regulator, keeps control of a Claude system.
| Gate | Pass condition | What it prevents |
|---|---|---|
| Plan to design | Every requirement is checkable and names its source | Designing against a wish such as "fast and accurate" |
| Design to build | The chosen platform meets residency, identity and feature needs | Building on a platform that fails security review |
| Test to release | The candidate meets or beats the pinned baseline on the eval suite | Shipping a regression the suite would have caught |
| Release to full traffic | A small share of live traffic shows no regression | A problem reaching every user at once |
| Operate to the next change | Production failures are turned into eval cases | Fixing the same failure twice |
A one-off experiment can merge stages. A regulated deployment cannot skip a gate, because the gate record is part of what the reviewer audits.
Rollout and rollback
Release a new model ID, prompt or tool set to a share of traffic first, compare its live metrics and sampled outputs with the current release, then promote it or roll it back. Rollback only works if the previous pinned version is still deployable: its model ID, prompt file and config are kept in version control and can be restored by a config change.
The failure this prevents is common. A service calls a moving alias, the alias advances, the new model formats one field differently and the parser starts raising errors. With no pinned previous version, the team cannot roll back; it can only hotfix the parser under pressure. Pinning, a retained previous version and the eval gate turn the same event into a planned upgrade.
Model versions are dependencies
Each Claude model ID identifies a fixed snapshot. Anthropic does not change the weights behind an existing ID; an updated model ships under a new ID. On newer model generations the dateless ID is itself the pinned snapshot. Some older models also had short aliases that point to the latest dated snapshot and move when a new one is released. For those, pin the dated ID in production.
Pinning does not freeze everything. Anthropic's infrastructure (request routing, safety classifiers, sampling) can be updated without changing the ID, so observable behaviour can shift slightly. Pinning plus monitoring is the pattern, not pinning alone.
Model IDs also differ by platform: Amazon Bedrock adds an anthropic. prefix, and Google Cloud uses its own format for older models. Keep the ID in configuration, never scattered through code. Where it lives in Claude Code and settings files is covered in Configuration Management.
Deprecation and retirement
Anthropic describes four states:
| State | What it means for your system |
|---|---|
| Active | Fully supported and recommended |
| Legacy | No longer updated; may be deprecated later |
| Deprecated | Still works, has a retirement date and a recommended replacement |
| Retired | Requests fail |
Anthropic gives at least 60 days' notice before retiring a publicly released model, by email and in the documentation. Amazon Bedrock and Google Cloud set their own retirement schedules. To find every place a model is still used, export usage from the Usage page in the Claude Console; the CSV breaks usage down by API key and model.
Upgrades can break requests, not only answers
A model change can change what the API accepts, so a migration needs more than a quality check:
- Newer models reject a prefilled assistant turn with a 400 error. The documented alternatives are structured outputs or system prompt instructions.
- Some newer models reject
tool_choiceofanyortoolwith a 400 error. Useautowith strict tool use, or structured outputs. - On newer models, thinking is set with adaptive thinking and an
effortlevel instead of a manualbudget_tokensvalue. - Newer models can use a different tokenizer, so the same text can count as a different number of tokens. Re-run your token counts and cost model.
Worked example: an eval gate in CI
Run this on every pull request that changes the model ID, the prompt or a tool. It fails the build if the candidate scores below the current release.
import json
import os
import sys
import anthropic
client = anthropic.Anthropic()
CANDIDATE = os.environ["CANDIDATE_MODEL"] # model ID under test
BASELINE = float(os.environ["BASELINE_PASS_RATE"]) # pass rate of the current release
SYSTEM = open("prompts/classifier_v7.txt").read() # versioned prompt file
def passes(case):
response = client.messages.create(
model=CANDIDATE,
max_tokens=200,
system=SYSTEM,
messages=[{"role": "user", "content": case["input"]}],
)
if response.stop_reason != "end_turn":
return False # cut off or refused: a fail
text = "".join(b.text for b in response.content if b.type == "text")
return text.strip() == case["expected"] # code-graded exact match
cases = [json.loads(line) for line in open("evals/classifier.jsonl")]
rate = sum(passes(c) for c in cases) / len(cases)
print(f"{CANDIDATE}: pass rate {rate:.1%} (baseline {BASELINE:.1%})")
sys.exit(0 if rate >= BASELINE else 1)
Exact match suits a classifier. For open text, swap passes for an LLM grader with a written rubric, and check a sample by hand.
Decisions the exam tests
| Situation | Choose | Why |
|---|---|---|
| Deprecation notice for your model | Run the eval suite on the named replacement, fix broken request shapes, release before the retirement date | Retired models fail every request |
| Config uses an alias on an older model | Pin the dated snapshot ID | Aliases move without a code change |
| Output drifts on a pinned ID | Compare against the eval baseline, check monitoring, report with request IDs | Infrastructure updates can shift behaviour without a new ID |
| A "small" prompt wording change | The same eval gate as a model change | Small wording changes can move results |
| A new model version passes the eval gate | Canary release with the previous pinned version still deployable | A live regression becomes a config rollback, not a hotfix |
| New model scores higher but uses more tokens | Check cost and latency requirements before switching | Most systems are judged on several criteria at once |
Rules that decide exam answers
- Pin the model ID and treat a change as a release. A model change goes through the same gate as a code change.
- No release without the eval gate. Model, prompt and tool changes all run the same suite against the same baseline.
- Deprecated still works; retired does not. Migrate between the notice and the retirement date, not after errors start.
- Upgrades can break request shapes. Prefill, forced
tool_choiceand manual thinking budgets are rejected on some newer models. Test the requests, not only the answers. - Pinning does not replace monitoring. Watch stop reasons, errors and sampled outputs even when the ID never changes.
- Keep the previous version ready. A rollback needs the old model ID, prompt and config still deployable; without them every regression becomes a hotfix.
Where it appears in the exam
Systems Life Cycle is 2.8% of the exam, inside Applications and Integration (33.1%), so expect one or two of the 53 items. The guide frames it as developing, implementing, operating and maintaining systems, so questions tend to describe a stage: a model upgrade, a deprecation notice, a prompt release or a change in production behaviour.
Two sample questions
These are original Timo practice questions. They are not official exam questions.
Build exercise
- Move the model ID and prompt version into a config file under version control and remove any hard-coded IDs.
- Build 30 eval cases from real inputs, including five edge cases, and add the gate script to CI.
- Run the gate against your current model to set the baseline, then against a candidate model or a changed prompt.
- Export usage from the Claude Console Usage page and check which model IDs each API key really calls.
Practise this topic
- Claude Certified Developer practice exam: free, 20 questions, no sign-up
- CCDV-F study guide: all topics
- Previous topic: Understanding Requirements
- Next topic: Claude API Mechanics
Sources
- Claude Certified Developer Foundations Exam Guide, version 1.0, effective July 2026 (Anthropic), Domain 2 topic: Systems Life Cycle
- Anthropic documentation: Model IDs and versioning
- Anthropic documentation: Model deprecations
- Anthropic documentation: Define success criteria and build evaluations
- Anthropic documentation: Using the Messages API
By Amotion AI