TimoBy Amotion AI

Systems life cycle: CCDV-F study guide

CCDV-F · Applications and Integration, topic weight 2.8% of the exam

Systems Life Cycle is a 2.8% topic in Applications and Integration, which makes up 33.1% of the CCDV-F exam. It tests whether you can take a Claude system through the usual stages of any IT system, and handle the parts that are specific to Claude: pinned model versions, evaluation gates before release, monitoring in production and model deprecation.

What the official guide covers

The Claude Certified Developer Foundations exam guide (version 1.0, effective July 2026) describes this topic as life cycle concepts and frameworks used to "develop, implement, operate, and maintain IT systems".

What the guide listsWhat it means in practice
DevelopBuild the prompt, tool definitions and an evaluation set together, all in version control
ImplementRelease behind an eval gate, with a pinned model ID and a staged rollout
OperateWatch stop reasons, error types, token use and sampled outputs; keep request IDs
MaintainHandle model deprecations, upgrades and prompt changes, re-running evals each time
Life cycle frameworksThe usual stages still apply; the model is a versioned dependency, like a library

The stages of a Claude system

StageWhat is specific to ClaudeWhat you produce
PlanSuccess criteria that are specific and measurable, and the real inputs, including edge cases, that will test themRequirements and success criteria
DesignPlatform, model tier and trust boundaries chosen against the requirements, such as residency and identityA design record
BuildThe system prompt, tool schemas and model ID live in config files under version controlA versioned prompt and config
TestCode-graded checks where possible, LLM grading with a clear rubric where needed, people for a sampleAn eval set and a report with a pass rate
ReleaseSame eval gate for model, prompt and tool changes; roll out to a small share of traffic firstA gated release
OperateTrack stop_reason (a rise in max_tokens or refusal is a signal), error types such as 429 and 529, and usageDashboards and alerts
Maintain and iterateDeprecation notices, migrations, prompt fixes; production failures become new eval cases and, where needed, new requirementsA new release through the same gate

Exam scenarios often ask where a piece of work belongs. A rule that data must be processed in a region is a requirement; choosing the platform that meets it is design. Writing the eval cases and rubric is test work; refusing to promote a version until it meets the pinned baseline is a release decision, as is pinning the model ID and keeping the previous one. Tracking what each production call costs in tokens and how long it takes is operate.

Log the request-id response header (in the Python SDK, response._request_id) with every call. It is the reference Anthropic support needs when you report a problem.

Gates between stages

A gate is the decision to move from one stage to the next. It is where a team, or a regulator, keeps control of a Claude system.

GatePass conditionWhat it prevents
Plan to designEvery requirement is checkable and names its sourceDesigning against a wish such as "fast and accurate"
Design to buildThe chosen platform meets residency, identity and feature needsBuilding on a platform that fails security review
Test to releaseThe candidate meets or beats the pinned baseline on the eval suiteShipping a regression the suite would have caught
Release to full trafficA small share of live traffic shows no regressionA problem reaching every user at once
Operate to the next changeProduction failures are turned into eval casesFixing the same failure twice

A one-off experiment can merge stages. A regulated deployment cannot skip a gate, because the gate record is part of what the reviewer audits.

Rollout and rollback

Release a new model ID, prompt or tool set to a share of traffic first, compare its live metrics and sampled outputs with the current release, then promote it or roll it back. Rollback only works if the previous pinned version is still deployable: its model ID, prompt file and config are kept in version control and can be restored by a config change.

The failure this prevents is common. A service calls a moving alias, the alias advances, the new model formats one field differently and the parser starts raising errors. With no pinned previous version, the team cannot roll back; it can only hotfix the parser under pressure. Pinning, a retained previous version and the eval gate turn the same event into a planned upgrade.

Model versions are dependencies

Each Claude model ID identifies a fixed snapshot. Anthropic does not change the weights behind an existing ID; an updated model ships under a new ID. On newer model generations the dateless ID is itself the pinned snapshot. Some older models also had short aliases that point to the latest dated snapshot and move when a new one is released. For those, pin the dated ID in production.

Pinning does not freeze everything. Anthropic's infrastructure (request routing, safety classifiers, sampling) can be updated without changing the ID, so observable behaviour can shift slightly. Pinning plus monitoring is the pattern, not pinning alone.

Model IDs also differ by platform: Amazon Bedrock adds an anthropic. prefix, and Google Cloud uses its own format for older models. Keep the ID in configuration, never scattered through code. Where it lives in Claude Code and settings files is covered in Configuration Management.

Deprecation and retirement

Anthropic describes four states:

StateWhat it means for your system
ActiveFully supported and recommended
LegacyNo longer updated; may be deprecated later
DeprecatedStill works, has a retirement date and a recommended replacement
RetiredRequests fail

Anthropic gives at least 60 days' notice before retiring a publicly released model, by email and in the documentation. Amazon Bedrock and Google Cloud set their own retirement schedules. To find every place a model is still used, export usage from the Usage page in the Claude Console; the CSV breaks usage down by API key and model.

Upgrades can break requests, not only answers

A model change can change what the API accepts, so a migration needs more than a quality check:

  • Newer models reject a prefilled assistant turn with a 400 error. The documented alternatives are structured outputs or system prompt instructions.
  • Some newer models reject tool_choice of any or tool with a 400 error. Use auto with strict tool use, or structured outputs.
  • On newer models, thinking is set with adaptive thinking and an effort level instead of a manual budget_tokens value.
  • Newer models can use a different tokenizer, so the same text can count as a different number of tokens. Re-run your token counts and cost model.

Worked example: an eval gate in CI

Run this on every pull request that changes the model ID, the prompt or a tool. It fails the build if the candidate scores below the current release.

import json
import os
import sys
import anthropic

client = anthropic.Anthropic()
CANDIDATE = os.environ["CANDIDATE_MODEL"]            # model ID under test
BASELINE = float(os.environ["BASELINE_PASS_RATE"])   # pass rate of the current release
SYSTEM = open("prompts/classifier_v7.txt").read()    # versioned prompt file

def passes(case):
    response = client.messages.create(
        model=CANDIDATE,
        max_tokens=200,
        system=SYSTEM,
        messages=[{"role": "user", "content": case["input"]}],
    )
    if response.stop_reason != "end_turn":
        return False                                   # cut off or refused: a fail
    text = "".join(b.text for b in response.content if b.type == "text")
    return text.strip() == case["expected"]            # code-graded exact match

cases = [json.loads(line) for line in open("evals/classifier.jsonl")]
rate = sum(passes(c) for c in cases) / len(cases)
print(f"{CANDIDATE}: pass rate {rate:.1%} (baseline {BASELINE:.1%})")
sys.exit(0 if rate >= BASELINE else 1)

Exact match suits a classifier. For open text, swap passes for an LLM grader with a written rubric, and check a sample by hand.

Decisions the exam tests

SituationChooseWhy
Deprecation notice for your modelRun the eval suite on the named replacement, fix broken request shapes, release before the retirement dateRetired models fail every request
Config uses an alias on an older modelPin the dated snapshot IDAliases move without a code change
Output drifts on a pinned IDCompare against the eval baseline, check monitoring, report with request IDsInfrastructure updates can shift behaviour without a new ID
A "small" prompt wording changeThe same eval gate as a model changeSmall wording changes can move results
A new model version passes the eval gateCanary release with the previous pinned version still deployableA live regression becomes a config rollback, not a hotfix
New model scores higher but uses more tokensCheck cost and latency requirements before switchingMost systems are judged on several criteria at once

Rules that decide exam answers

  • Pin the model ID and treat a change as a release. A model change goes through the same gate as a code change.
  • No release without the eval gate. Model, prompt and tool changes all run the same suite against the same baseline.
  • Deprecated still works; retired does not. Migrate between the notice and the retirement date, not after errors start.
  • Upgrades can break request shapes. Prefill, forced tool_choice and manual thinking budgets are rejected on some newer models. Test the requests, not only the answers.
  • Pinning does not replace monitoring. Watch stop reasons, errors and sampled outputs even when the ID never changes.
  • Keep the previous version ready. A rollback needs the old model ID, prompt and config still deployable; without them every regression becomes a hotfix.

Where it appears in the exam

Systems Life Cycle is 2.8% of the exam, inside Applications and Integration (33.1%), so expect one or two of the 53 items. The guide frames it as developing, implementing, operating and maintaining systems, so questions tend to describe a stage: a model upgrade, a deprecation notice, a prompt release or a change in production behaviour.

Two sample questions

These are original Timo practice questions. They are not official exam questions.

Question 1

A ticket classifier's config uses a short model alias for an older model generation, set when the project started. One morning the share of tickets labelled "Billing" jumps, though nobody changed the code or the prompt. What should the team change so this cannot happen silently again?

Answer: B. Aliases on older models move to new snapshots automatically, so the model changed under the team. A pinned ID plus the gate makes every change deliberate. A cannot pin a model. C reduces sampling randomness but does not stop the alias moving. D repeats the same unplanned change.

Question 2

A team gets a deprecation notice for the model behind its contract-review service, with a named replacement. In staging, every request that prefills the start of Claude's reply with { now returns a 400 error on the replacement. What is the right next step?

Answer: C. Newer models reject prefill, and structured outputs is the documented replacement; the eval suite then checks quality. A fails because deprecated models have a retirement date. B retries a request the API will always reject. D skips the gate and breaks the service.

Build exercise

  1. Move the model ID and prompt version into a config file under version control and remove any hard-coded IDs.
  2. Build 30 eval cases from real inputs, including five edge cases, and add the gate script to CI.
  3. Run the gate against your current model to set the baseline, then against a candidate model or a changed prompt.
  4. Export usage from the Claude Console Usage page and check which model IDs each API key really calls.

Practise this topic

Sources