TimoBy Amotion AI

Technical fundamentals: CCDV-F study guide

CCDV-F · Model Selection and Optimization, topic weight 6.1% of the exam

Technical Fundamentals is a 6.1% topic in Model Selection and Optimization, which makes up 16.8% of the CCDV-F exam. It tests the plumbing under every Claude application: what the client SDKs do on top of the REST API, how to handle each error type, and how streaming reaches a user's screen.

What the official guide covers

The Claude Certified Developer Foundations exam guide (version 1.0, effective July 2026) describes this topic as foundational technical concepts, including basic engineering practices such as "integrating with SDKs that wrap REST APIs, websockets".

What the guide listsWhat it means in practice
SDKs that wrap REST APIsWhat the Claude SDK adds to raw HTTP: headers, retries, timeouts, typed errors, request IDs, streaming helpers
WebsocketsHow the Claude API's server-sent events differ from WebSockets, and where each belongs in your app
Basic engineering practicesRetrying only what can succeed, setting timeouts, logging request IDs, keeping keys on the server

The REST call under the SDK

The Claude API is plain HTTPS and JSON. This is a complete request:

curl https://api.anthropic.com/v1/messages \
  -H "x-api-key: $ANTHROPIC_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "content-type: application/json" \
  -d '{"model": "'"$CLAUDE_MODEL"'", "max_tokens": 256,
       "messages": [{"role": "user", "content": "Hello, Claude"}]}'

Every response carries a request-id header. Errors come back as JSON with a type of "error", an error object holding its own type and message, and the request_id.

Official client SDKs exist for Python, TypeScript, Java, Go, C#, Ruby and PHP. They send the same requests; the difference is what you no longer write yourself.

What the SDK adds

ConcernWith raw HTTPWith the Python SDK (pip install anthropic)
Auth and version headersYou set themReads ANTHROPIC_API_KEY and sets the headers
RetriesYou write backoffRetries connection errors, 408, 409, 429 and 5xx errors twice by default; change with max_retries
TimeoutsYou choose one10 minutes by default; set timeout on the client or per request with client.with_options(...)
ErrorsParse the bodyTyped exceptions such as BadRequestError, AuthenticationError, RateLimitError, InternalServerError and APIConnectionError
Request IDRead the headerresponse._request_id
StreamingParse server-sent events yourselfclient.messages.stream(), text_stream, get_final_message()
AsyncYour own HTTP clientAsyncAnthropic with the same methods; in TypeScript the single client is Promise-based, so you await it

Error types and what to do with each

StatusError typeRetry?What to do
400invalid_request_errorNoFix the request; the same request fails again
401authentication_errorNoFix or rotate the key
403permission_errorNoThe key lacks access to that resource
404not_found_errorNoCheck the model ID or resource ID
413request_too_largeNoShrink or split the request
429rate_limit_errorYes, unless it is a spend capWait for the retry-after header; smooth bursts and ramp traffic up gradually. A spend-cap 429 has no retry-after and keeps failing until access resumes
500api_errorYesRetry with exponential backoff
504timeout_errorYesRetry; if it repeats on the same large request, stream it or make it smaller
529overloaded_errorYesRetry with backoff

The API also enforces acceleration limits when an organisation's usage rises sharply, which is another reason to ramp traffic up rather than switch it on all at once.

Retriable or terminal: one question

For any failure, ask: would the identical request plausibly succeed if sent again later? If yes, it is retriable (rate limits, overload, timeouts, dropped connections). If no, it is terminal (a malformed body, a bad key, a missing model), and every retry wastes time and budget while hiding the real fault. When you cannot tell, treat it as terminal: a wrong terminal call fails loudly and gets fixed, while a wrong retriable call hammers the service.

Two more traps sit next to this decision:

  • One retry layer, not two. The SDK already retries connection errors, 408, 409, 429 and 5xx responses. A hand-written loop around the same call multiplies attempts against a rate limit. Either keep the SDK's retries and add only application fallbacks, or set max_retries=0 and own the whole path.
  • A refusal is not an error. It arrives as HTTP 200 with stop_reason: "refusal", so no retry logic sees it. Check stop_reason, log the refusal and return it to the caller rather than retrying or treating it as valid output.

Worked example: owning the retry path

A common broken helper catches every exception and retries at once with no wait. It retries 400s that can never succeed and hammers a rate limit it should wait out. When you need your own policy, turn the SDK's retries off so there is only one layer, then classify, wait and give up:

import random
import time
import anthropic

client = anthropic.Anthropic(max_retries=0)       # this function is the only retry layer
RETRIABLE = {408, 409, 429, 500, 502, 503, 504, 529}

def create_with_retry(max_attempts=5, cap=30, **params):
    for attempt in range(1, max_attempts + 1):
        try:
            return client.messages.create(**params)
        except anthropic.APIStatusError as e:
            if e.status_code not in RETRIABLE or attempt == max_attempts:
                raise                              # terminal, or out of attempts
            wait = e.response.headers.get("retry-after")
        except anthropic.APIConnectionError:       # dropped connection or timeout
            if attempt == max_attempts:
                raise
            wait = None
        delay = float(wait) if wait else min(cap, 2 ** attempt) + random.uniform(0, 1)
        time.sleep(delay)                          # server's wait first, then capped backoff with jitter

The attempt cap matters as much as the wait: a spend-cap 429 has no retry-after and will not clear within a job, so the loop must end and report it.

Server-sent events and WebSockets

The Claude API streams with server-sent events (SSE). Your client makes one HTTP request with stream: true, and the server pushes events one way down that response: message_start, content block events, message_delta, message_stop, with ping events in between. There is no WebSocket connection to the Claude API.

WebSockets are two-way and long-lived. They belong between your users and your server, when the browser needs to send as well as receive: a chat that can be interrupted, or an agent the user steers. Anthropic's Agent SDK hosting guide describes the same split: the container exposes an HTTP or WebSocket endpoint, and the agent subprocess itself does not listen on the network.

The usual shape is: browser to your server over a WebSocket, your server to Claude over an SSE stream. The API key stays on the server.

Worked example: relay a Claude stream over a WebSocket

import os
import anthropic
from anthropic import AsyncAnthropic
from fastapi import FastAPI, WebSocket

app = FastAPI()
client = AsyncAnthropic()                 # the API key never reaches the browser
MODEL = os.environ["CLAUDE_MODEL"]

@app.websocket("/chat")
async def chat(ws: WebSocket):
    await ws.accept()
    history = []
    while True:
        history.append({"role": "user", "content": await ws.receive_text()})
        try:
            async with client.messages.stream(model=MODEL, max_tokens=1024, messages=history) as stream:
                async for text in stream.text_stream:          # SSE from Claude
                    await ws.send_json({"type": "delta", "text": text})
                final = await stream.get_final_message()
        except anthropic.RateLimitError:                       # raised after the SDK's own retries
            history.pop()
            await ws.send_json({"type": "error", "message": "Busy. Please try again shortly."})
            continue
        except anthropic.APIError:                             # includes errors sent mid-stream
            history.pop()
            await ws.send_json({"type": "error", "message": "The request failed."})
            continue
        history.append({"role": "assistant", "content": final.content})
        await ws.send_json({"type": "done", "stop_reason": final.stop_reason})

Two details matter. An error can arrive as an SSE error event after the HTTP status was already 200, so the relay must report failures inside the stream, not only before it. And stop_reason tells the browser whether the reply finished (end_turn) or was cut off (max_tokens).

Decisions the exam tests

SituationChooseWhy
Browser chat with live text and a stop buttonWebSocket to your server, SSE from your server to ClaudeTwo-way for the user; the key stays server-side
Server-to-server call with a short replyA plain SDK callSimplest; nothing to relay
A reply with a very large max_tokensStreamingThe SDK requires it to avoid HTTP timeouts
Requests that may run over 10 minutesStreaming or the Message Batches APIIdle connections can be dropped on long non-streaming calls
Bursty traffic hitting 429SDK retries plus your own queue or concurrency capRetries absorb spikes; the cap prevents them
An unexplained bad response in productionLook up the logged request IDIt is what Anthropic support asks for

Rules that decide exam answers

  • Retry only what can succeed, in one layer. 429, 500, 504, 529 and connection errors are retryable. 400, 401, 403, 404 and 413 fail again until you change something. Do not wrap your own retry loop around the SDK's.
  • Honour retry-after. On a 429, wait as long as the header says. Retrying at once makes it worse.
  • Claude streams over SSE. WebSockets sit between your users and your server, not between your server and Claude.
  • Keep the API key on the server. A browser never calls the Claude API directly with your key.
  • A 200 can still fail. In a stream, handle error events and read stop_reason at the end.
  • Log the request ID. Every response has one, and support needs it.

Where it appears in the exam

Technical Fundamentals is 6.1% of the exam, inside Model Selection and Optimization (16.8%), so expect about three of the 53 items. The guide names SDKs that wrap REST APIs and WebSockets, so questions describe an integration that fails, retries badly, times out or streams to the wrong place, and ask for the sound engineering fix.

Two sample questions

These are original Timo practice questions. They are not official exam questions.

Question 1

A team is building a browser chat for Claude and wants text to appear as it is generated. One developer proposes connecting the browser straight to the Claude API so tokens stream into the page. What is the better design?

Answer: D. The server holds the key, receives Claude's SSE stream and relays it over a two-way channel the browser can also use to stop or reply. A exposes the key to every visitor. B gives up the live text the team wants. C relies on a polling interface the Messages API does not offer.

Question 2

A nightly job calls Claude through the Python SDK. Some calls fail with BadRequestError and others with RateLimitError. The developer wraps every call in a loop that retries any exception ten times at once. What is wrong with this?

Answer: B. Bad requests need a fix, not a retry, and immediate retries on top of the SDK's own backoff add load during a rate limit. A wastes calls on errors that cannot succeed. C is wrong because 429s clear once you slow down. D misreads a 400, which means the request itself is invalid.

Build exercise

  1. Send the curl request above, then the same request with the Python SDK. Compare the headers and print response._request_id.
  2. Send a request with no max_tokens, catch the BadRequestError and print its status code and message.
  3. Add "stream": true to the curl body and match each event in the raw output to the list above.
  4. Run the WebSocket relay locally with uvicorn, open the chat from two browser tabs and confirm each tab streams its own reply.

Practise this topic

Sources