Debugging and Error Handling: CCDV-F study guide
CCDV-F · Eval, Testing, and Debugging, topic weight 2.6% of the exam
Debugging and Error Handling is the only topic in Domain 4, Eval, Testing, and Debugging, so its 2.6% is the whole domain. It tests one skill: when a Claude application fails, name the kind of failure, pick the recovery, and find out whether the fault is in your integration code or in the model's output.
What the official guide covers
The Claude Certified Developer Foundations exam guide (version 1.0, effective July 2026) describes this topic as debugging and error handling techniques for Claude applications:
| What the guide lists | What it means in practice |
|---|---|
| Error type identification | Tell an HTTP error from an early stop, a failed tool and a wrong answer |
| Recovery strategy selection | Retry what is temporary; fix what is your fault |
| Trace analysis to identify failure modes | Log each request, tool call and stop reason so you can replay a failure |
| Problem origin isolation between the integration layer and model output | Decide whether the bug is in your code, tools, prompt or the model |
Four kinds of failure
Sort the symptom first; most debugging time is lost treating one kind of failure as another:
| Kind | How it shows | Example | First move |
|---|---|---|---|
| Request failed | An HTTP error status; the SDK raises an exception | 429 rate_limit_error, 529 overloaded_error | Read the error type; retry only if it is temporary |
| Response stopped early | HTTP 200, but stop_reason is not end_turn | max_tokens cut the JSON in half | Handle each stop_reason explicitly |
| Tool failed | A tool_result with is_error: true, or a tool that returned wrong data | Order lookup hit the wrong database | Fix the tool; give Claude a clear error message |
| Output is wrong | Valid format, wrong content | A plausible but wrong summary | Check what Claude actually received, then the prompt |
Stop reasons are not errors: a max_tokens stop is a successful request with cut-off output, so try/except never sees it.
API error types and the right recovery
| Status | Error type | Meaning | Recovery |
|---|---|---|---|
| 400 | invalid_request_error | The request format or content is wrong | Fix the payload; do not retry unchanged |
| 401 | authentication_error | Key missing, malformed, revoked or expired | Fix credentials |
| 403 | permission_error | The key cannot use this resource | Fix access, not the code |
| 404 | not_found_error | Resource not found, such as a wrong model ID | Check names and IDs |
| 413 | request_too_large | The request is over the size limit | Send less: split documents, use files |
| 429 | rate_limit_error | A rate or usage limit was hit | Back off and retry, honouring retry-after |
| 500 | api_error | Unexpected error on Anthropic's side | Retry with backoff |
| 529 | overloaded_error | The API is temporarily overloaded | Retry with backoff; ramp traffic up gradually |
Every response carries a request-id header; log it.
The Python SDK already retries connection errors, 408, 409, 429 and 5xx responses twice by default with exponential backoff (max_retries changes it), and times out after 10 minutes by default with APITimeoutError. Each status has an exception class: BadRequestError (400), AuthenticationError (401), PermissionDeniedError (403), NotFoundError (404), RateLimitError (429), InternalServerError (5xx) and APIConnectionError for network failures. ANTHROPIC_LOG=debug shows the SDK's own logging.
In streaming, an error can arrive as an error event after the HTTP 200, such as overloaded_error. Partial text can be kept and continued with a follow-up request that asks Claude to carry on from where it stopped, but tool_use and thinking blocks cannot be partly recovered: discard an unfinished one and resume from the last complete text block, or retry the request.
A call wrapper that sorts failures (Python)
import logging
import os
import anthropic
log = logging.getLogger("claude")
client = anthropic.Anthropic(max_retries=3, timeout=60.0) # SDK retries 429, 5xx and connection errors
MODEL = os.environ["CLAUDE_MODEL"]
class TruncatedOutput(Exception):
pass
def ask(messages, max_tokens=2048):
try:
response = client.messages.create(model=MODEL, max_tokens=max_tokens, messages=messages)
except anthropic.BadRequestError as e: # 400: fix the payload, do not retry
log.error("invalid request: %s", e)
raise
except (anthropic.AuthenticationError, anthropic.PermissionDeniedError) as e:
log.error("credential or access problem: %s", e)
raise
except anthropic.RateLimitError as e: # still 429 after SDK retries
log.warning("rate limited, queue for later: %s", e)
raise
except anthropic.APIConnectionError as e: # network failure or timeout
log.warning("network problem: %s", e)
raise
except anthropic.APIStatusError as e: # any other status, including 5xx
log.error("status=%s request_id=%s", e.status_code, e.response.headers.get("request-id"))
raise
log.info("request_id=%s stop_reason=%s input_tokens=%s output_tokens=%s",
response._request_id, response.stop_reason,
response.usage.input_tokens, response.usage.output_tokens)
if response.stop_reason == "max_tokens": # HTTP 200, output cut off
raise TruncatedOutput("Raise max_tokens or split the input")
if response.stop_reason == "refusal":
raise RuntimeError("Claude declined; inspect stop_details")
return response
Specific exception classes come first, because RateLimitError is a subclass of APIStatusError.
Trace analysis
A trace is the full record of one run: every request, response, stop_reason, tool input and tool result, with request IDs and token counts. It shows the exact step where a run went wrong. Typical findings:
- A tool returned an error, but the error message was too vague for Claude to correct the input.
- A
max_tokensstop cut off atool_useblock, so the tool never ran. The fix is a retry with a highermax_tokens. - The loop ended early because code treated text as a stop signal instead of reading
stop_reason. - A key fact was cleared or compacted out of context before Claude needed it.
- Tool choice was right for several turns and wrong from one turn onward. Large tool results are piling up in the window; the tool description is fine, since it worked earlier. Prune old results or compact. Raising
max_tokenswill not help: it limits what Claude writes, not what it reads.
Agents on the Claude Agent SDK or claude -p give you much of the trace for free. The ResultMessage subtype says how the run ended (success, error_max_turns, error_during_execution and others) with num_turns and usage. In stream-json output, a system/api_retry event reports each retried API failure. CLAUDE_CODE_ENABLE_TELEMETRY=1 with the OpenTelemetry exporter variables sends traces, metrics and logs to your collector; prompt text and tool inputs are excluded unless you opt in.
Errors that surface one request late
Some faults pass silently on the request that causes them and fail the next one with a 400 invalid_request_error. Teams then debug the wrong request.
| Next request fails because | Where the fault really is | Fix |
|---|---|---|
A tool_result carries a tool_use_id that matches no tool_use in the turn before | The code that builds tool results | Copy each block's id exactly, answer every call in one user message straight after, and put tool_result blocks before any text |
History holds a tool_use block with truncated JSON input | A stream dropped mid-block and the handler saved the turn anyway | Append a streamed turn only after message_stop; on interruption, drop the partial turn and retry |
| A thinking block in a tool-use turn was edited, reordered or dropped | Code that filters or rewrites history | Pass thinking blocks back exactly as received within a tool-use turn |
The second row is the classic trap: a stream that ends is not a message that completed. Gate history writes on the end of the message, not on the end of your read loop.
Match the test to the failure
A trace shows where a run failed; tests decide which failures you hear about before users do. Each level sees a different break:
| Level | Checks | Cannot see |
|---|---|---|
| Unit | A single function, for example a parser or a tool wrapper | How pieces fit together |
| Functional | A single Claude call gives back the expected shape for an input | The system around that call |
| Integration | The hand-off between two components, such as retrieval into the prompt builder | Behaviour that only appears end to end |
| End-to-end | The whole flow from user input to output | Which step broke |
Most silent failures live at the integration seam. A typical case: retrieval returns a list of chunk objects, the prompt builder expects a string, the context arrives malformed, and Claude answers from general knowledge. The unit and functional tests both pass because each side works alone. Only a test that drives real retrieval output into the prompt builder catches it, and the fix is to agree the data format at that boundary, not to reword the prompt.
Isolate the origin: integration layer or model output
Work from the outside in. Each step either finds the fault or clears one layer:
| Check | If it fails, the problem is in | Fix |
|---|---|---|
| Did the request leave your code as intended? Log the payload | Integration layer | Fix message building, tool schemas, model ID |
| Did the API return an error status? | Integration layer or service | Recover by error type (table above) |
Did stop_reason show a cut-off or refusal? | Request settings | Raise max_tokens, split input, handle refusal |
| Did each tool receive the right input and return correct data? | Tool implementation | Fix the tool or its error messages |
| Was every fact the answer needed actually in the context? | Context assembly | Fix retrieval, pruning or the prompt |
| All of the above are correct, and the answer is still wrong | Model output | Improve instructions or examples, add validation, or change model tier |
Replay the logged request to see whether the failure repeats, and test a fix on several runs, since outputs vary.
Rules that decide exam answers
- Classify before you retry. Retry 429, 500 and 529; never retry a 400 or 401 unchanged.
- Stop reasons are not exceptions. Check
stop_reasonon every successful response;max_tokensmeans truncated output. - Do not stack retries blindly. The SDK already retries transient errors; your own loop multiplies the attempts. If you own the retry, wait longer each time with jitter and a cap, honour
retry-after, and raise at once on a terminal error such as 400 or 401. - Prove the integration layer first. Log request IDs, stop reasons, the exact request and tool results before blaming the model.
- A 400 on a retry often started one request earlier. Check tool result IDs and any turn assembled from an interrupted stream before touching the schema.
- Give tools useful error messages. Claude recovers from an
is_errorresult only if it says what went wrong.
Where it appears in the exam
Domain 4, Eval, Testing, and Debugging, is 2.6% of the exam and is this one topic, so expect about one item out of 53. Questions describe a symptom with log or trace details and ask for the cause or the right recovery.
Two sample questions
These are original Timo practice questions. They are not official exam questions.
Build exercise
- Use the
ask()wrapper above. Force a 401 with a bad key and a 400 with an empty message list, and check each branch. - Set
max_tokensto 50 on a long answer and confirm theTruncatedOutputpath runs. - Give a tool a vague error ("failed"), then a specific one ("IDs look like A100"), and compare Claude's recovery.
- Run
claude -p --output-format stream-json --verboseand read thesystem/initand finalresultevents.
Practise this topic
- Claude Certified Developer practice exam: free, 20 questions, no sign-up
- CCDV-F study guide: all topics
- Worked example: Production Claude tool loop
- Previous topic: Claude Code Operation
- Next topic: LLM Fundamentals
Sources
- Claude Certified Developer Foundations Exam Guide, version 1.0, effective July 2026 (Anthropic), Domain 4 topic: Debugging and Error Handling
- Claude Platform documentation: Claude API errors
- Claude Platform documentation: Python SDK
- Claude Platform documentation: Stop reasons
- Claude Agent SDK documentation: Hosting the Agent SDK
By Amotion AI