TimoBy Amotion AI

Debugging and Error Handling: CCDV-F study guide

CCDV-F · Eval, Testing, and Debugging, topic weight 2.6% of the exam

Debugging and Error Handling is the only topic in Domain 4, Eval, Testing, and Debugging, so its 2.6% is the whole domain. It tests one skill: when a Claude application fails, name the kind of failure, pick the recovery, and find out whether the fault is in your integration code or in the model's output.

What the official guide covers

The Claude Certified Developer Foundations exam guide (version 1.0, effective July 2026) describes this topic as debugging and error handling techniques for Claude applications:

What the guide listsWhat it means in practice
Error type identificationTell an HTTP error from an early stop, a failed tool and a wrong answer
Recovery strategy selectionRetry what is temporary; fix what is your fault
Trace analysis to identify failure modesLog each request, tool call and stop reason so you can replay a failure
Problem origin isolation between the integration layer and model outputDecide whether the bug is in your code, tools, prompt or the model

Four kinds of failure

Sort the symptom first; most debugging time is lost treating one kind of failure as another:

KindHow it showsExampleFirst move
Request failedAn HTTP error status; the SDK raises an exception429 rate_limit_error, 529 overloaded_errorRead the error type; retry only if it is temporary
Response stopped earlyHTTP 200, but stop_reason is not end_turnmax_tokens cut the JSON in halfHandle each stop_reason explicitly
Tool failedA tool_result with is_error: true, or a tool that returned wrong dataOrder lookup hit the wrong databaseFix the tool; give Claude a clear error message
Output is wrongValid format, wrong contentA plausible but wrong summaryCheck what Claude actually received, then the prompt

Stop reasons are not errors: a max_tokens stop is a successful request with cut-off output, so try/except never sees it.

API error types and the right recovery

StatusError typeMeaningRecovery
400invalid_request_errorThe request format or content is wrongFix the payload; do not retry unchanged
401authentication_errorKey missing, malformed, revoked or expiredFix credentials
403permission_errorThe key cannot use this resourceFix access, not the code
404not_found_errorResource not found, such as a wrong model IDCheck names and IDs
413request_too_largeThe request is over the size limitSend less: split documents, use files
429rate_limit_errorA rate or usage limit was hitBack off and retry, honouring retry-after
500api_errorUnexpected error on Anthropic's sideRetry with backoff
529overloaded_errorThe API is temporarily overloadedRetry with backoff; ramp traffic up gradually

Every response carries a request-id header; log it.

The Python SDK already retries connection errors, 408, 409, 429 and 5xx responses twice by default with exponential backoff (max_retries changes it), and times out after 10 minutes by default with APITimeoutError. Each status has an exception class: BadRequestError (400), AuthenticationError (401), PermissionDeniedError (403), NotFoundError (404), RateLimitError (429), InternalServerError (5xx) and APIConnectionError for network failures. ANTHROPIC_LOG=debug shows the SDK's own logging.

In streaming, an error can arrive as an error event after the HTTP 200, such as overloaded_error. Partial text can be kept and continued with a follow-up request that asks Claude to carry on from where it stopped, but tool_use and thinking blocks cannot be partly recovered: discard an unfinished one and resume from the last complete text block, or retry the request.

A call wrapper that sorts failures (Python)

import logging
import os
import anthropic

log = logging.getLogger("claude")
client = anthropic.Anthropic(max_retries=3, timeout=60.0)   # SDK retries 429, 5xx and connection errors
MODEL = os.environ["CLAUDE_MODEL"]

class TruncatedOutput(Exception):
    pass

def ask(messages, max_tokens=2048):
    try:
        response = client.messages.create(model=MODEL, max_tokens=max_tokens, messages=messages)
    except anthropic.BadRequestError as e:              # 400: fix the payload, do not retry
        log.error("invalid request: %s", e)
        raise
    except (anthropic.AuthenticationError, anthropic.PermissionDeniedError) as e:
        log.error("credential or access problem: %s", e)
        raise
    except anthropic.RateLimitError as e:               # still 429 after SDK retries
        log.warning("rate limited, queue for later: %s", e)
        raise
    except anthropic.APIConnectionError as e:           # network failure or timeout
        log.warning("network problem: %s", e)
        raise
    except anthropic.APIStatusError as e:               # any other status, including 5xx
        log.error("status=%s request_id=%s", e.status_code, e.response.headers.get("request-id"))
        raise

    log.info("request_id=%s stop_reason=%s input_tokens=%s output_tokens=%s",
             response._request_id, response.stop_reason,
             response.usage.input_tokens, response.usage.output_tokens)

    if response.stop_reason == "max_tokens":           # HTTP 200, output cut off
        raise TruncatedOutput("Raise max_tokens or split the input")
    if response.stop_reason == "refusal":
        raise RuntimeError("Claude declined; inspect stop_details")
    return response

Specific exception classes come first, because RateLimitError is a subclass of APIStatusError.

Trace analysis

A trace is the full record of one run: every request, response, stop_reason, tool input and tool result, with request IDs and token counts. It shows the exact step where a run went wrong. Typical findings:

  • A tool returned an error, but the error message was too vague for Claude to correct the input.
  • A max_tokens stop cut off a tool_use block, so the tool never ran. The fix is a retry with a higher max_tokens.
  • The loop ended early because code treated text as a stop signal instead of reading stop_reason.
  • A key fact was cleared or compacted out of context before Claude needed it.
  • Tool choice was right for several turns and wrong from one turn onward. Large tool results are piling up in the window; the tool description is fine, since it worked earlier. Prune old results or compact. Raising max_tokens will not help: it limits what Claude writes, not what it reads.

Agents on the Claude Agent SDK or claude -p give you much of the trace for free. The ResultMessage subtype says how the run ended (success, error_max_turns, error_during_execution and others) with num_turns and usage. In stream-json output, a system/api_retry event reports each retried API failure. CLAUDE_CODE_ENABLE_TELEMETRY=1 with the OpenTelemetry exporter variables sends traces, metrics and logs to your collector; prompt text and tool inputs are excluded unless you opt in.

Errors that surface one request late

Some faults pass silently on the request that causes them and fail the next one with a 400 invalid_request_error. Teams then debug the wrong request.

Next request fails becauseWhere the fault really isFix
A tool_result carries a tool_use_id that matches no tool_use in the turn beforeThe code that builds tool resultsCopy each block's id exactly, answer every call in one user message straight after, and put tool_result blocks before any text
History holds a tool_use block with truncated JSON inputA stream dropped mid-block and the handler saved the turn anywayAppend a streamed turn only after message_stop; on interruption, drop the partial turn and retry
A thinking block in a tool-use turn was edited, reordered or droppedCode that filters or rewrites historyPass thinking blocks back exactly as received within a tool-use turn

The second row is the classic trap: a stream that ends is not a message that completed. Gate history writes on the end of the message, not on the end of your read loop.

Match the test to the failure

A trace shows where a run failed; tests decide which failures you hear about before users do. Each level sees a different break:

LevelChecksCannot see
UnitA single function, for example a parser or a tool wrapperHow pieces fit together
FunctionalA single Claude call gives back the expected shape for an inputThe system around that call
IntegrationThe hand-off between two components, such as retrieval into the prompt builderBehaviour that only appears end to end
End-to-endThe whole flow from user input to outputWhich step broke

Most silent failures live at the integration seam. A typical case: retrieval returns a list of chunk objects, the prompt builder expects a string, the context arrives malformed, and Claude answers from general knowledge. The unit and functional tests both pass because each side works alone. Only a test that drives real retrieval output into the prompt builder catches it, and the fix is to agree the data format at that boundary, not to reword the prompt.

Isolate the origin: integration layer or model output

Work from the outside in. Each step either finds the fault or clears one layer:

CheckIf it fails, the problem is inFix
Did the request leave your code as intended? Log the payloadIntegration layerFix message building, tool schemas, model ID
Did the API return an error status?Integration layer or serviceRecover by error type (table above)
Did stop_reason show a cut-off or refusal?Request settingsRaise max_tokens, split input, handle refusal
Did each tool receive the right input and return correct data?Tool implementationFix the tool or its error messages
Was every fact the answer needed actually in the context?Context assemblyFix retrieval, pruning or the prompt
All of the above are correct, and the answer is still wrongModel outputImprove instructions or examples, add validation, or change model tier

Replay the logged request to see whether the failure repeats, and test a fix on several runs, since outputs vary.

Rules that decide exam answers

  • Classify before you retry. Retry 429, 500 and 529; never retry a 400 or 401 unchanged.
  • Stop reasons are not exceptions. Check stop_reason on every successful response; max_tokens means truncated output.
  • Do not stack retries blindly. The SDK already retries transient errors; your own loop multiplies the attempts. If you own the retry, wait longer each time with jitter and a cap, honour retry-after, and raise at once on a terminal error such as 400 or 401.
  • Prove the integration layer first. Log request IDs, stop reasons, the exact request and tool results before blaming the model.
  • A 400 on a retry often started one request earlier. Check tool result IDs and any turn assembled from an interrupted stream before touching the schema.
  • Give tools useful error messages. Claude recovers from an is_error result only if it says what went wrong.

Where it appears in the exam

Domain 4, Eval, Testing, and Debugging, is 2.6% of the exam and is this one topic, so expect about one item out of 53. Questions describe a symptom with log or trace details and ask for the cause or the right recovery.

Two sample questions

These are original Timo practice questions. They are not official exam questions.

Question 1

An extraction service asks Claude for a JSON object from each contract. For long contracts, about one in five responses fails to parse. No exceptions are raised, and the logs show stop_reason: "max_tokens" on every failed response. What should the developer do?

Answer: D. The responses were cut off at the token limit, so the output needs more room or the input needs splitting. A repeats the same truncated request. B cannot help when generation stops at the limit. C confuses a stop reason on a successful response with an HTTP error.

Question 2

A support agent tells customers that valid orders do not exist. The trace shows Claude calling get_order with the correct order ID, and the tool returning "not found". The same ID is in the production database. Where is the fault, and what is the fix?

Answer: A. Claude passed the right ID and reported what the tool returned, so the tool or its configuration is wrong. B changes the prompt when the model behaved correctly. C does not match the trace, which shows a complete tool call. D addresses context loss, but the ID reached the tool intact.

Build exercise

  1. Use the ask() wrapper above. Force a 401 with a bad key and a 400 with an empty message list, and check each branch.
  2. Set max_tokens to 50 on a long answer and confirm the TruncatedOutput path runs.
  3. Give a tool a vague error ("failed"), then a specific one ("IDs look like A100"), and compare Claude's recovery.
  4. Run claude -p --output-format stream-json --verbose and read the system/init and final result events.

Practise this topic

Sources

  • Claude Certified Developer Foundations Exam Guide, version 1.0, effective July 2026 (Anthropic), Domain 4 topic: Debugging and Error Handling
  • Claude Platform documentation: Claude API errors
  • Claude Platform documentation: Python SDK
  • Claude Platform documentation: Stop reasons
  • Claude Agent SDK documentation: Hosting the Agent SDK