TimoBy Amotion AI

Validation and retry loops: CCAR-F task statement 4.4

CCAR-F · Prompt Engineering & Structured Output (20% of the exam)

Task statement 4.4 sits in Prompt Engineering & Structured Output, 20% of the CCAR-F exam. It tests what happens after extraction: checking the data in code, sending the specific errors back to Claude for a corrected attempt, and recognising the cases where a retry cannot help.

What the official guide covers

The Claude Certified Architect Foundations exam guide (version 1.0, effective July 2026) lists this under task statement 4.4, "Implement validation, retry, and feedback loops for extraction quality":

Knowledge ofSkills in
Retry with error feedback: adding the specific validation errors to the retry prompt so Claude can correct themSending follow-up requests that contain the original document, the failed extraction and the specific validation errors
Retries do not help when the information is not in the source document, unlike format or structure errorsTelling when a retry will fail (the data is only in an external document you did not provide) and when it will work (format mismatches, structural errors)
Feedback loops: a detected_pattern field records which code construct triggered a finding, so dismissals can be analysedAdding detected_pattern to structured findings to study the false positives developers dismiss
Semantic validation errors (values that do not add up, values in the wrong field) differ from schema syntax errors, which tool use removesDesigning self-checks: extracting calculated_total next to stated_total, and adding a conflict_detected flag for inconsistent source data

How the loop works

  1. Extract with a schema. Use strict tool use or JSON outputs (task statement 4.3). Syntax errors and wrong types are gone before your code sees the data.
  2. Validate the meaning in code. Check what a schema cannot: totals, date order, cross-field rules, values in the right field.
  3. If a check fails, send one follow-up request with the original document, the extraction that failed and each error written out with its numbers, such as "line items add up to 1,180.00 but calculated_total is 1,108.00".
  4. Validate again. Stop after one or two retries. Anything still failing goes to a person with the errors attached.

"Try again" alone gives Claude nothing to fix. The specific error tells it where to look in the document.

Keep "Use null for values the document does not contain" in the retry prompt too. Anthropic's guide to reducing hallucinations recommends giving Claude explicit permission to say it does not know. Without it, a retry that demands a value can turn a missing field into an invented one.

Will a retry help?

ErrorRetry helps?Why
Date written as "03/04" instead of YYYY-MM-DDYesFormat problem; the value is in the document
Tax amount placed in the discount fieldYesStructural problem; the error message says which field
calculated_total does not match the line items Claude extractedUsuallyClaude missed or misread a line it can find again
The invoice's own printed total disagrees with its linesNoThe source conflicts; set conflict_detected and send it to a person
A contract says "terms as per the master agreement", which you did not provideNoThe data is not in the input; fetch that document or escalate
A field is null because the document does not contain itNonull is the correct answer
Reply cut off with stop_reason "max_tokens", so it does not match the schemaYes, with a higher max_tokensNot a content error; error feedback adds nothing
stop_reason is "refusal"NoA content decision, returned with HTTP 200; log it and route it to a person

Check stop_reason before you validate. A truncated or refused reply fails validation for reasons that have nothing to do with the extraction.

Which retries your code owns

Three different retry mechanisms show up in extraction systems, and exam options often mix them up. First ask whether the failure is retriable at all, then who does the retry.

FailureWho retriesHow
Rate limit, timeout, connection or server errorThe Python SDK, automaticallySame request again with backoff; 2 retries by default, changed with max_retries
A tool the agent called raised an errorClaude, on its next turnReturn a tool_result with is_error set and a message that says what went wrong (task 2.2)
Output matches the schema but fails a business checkYour codeA new request with the document, the failed output and the errors
The value is not in the sourceNobodyFetch the other document or send the record to a person

Raising the SDK's retry count does nothing for a wrong total, and a validation retry does nothing for a value that is not there.

Self-check fields: calculated_total and conflict_detected

Ask for two totals. stated_total is the number printed on the document. calculated_total is the sum of the line items Claude extracted. Your code then checks two things: does calculated_total really equal the sum of the line items, and does it equal stated_total?

If Claude's own sum is wrong, a retry with that error usually fixes it. If the sum is right but the printed total differs, the document itself is inconsistent. A conflict_detected flag lets Claude say so directly, and your code routes that record to review instead of retrying. Retrying a document that really is inconsistent only invites Claude to "fix" one of the numbers.

Retry with the document, the failed output and the errors (Python)

import json
import re
import anthropic

client = anthropic.Anthropic()
MODEL = "your-model-id"

SCHEMA = {
    "type": "object",
    "properties": {
        "invoice_number": {"type": "string"},
        "invoice_date": {"type": ["string", "null"]},
        "line_items": {"type": "array", "items": {
            "type": "object",
            "properties": {"description": {"type": "string"}, "amount": {"type": "number"}},
            "required": ["description", "amount"],
            "additionalProperties": False}},
        "stated_total": {"type": ["number", "null"]},
        "calculated_total": {"type": "number"},
        "conflict_detected": {"type": "boolean"},
    },
    "required": ["invoice_number", "invoice_date", "line_items",
                 "stated_total", "calculated_total", "conflict_detected"],
    "additionalProperties": False,
}

def extract(document, failed=None, errors=None):
    prompt = (f"<document>\n{document}\n</document>\n"
              "Extract the invoice. Dates as YYYY-MM-DD. Use null for values "
              "the document does not contain.")
    if errors:
        prompt += (f"\n<failed_extraction>\n{json.dumps(failed)}\n</failed_extraction>\n"
                   "<validation_errors>\n" + "\n".join(errors) + "\n</validation_errors>\n"
                   "Correct these errors using the document.")
    response = client.messages.create(
        model=MODEL,
        max_tokens=2048,
        output_config={"format": {"type": "json_schema", "schema": SCHEMA}},
        messages=[{"role": "user", "content": prompt}],
    )
    return json.loads(next(b.text for b in response.content if b.type == "text"))

def validate(data):
    errors = []
    items_sum = round(sum(item["amount"] for item in data["line_items"]), 2)
    if abs(items_sum - data["calculated_total"]) > 0.01:
        errors.append(f"line_items add up to {items_sum} but calculated_total is "
                      f"{data['calculated_total']}.")
    if data["invoice_date"] and not re.fullmatch(r"\d{4}-\d{2}-\d{2}", data["invoice_date"]):
        errors.append(f"invoice_date '{data['invoice_date']}' is not YYYY-MM-DD.")
    if (data["stated_total"] is not None and not data["conflict_detected"]
            and abs(data["stated_total"] - data["calculated_total"]) > 0.01):
        errors.append("stated_total differs from calculated_total. Recheck the line "
                      "items; if the document itself disagrees, set conflict_detected.")
    return errors

def extract_with_retry(document, max_retries=2):
    data = extract(document)
    for _ in range(max_retries):
        errors = validate(data)
        if not errors:
            return data, []
        data = extract(document, failed=data, errors=errors)
    return data, validate(data)   # errors left: route to human review

A record with conflict_detected set to true passes these checks, so route it to review in the calling code.

Valid is not the same as right: test with an eval set

Validation code can only check what it can compute. Some wrong extractions pass every check. A message says "ordered on 3 March, delivered on 12 April" and the pipeline returns 12 April as the order date: the date is well-formed, the field is filled, and nothing fails. Only a graded set of cases with known answers catches this before a customer does.

Build one before launch:

  1. Collect real documents and write the expected output for each.
  2. Add edge cases: two dates in one sentence, no date, a relative date such as "next Tuesday". Claude can draft more edge cases from a starting set; a person checks every expected answer.
  3. Grade format with code (parse, range, required fields) and meaning with a model grader that has a clear rubric, with each score band defined (for example, low means required content is missing). Anthropic's guide to building evaluations favours more cases with automated grading over a few graded by hand.
  4. Re-run the set after every prompt change, one change at a time, and read the failures case by case.

The eval does not fix the extraction. It finds the failure and keeps it from coming back after the next change.

Feedback loops for review findings

The same idea works for code review in CI. Give every structured finding a detected_pattern field, such as string-built-sql or broad-except. When developers dismiss a finding, log the pattern. After a few weeks you can see which patterns are dismissed most and fix them: tighter criteria (task statement 4.1) or a counter-example (task statement 4.2). Free-text findings cannot be grouped this way.

Rules that decide exam answers

  • Send the specific error, not "try again". The follow-up needs the document, the failed extraction and each validation error.
  • Do not retry for information that is not there. If the data lives in a document you did not provide, fetch it or escalate.
  • Schema, validation and evals catch different errors. A schema removes syntax errors, validation code finds wrong sums and misplaced values, and only a graded eval set catches a valid-looking wrong value.
  • Separate Claude's mistakes from the source's mistakes. calculated_total against stated_total, plus conflict_detected, tells you which one you have.
  • Cap retries and route the rest to a person. Unlimited retries waste calls on cases that cannot succeed.
  • Make findings analysable. A detected_pattern field turns dismissals into data you can act on.

Where it appears in the exam

Prompt Engineering & Structured Output is a primary domain in two of the six exam scenarios: Claude Code for Continuous Integration and Structured Data Extraction. Retry loops and self-check fields fit the extraction scenario. The detected_pattern feedback loop fits review findings in the CI scenario.

Two sample questions

These are original Timo practice questions. They are not official exam questions.

Question 1

A pipeline extracts renewal dates from supplier contracts, and validation fails on about 6% of them. The logs show two groups: dates written as "1st of March next year" that fail the YYYY-MM-DD check, and contracts that say "renewal terms as per the Master Services Agreement", which is not in the input. The team plans to retry every failure up to five times. What should they do instead?

Answer: B. Format errors are fixable with specific feedback, while missing information needs the other document or a person. A repeats calls that cannot succeed, C pushes Claude to invent dates, and D silently loses records.

Question 2

A CI review bot posts findings as free text, and developers dismiss about 40% of them. The team wants to know which kinds of finding are dismissed most, so it can fix the right criteria. What change makes that analysis possible?

Answer: C. A pattern field on every finding lets you count dismissals per pattern and target the fix. A depends on developers writing consistent reasons, B changes what is posted without explaining dismissals, and D filters on confidence without telling you which patterns fail.

Build exercise

  1. Run the code above on 10 invoices, including one whose printed total disagrees with its line items. Log every validation error.
  2. Compare a retry that says only "try again" with one that includes the document, the failed extraction and the errors. Count which fixes more records.
  3. Add a contract that refers to a document you did not supply. Confirm that retries do not fix it and route it to review instead.
  4. Add detected_pattern to a set of review findings, mark some as dismissed, and group the dismissals by pattern.

Practise this topic

Sources