TimoBy Amotion AI

Information provenance: CCAR-F task statement 5.6

CCAR-F · Context Management & Reliability (15% of the exam)

Task statement 5.6 sits in Context Management & Reliability, 15% of the CCAR-F exam. It tests how a multi-agent system keeps every claim tied to its source through summarising and synthesis, and how it reports sources that disagree or describe different periods.

What the official guide covers

The Claude Certified Architect Foundations exam guide (version 1.0, effective July 2026) lists this under task statement 5.6, "Preserve information provenance and handle uncertainty in multi-source synthesis":

Knowledge ofSkills in
Source attribution is lost when summarising compresses findings without keeping claim-source mappingsRequiring subagents to output claim-source mappings (source URLs, document names, relevant excerpts) that later agents keep through synthesis
The synthesis agent must preserve and merge structured claim-source mappingsStructuring reports with separate sections for well-established and contested findings, keeping each source's own wording and method
Conflicting statistics from credible sources are annotated with their sources, not settled by picking oneFinishing document analysis with conflicting values included and labelled, and letting the coordinator decide how to reconcile them
Publication and collection dates stop differences in time being read as contradictionsRequiring subagents to include publication or data collection dates in structured output
Presenting each content type in a suitable form: financial data as tables, news as prose, technical findings as structured lists

Where attribution gets lost

A research system passes findings along a chain: a search subagent finds pages, an analysis subagent reads them, a synthesis agent combines the results, and a report comes out. Each step compresses. If a finding travels as a sentence ("dwell times rose to about six days"), the source, the exact figure, the period and the method fall away at the first summary. The synthesis agent then cannot say who said what, and a reader cannot check it.

The fix is to make the source part of the data, not part of the prose. The same applies to what goes into each agent: Anthropic's prompting guidance recommends wrapping each document in <document> tags with a <source> tag and other metadata, so Claude can tell documents apart. Anthropic's own research system goes one step further, with a separate citation agent that reads the documents and the draft report to find the exact places to cite.

Provenance starts when documents are indexed

In a retrieval system, attribution can be lost before any agent runs. If the index stores only text chunks and their embeddings, a retrieved passage arrives with no title, date or location, and no later step can add them back. Store the source fields with every chunk:

def make_chunks(doc: dict, sections: list[dict]) -> list[dict]:
    """doc: title, url, published. sections: heading, text, page, split on document headings."""
    return [
        {
            "text": f"{s['heading']}\n{s['text']}",  # keep the heading with its text
            "source_title": doc["title"],
            "source_url": doc["url"],
            "published": doc["published"],
            "page": s["page"],
            "chunk_id": f"{doc['url']}#p{s['page']}-{i}",
        }
        for i, s in enumerate(sections)
    ]

Keeping the heading with its text matters more than it looks. A heading such as "Second quarter, import containers" often carries the period and scope of the figures under it. A fixed-size split can cut a table away from that heading, and the number then travels without the facts that make it comparable. Splitting on the document's own sections, or adding overlap, keeps them together. When a chunk only makes sense with the rest of its document, Anthropic's contextual retrieval technique goes further: Claude writes a short note that places the chunk in its document (which company's filing, which quarter), and the note is added to the chunk before it is embedded and indexed for keyword search. The retrieved chunk then hands its source fields straight to the subagent's claim-source output.

Test that hand-off with real retrieved chunks, not only each side on its own. If the prompt builder drops the chunk text or its source fields, Claude answers from its own knowledge, and the report cites nothing it actually read.

A claim-source mapping

Ask every subagent to return findings in a fixed structure:

{
  "claim_id": "c-014",
  "claim": "Average import container dwell time at the port reached 6.1 days",
  "excerpt": "average import dwell time reached 6.1 days in the second quarter",
  "source": {
    "title": "Port authority quarterly statistics",
    "url": "https://example.org/port-q2-2026",
    "type": "official statistics"
  },
  "published": "2026-07-15",
  "data_period": "2026-Q2",
  "method": "mean across all import containers"
}

The synthesis agent's instructions then say: merge findings by topic, keep every claim_id and its source fields, and make every sentence in the report trace back to at least one claim. A finding with no mapping does not go in the report.

Conflicting figures and different dates

Two credible sources give different numbers. Before calling it a conflict, compare the dates and methods:

SourceFigurePeriodMethod
Port authority6.1 days2026 Q2Mean, all import containers
Shipping association4.8 days2025, full yearMedian, refrigerated containers only

This is not one fact with two values. It is two different measures from two periods. Without published, data_period and method in the output, the synthesis agent would see a contradiction, or quietly pick one.

When two sources really do disagree about the same measure, the analysis step keeps both values, each with its source, and marks them as conflicting. The coordinator decides how to reconcile them before synthesis. The report then shows both and says the finding is contested. Averaging, picking the newer one or picking the "more official" one all throw away information the reader needs.

Write the coordinator's conflict rule down before the run, not in the moment. For example: same measure, same period, both credible, so report both as contested; same measure, different periods, so report as a change over time; one source cannot be checked, so report it with that caveat. When a conflict matters to a decision and no rule covers it, send it to a person rather than letting the synthesis step settle it. Log which claim IDs and chunks fed each finding, so anyone can later trace a sentence in the report back to its sources.

Citations when Claude reads your documents

When Claude answers from documents you pass in, the Citations feature returns the exact passages it relied on. Turn it on per document block:

import os
import anthropic

client = anthropic.Anthropic()
report_text = open("port_q2_report.txt").read()

response = client.messages.create(
    model=os.environ["CLAUDE_MODEL"],   # a current model ID
    max_tokens=1024,
    messages=[{
        "role": "user",
        "content": [
            {
                "type": "document",
                "source": {"type": "text", "media_type": "text/plain", "data": report_text},
                "title": "Port authority quarterly statistics",
                "context": "Published 2026-07-15. Covers 2026 Q2.",
                "citations": {"enabled": True},
            },
            {"type": "text", "text": "What was the average import dwell time? Cite the passage."},
        ],
    }],
)

for block in response.content:
    if block.type == "text":
        for c in block.citations or []:
            print(f"{c.document_title}: {c.cited_text}")

Each citation carries cited_text, document_index and document_title, plus a location: character ranges for plain text, page numbers for PDFs, block indices for custom content. The context field gives Claude information about the document that is not itself cited. One limit matters for pipeline design: citations cannot be combined with structured outputs (output_config.format) in the same request, and the API returns an error if you try. A step that must return JSON should carry the excerpt and source fields in its schema, as in the mapping above.

How the report should look

  • A section for well-established findings, supported by several independent sources.
  • A section for contested findings, showing each value with its source, date and method, in the source's own terms.
  • A gaps note for areas with little or no evidence.
  • Each kind of content in a suitable form: financial figures in tables, news events as prose, technical findings as structured lists. Forcing everything into bullet points loses the comparison a table gives.

Which approach fits

SituationChooseWhy
The final report cannot say where a figure came fromClaim-source mappings from every subagent, kept through synthesisThe source travels with the claim
Two credible sources give different valuesKeep both with attribution; mark as contestedPicking one hides real disagreement
Old and new figures look contradictoryRequire publication and data period datesChange over time is not a contradiction
Claude answers from documents you supplyCitations enabled on the document blocksReturns the exact supporting passages
A synthesis step must return strict JSONExcerpts and source fields in the schemaCitations and structured outputs cannot be combined
Retrieved passages arrive without title or dateStore source fields with every chunk at indexing timeLost metadata cannot be added back later
A conflict affects a decision and no rule covers itEscalate to a personSynthesis should not settle it silently
Retrieved passages give different values for something that changes, such as an order statusQuery the system of record with a tool; keep retrieval for stable documentsThe most similar passage is not the newest

Rules that decide exam answers

  • The source is data, not decoration. Subagents return claim, excerpt, source and date as fields; synthesis must keep them.
  • Do not resolve conflicts by picking. Report both values with their sources and let the coordinator decide; never average or choose silently.
  • Dates first, contradiction second. Many apparent conflicts are different periods or methods.
  • Separate settled from contested. A report that mixes them overstates certainty.
  • Match format to content. Tables for figures, prose for news, lists for technical findings.
  • Attribution is set at the start of the chain. If chunks or subagent outputs lack source fields, no later prompt can recover them.

Where it appears in the exam

Context Management & Reliability is a primary domain in four of the six exam scenarios: Customer Support Resolution Agent, Code Generation with Claude Code, Multi-Agent Research System and Structured Data Extraction. Provenance questions fit the Multi-Agent Research System scenario, which produces cited reports from several subagents. The guide's preparation exercise 4 tests source attribution and conflicting statistics directly and lists Domain 5 among the domains it reinforces.

Two sample questions

These are original Timo practice questions. They are not official exam questions.

Question 1

A research system finds two credible figures for the same market size: an industry body reports 4.2 billion and a national statistics office reports 3.6 billion, both for the same year. The synthesis agent currently keeps the higher figure and cites one source. What should the system do?

Answer: B. Both values stay visible with attribution, and the reader sees the disagreement. A creates a number neither source published, C picks one by rule and hides the conflict, and D drops a finding the reader needs.

Question 2

A report lists "share of staff working remotely" under contested findings, citing one survey at 38% and another at 22%. Reviewers find both are correct: one was collected in 2021 and the other in 2025. Subagents return claims without dates. What change prevents this?

Answer: C. With dates in the structured output, the synthesis agent sees a change over time, not a conflict. A hides the older figure, which is still a valid finding, B adds cost and still lacks dates, and D may or may not include the date and does not require it.

Build exercise

  1. Build a two-subagent research pipeline (search and document analysis) that returns findings in the claim-source format above.
  2. Give it two sources that disagree on one figure and two sources from different years on another. Check that the report separates the true conflict from the change over time.
  3. Remove published and data_period from the format and run it again. Note what the synthesis agent does with the dated pair.
  4. Call the Messages API on one source document with citations enabled and compare the cited_text it returns with the excerpt your subagent recorded.

Practise this topic

Sources