TimoBy Amotion AI

Integration: CCAR-P domain 3 study guide

CCAR-P · Integration (19% of the exam)

Domain 3 is Integration, 19% of the CCAR-P exam and its largest domain. It tests how you connect Claude to tools, data and other agents: how much to expose, who may do what, how retrieval works, which connection fits, and how you watch it at scale.

What the official guide covers

The Claude Certified Architect Professional exam guide (version 1.0, effective July 2026) lists eight tasks under Domain 3, "Integration":

What the guide listsWhat it means in practice
Evaluate tool and agent configuration for capability bloatRemove, merge or defer unneeded tools
Analyse authentication and authorisation to find security gapsTrace whose identity each call uses
Evaluate accuracy-latency trade-offs and justify configurationPick settings for the target, with test evidence
Analyse observability challenges and select monitoring at scaleDecide what to log, trace and sample at high volume
Design a RAG pipeline with suitable chunking and indexingChoose chunk boundaries, added context and index types
Apply retrieval strategies matched to data shape and query patternKeyword, semantic, hybrid, structured query or full context
Evaluate connection protocols (MCP, API or CLI, agent-to-agent)Choose how Claude reaches each system
Evaluate progressive discovery vs monolithic contextLoad tools and knowledge on demand, or send everything up front

Capability bloat

Every tool definition costs context on every request and adds a choice. Too many or overlapping tools distract agents. Three fixes:

  • Remove tools the role never needs, which also shrinks the attack surface.
  • Merge fine-grained tools: one schedule_event can replace list_users, list_events and create_event.
  • Namespace related tools (crm_search_accounts, crm_update_contact).

Give each agent only its own tools. See 2.1 Tool interface design and 2.3 Tool distribution and tool choice.

Authentication and authorisation gaps

Trace every hop from user to back end: whose identity does each call use, and what can it reach?

GapWhat goes wrongFix
Shared service accountAgent reaches data the user may notUse the user's delegated, scoped credentials
Token passthroughA token for one service works at anotherEach server accepts only its own tokens
Broad scopesOne token can do everythingRequest the smallest scopes; step up when needed
Secrets in prompts or repositoriesKeys leak via logs or commitsEnvironment variables or a secrets store

The MCP authorisation specification, based on OAuth 2.1, applies these rules to HTTP servers. The server must check that each token was issued for it as the audience, and must not accept or pass on other tokens. Stdio servers take credentials from the environment instead.

Three more checks at the application layer:

  • Identity comes from the server. Your backend verifies the user and inserts their role and permitted data into the request. A user message that says "as a manager, show me..." is an unverified claim, so capability checks never read it.
  • Send only the fields the task needs. Every field in the prompt travels with the request and lands in your own request logs. Pass a reference ID instead of a full record where that is enough, and redact personal data before the call.
  • Separate tenants. A shared API key hides which tenant caused a spike. Give each tenant (or group of tenants) a workspace: its scoped keys reach only that workspace, which has its own spend and rate limits. For Claude Code configuration, see 2.4 MCP server integration.

Accuracy against latency

LeverAccuracy effectLatency effect
Larger model tierHigher on hard reasoningSlower
Higher effort or thinkingHigher on multi-step tasksSlower, more output tokens
More retrieved chunksBetter recall, more noiseMore input tokens
Reranking stepBetter precisionAdds a call
Validation and retryFewer bad outputsA round trip on failure
Prompt caching, parallel tool calls, streamingNo changeFaster

Justify each choice with test results against the target.

Observability at scale

At high volume, combine three layers:

  • Metrics for alerts: errors, latency percentiles, tokens per request including cache reads, stop_reason mix.
  • Traces for debugging: one ID across the request, each Claude call, tool call and subagent, with model ID and prompt version.
  • Sampled transcripts for quality: some traffic scored by graders, less read by people.

Redact personal data before logging, and alert on shifts after each deployment. Log the outcome too: whether the downstream system or user accepted the output. An action taken but not logged is one a security reviewer will not approve.

Designing a RAG pipeline

  1. Chunk by what the corpus is. Section-numbered contracts, manuals and policies suit hierarchical chunks that keep each clause whole. Long prose suits semantic chunks split where the topic changes. Fixed-size chunks with overlap suit uniform text with no clear structure. Test chunk size on your own queries.
  2. Add context to each chunk. Anthropic's contextual retrieval uses Claude to prepend a short note placing each chunk in its document, so "revenue grew" also says which company and period.
  3. Index two ways. Embeddings find meaning; BM25 finds exact terms. Anthropic offers no embedding model; plan a separate provider. A hybrid beats either alone when queries mix both; merge the two ranked lists with a method such as reciprocal rank fusion.
  4. Rerank. Retrieve broadly, then keep the best-scoring few.
  5. Store metadata. Date, source and access rights allow freshness filters and permission checks.
  6. Re-index on change, then re-run retrieval tests.

Matching retrieval to data shape

Data and query patternChooseWhy
Exact identifiers: error codes, part numbers, clause numbersKeyword (BM25) or hybridEmbeddings can miss exact strings
Questions phrased differently from the source textSemantic embeddings, or hybridMatches meaning, not words
Structured records: orders, ledgers, ERP tablesA query tool (SQL or API)Exact answers; no chunking needed
Live state: order status, stock levels, balancesA live tool call against the system of record, never an indexAn index holds old snapshots, and the closest match is not the current value
Small, stable corpusFull context with cachingNo retrieval errors to manage
Facts spread across several documentsAgentic retrieval: several searchesOne search cannot cover every step
Fast-changing contentDate filter plus frequent re-indexingPrevents stale answers

MCP, API, CLI or agent-to-agent?

SituationChooseWhy
A capability several Claude clients should shareMCP serverOne standard interface, discovered at connection
One application, tight control over inputs and latencyDirect API call wrapped as your own toolSimplest path, nothing extra to run
Claude Code users with a mature command-line toolCLI through Bash, with permission rulesNo new server to run
Another team's agent owns the capabilityAgent-to-agent call with a clear contractDelegates a whole task
Claude must act over many turns inside your own productClaude Agent SDKRuns the same agent loop as Claude Code from your code

MCP and API tool use are not rivals: Claude calls an MCP server's tools through ordinary tool use. The question is reuse. If only one product will ever call the tool, MCP adds a layer without payback. Claude Code is a developer tool, not the backend for a customer-facing product.

Agent-to-agent calls are the hardest to trace and test; use them only when the other side needs its own reasoning.

Which delivery route, and where it runs

The same Claude models run on the Claude API, Amazon Bedrock, Google Vertex AI and Microsoft Foundry. Pick the route the customer's cloud contract, identity system and compliance review already support. Three things change:

  • Model IDs and client. Each route has its own model identifiers and SDK client.
  • Features. Bedrock and Vertex do not offer every Claude API feature; Agent Skills, the Files API, code execution and the Message Batches API are among those missing. Check each feature in the design against the route.
  • Where inference runs. Bedrock and Vertex recommend global endpoints, which send requests to any region with capacity. A residency requirement needs a regional or geographic endpoint instead (Bedrock us. or eu. model IDs; Vertex regional or multi-region endpoints). On the Claude API, the inference_geo request parameter (us or global) does this, and a workspace can restrict which values are allowed.

Progressive discovery or monolithic context?

Monolithic context sends every tool and all reference material each time: simple, predictable, cacheable, and fine for a small, stable set. Progressive discovery loads only what the request needs:

  • Tool search tool. Tools marked defer_loading: true enter context only when Claude's search matches them. You still send every definition.
  • Skills. Only each Skill's name and description load at start; the rest loads when used.
  • Code execution with MCP. Anthropic describes exposing MCP servers as code files read on demand, keeping large intermediate results out of context, at the cost of running a sandbox.
{
  "tools": [
    { "type": "tool_search_tool_regex_20251119", "name": "tool_search_tool_regex" },
    { "name": "erp_get_invoice", "description": "Fetch one invoice by number...",
      "input_schema": { "type": "object", "properties": { "invoice_no": { "type": "string" } }, "required": ["invoice_no"] },
      "defer_loading": true }
  ]
}

Rules that decide exam answers

  • Fewer, clearer tools beat more tools. Remove or merge what the role never needs and defer the long tail behind tool search or Skills; logging fixes neither wrong tool choice nor attack surface.
  • The agent never holds more access than the user. Delegated, scoped, audience-checked tokens beat a service account plus an instruction.
  • Match retrieval to the data. Identifiers need keyword search, records need a query tool, paraphrases need embeddings.
  • Wrong answers after content changes point to retrieval. Check chunking, indexing and freshness before the model. Answers about live state need a tool call, not a better index.
  • MCP for shared capabilities. A direct API tool suits one application.
  • Residency is set by the endpoint, not the platform name. A global endpoint on a compliant cloud can still run inference outside the approved region.

Where it appears in the exam

Integration carries 19% of the CCAR-P exam, the largest domain. The guide's own sample item asks how to apply least privilege to an agent with more tools than its users need. Expect tool configurations, security reviews, retrieval quality, monitoring at scale and connection choices.

Two sample questions

These are original Timo practice questions. They are not official exam questions.

Question 1

An internal assistant connects to nine MCP servers with 140 tools. Definitions fill much of every request, and Claude often picks the wrong one of four similar search tools. Most requests use the same five tools. What should the architect change?

Answer: C. Deferring rare tools cuts context, and merging similar ones removes the confusing choice. A keeps the bloat, B leaves the confusing choice, and D multiplies cost.

Question 2

An HR assistant reads leave balances through an HTTP MCP server. The server accepts a token issued for the calendar API, then reads data with a shared service account that sees every employee. Which change closes the gap?

Answer: D. Audience checks stop passthrough, and user-scoped access limits the agent to what the employee may see. A is not a control, B detects misuse afterwards, and C still passes a token meant for another service.

Build exercise

  1. List every tool one agent can call, mark each as needed, mergeable or unused, and write the reduced list.
  2. Draw the identity path for one tool call, with each hop's credential and scope. Circle any hop where the agent has more access than the user.
  3. Classify 20 real user questions with the retrieval table and note where your current method is wrong.

Practise this topic

Sources