Integration: CCAR-P domain 3 study guide
CCAR-P · Integration (19% of the exam)
Domain 3 is Integration, 19% of the CCAR-P exam and its largest domain. It tests how you connect Claude to tools, data and other agents: how much to expose, who may do what, how retrieval works, which connection fits, and how you watch it at scale.
What the official guide covers
The Claude Certified Architect Professional exam guide (version 1.0, effective July 2026) lists eight tasks under Domain 3, "Integration":
| What the guide lists | What it means in practice |
|---|---|
| Evaluate tool and agent configuration for capability bloat | Remove, merge or defer unneeded tools |
| Analyse authentication and authorisation to find security gaps | Trace whose identity each call uses |
| Evaluate accuracy-latency trade-offs and justify configuration | Pick settings for the target, with test evidence |
| Analyse observability challenges and select monitoring at scale | Decide what to log, trace and sample at high volume |
| Design a RAG pipeline with suitable chunking and indexing | Choose chunk boundaries, added context and index types |
| Apply retrieval strategies matched to data shape and query pattern | Keyword, semantic, hybrid, structured query or full context |
| Evaluate connection protocols (MCP, API or CLI, agent-to-agent) | Choose how Claude reaches each system |
| Evaluate progressive discovery vs monolithic context | Load tools and knowledge on demand, or send everything up front |
Capability bloat
Every tool definition costs context on every request and adds a choice. Too many or overlapping tools distract agents. Three fixes:
- Remove tools the role never needs, which also shrinks the attack surface.
- Merge fine-grained tools: one
schedule_eventcan replacelist_users,list_eventsandcreate_event. - Namespace related tools (
crm_search_accounts,crm_update_contact).
Give each agent only its own tools. See 2.1 Tool interface design and 2.3 Tool distribution and tool choice.
Authentication and authorisation gaps
Trace every hop from user to back end: whose identity does each call use, and what can it reach?
| Gap | What goes wrong | Fix |
|---|---|---|
| Shared service account | Agent reaches data the user may not | Use the user's delegated, scoped credentials |
| Token passthrough | A token for one service works at another | Each server accepts only its own tokens |
| Broad scopes | One token can do everything | Request the smallest scopes; step up when needed |
| Secrets in prompts or repositories | Keys leak via logs or commits | Environment variables or a secrets store |
The MCP authorisation specification, based on OAuth 2.1, applies these rules to HTTP servers. The server must check that each token was issued for it as the audience, and must not accept or pass on other tokens. Stdio servers take credentials from the environment instead.
Three more checks at the application layer:
- Identity comes from the server. Your backend verifies the user and inserts their role and permitted data into the request. A user message that says "as a manager, show me..." is an unverified claim, so capability checks never read it.
- Send only the fields the task needs. Every field in the prompt travels with the request and lands in your own request logs. Pass a reference ID instead of a full record where that is enough, and redact personal data before the call.
- Separate tenants. A shared API key hides which tenant caused a spike. Give each tenant (or group of tenants) a workspace: its scoped keys reach only that workspace, which has its own spend and rate limits. For Claude Code configuration, see 2.4 MCP server integration.
Accuracy against latency
| Lever | Accuracy effect | Latency effect |
|---|---|---|
| Larger model tier | Higher on hard reasoning | Slower |
| Higher effort or thinking | Higher on multi-step tasks | Slower, more output tokens |
| More retrieved chunks | Better recall, more noise | More input tokens |
| Reranking step | Better precision | Adds a call |
| Validation and retry | Fewer bad outputs | A round trip on failure |
| Prompt caching, parallel tool calls, streaming | No change | Faster |
Justify each choice with test results against the target.
Observability at scale
At high volume, combine three layers:
- Metrics for alerts: errors, latency percentiles, tokens per request including cache reads,
stop_reasonmix. - Traces for debugging: one ID across the request, each Claude call, tool call and subagent, with model ID and prompt version.
- Sampled transcripts for quality: some traffic scored by graders, less read by people.
Redact personal data before logging, and alert on shifts after each deployment. Log the outcome too: whether the downstream system or user accepted the output. An action taken but not logged is one a security reviewer will not approve.
Designing a RAG pipeline
- Chunk by what the corpus is. Section-numbered contracts, manuals and policies suit hierarchical chunks that keep each clause whole. Long prose suits semantic chunks split where the topic changes. Fixed-size chunks with overlap suit uniform text with no clear structure. Test chunk size on your own queries.
- Add context to each chunk. Anthropic's contextual retrieval uses Claude to prepend a short note placing each chunk in its document, so "revenue grew" also says which company and period.
- Index two ways. Embeddings find meaning; BM25 finds exact terms. Anthropic offers no embedding model; plan a separate provider. A hybrid beats either alone when queries mix both; merge the two ranked lists with a method such as reciprocal rank fusion.
- Rerank. Retrieve broadly, then keep the best-scoring few.
- Store metadata. Date, source and access rights allow freshness filters and permission checks.
- Re-index on change, then re-run retrieval tests.
Matching retrieval to data shape
| Data and query pattern | Choose | Why |
|---|---|---|
| Exact identifiers: error codes, part numbers, clause numbers | Keyword (BM25) or hybrid | Embeddings can miss exact strings |
| Questions phrased differently from the source text | Semantic embeddings, or hybrid | Matches meaning, not words |
| Structured records: orders, ledgers, ERP tables | A query tool (SQL or API) | Exact answers; no chunking needed |
| Live state: order status, stock levels, balances | A live tool call against the system of record, never an index | An index holds old snapshots, and the closest match is not the current value |
| Small, stable corpus | Full context with caching | No retrieval errors to manage |
| Facts spread across several documents | Agentic retrieval: several searches | One search cannot cover every step |
| Fast-changing content | Date filter plus frequent re-indexing | Prevents stale answers |
MCP, API, CLI or agent-to-agent?
| Situation | Choose | Why |
|---|---|---|
| A capability several Claude clients should share | MCP server | One standard interface, discovered at connection |
| One application, tight control over inputs and latency | Direct API call wrapped as your own tool | Simplest path, nothing extra to run |
| Claude Code users with a mature command-line tool | CLI through Bash, with permission rules | No new server to run |
| Another team's agent owns the capability | Agent-to-agent call with a clear contract | Delegates a whole task |
| Claude must act over many turns inside your own product | Claude Agent SDK | Runs the same agent loop as Claude Code from your code |
MCP and API tool use are not rivals: Claude calls an MCP server's tools through ordinary tool use. The question is reuse. If only one product will ever call the tool, MCP adds a layer without payback. Claude Code is a developer tool, not the backend for a customer-facing product.
Agent-to-agent calls are the hardest to trace and test; use them only when the other side needs its own reasoning.
Which delivery route, and where it runs
The same Claude models run on the Claude API, Amazon Bedrock, Google Vertex AI and Microsoft Foundry. Pick the route the customer's cloud contract, identity system and compliance review already support. Three things change:
- Model IDs and client. Each route has its own model identifiers and SDK client.
- Features. Bedrock and Vertex do not offer every Claude API feature; Agent Skills, the Files API, code execution and the Message Batches API are among those missing. Check each feature in the design against the route.
- Where inference runs. Bedrock and Vertex recommend global endpoints, which send requests to any region with capacity. A residency requirement needs a regional or geographic endpoint instead (Bedrock
us.oreu.model IDs; Vertex regional or multi-region endpoints). On the Claude API, theinference_georequest parameter (usorglobal) does this, and a workspace can restrict which values are allowed.
Progressive discovery or monolithic context?
Monolithic context sends every tool and all reference material each time: simple, predictable, cacheable, and fine for a small, stable set. Progressive discovery loads only what the request needs:
- Tool search tool. Tools marked
defer_loading: trueenter context only when Claude's search matches them. You still send every definition. - Skills. Only each Skill's name and description load at start; the rest loads when used.
- Code execution with MCP. Anthropic describes exposing MCP servers as code files read on demand, keeping large intermediate results out of context, at the cost of running a sandbox.
{
"tools": [
{ "type": "tool_search_tool_regex_20251119", "name": "tool_search_tool_regex" },
{ "name": "erp_get_invoice", "description": "Fetch one invoice by number...",
"input_schema": { "type": "object", "properties": { "invoice_no": { "type": "string" } }, "required": ["invoice_no"] },
"defer_loading": true }
]
}
Rules that decide exam answers
- Fewer, clearer tools beat more tools. Remove or merge what the role never needs and defer the long tail behind tool search or Skills; logging fixes neither wrong tool choice nor attack surface.
- The agent never holds more access than the user. Delegated, scoped, audience-checked tokens beat a service account plus an instruction.
- Match retrieval to the data. Identifiers need keyword search, records need a query tool, paraphrases need embeddings.
- Wrong answers after content changes point to retrieval. Check chunking, indexing and freshness before the model. Answers about live state need a tool call, not a better index.
- MCP for shared capabilities. A direct API tool suits one application.
- Residency is set by the endpoint, not the platform name. A global endpoint on a compliant cloud can still run inference outside the approved region.
Where it appears in the exam
Integration carries 19% of the CCAR-P exam, the largest domain. The guide's own sample item asks how to apply least privilege to an agent with more tools than its users need. Expect tool configurations, security reviews, retrieval quality, monitoring at scale and connection choices.
Two sample questions
These are original Timo practice questions. They are not official exam questions.
Build exercise
- List every tool one agent can call, mark each as needed, mergeable or unused, and write the reduced list.
- Draw the identity path for one tool call, with each hop's credential and scope. Circle any hop where the agent has more access than the user.
- Classify 20 real user questions with the retrieval table and note where your current method is wrong.
Practise this topic
- Claude Certified Architect Professional practice exam: free, 20 questions, no sign-up
- Claude Certified Architect hub
- CCAR-P study guide: all topics
- Worked example: Claude MCP server integration
- Previous topic: Models, Prompting and Context
- Next topic: Evaluation, Testing and Optimization
Sources
- Claude Certified Architect, Professional Exam Guide, version 1.0, effective July 2026 (Anthropic), Domain 3: Integration
- Anthropic documentation: Tool search tool
- Model Context Protocol: Authorization
- Anthropic: Introducing Contextual Retrieval
- Anthropic documentation: Data residency
By Amotion AI