Governance, Safety and Risk: CCAR-P domain 5 study guide
CCAR-P · Governance, Safety & Risk Management (14% of the exam)
Domain 5 is Governance, Safety & Risk Management, 14% of the CCAR-P exam. It tests whether you place the right control at the right layer: what a prompt can steer, what code must enforce, where a person decides, and what regulation requires of the design.
What the official guide covers
The Claude Certified Architect Professional exam guide (version 1.0, effective July 2026) lists five tasks under Domain 5, "Governance, Safety & Risk Management":
| What the guide lists | What it means in practice |
|---|---|
| Implement guardrails and safety controls | Layer controls and know which are deterministic |
| Identify risks, limitations and failure modes of LLM systems | Name each failure and its control |
| Apply human-in-the-loop validation strategies | Decide which outputs a person approves before, after or by sample |
| Ensure compliance with regulations such as GDPR, HIPAA and FedRAMP | Map data flows to obligations, platforms and contracts |
| Address ethical AI considerations (bias, fairness, transparency) | Test outcomes by group; make AI use visible |
Guardrails in layers
Start by drawing the boundary. Claude's training reduces broad classes of harm on every request, but it has never seen your tenancy rules, approved scripts or refund limits. A domain rule that lives in no layer you built is not enforced, however safe Claude looked in testing.
Anthropic's guidance on jailbreaks and prompt injection recommends chaining several controls. Know which layers are guaranteed.
| Layer | Control | Type |
|---|---|---|
| Input | Screen input with a lightweight model; filter known injection patterns | Probabilistic |
| Processing | Rules and refusals in the system prompt; third-party content only in tool_result blocks, labelled untrusted | Probabilistic |
| Actions | Least-privilege tools; a PreToolUse hook or permission rule that blocks forbidden calls | Deterministic |
| Output | Schema validation, checks for personal data, citation checks | Deterministic or graded |
| Monitoring | Review outputs for signs of successful injection; throttle repeat offenders | Detective |
Screening inputs, screening outputs and authorizing tool calls each answer a separate question, so a filter on the output does nothing about a refund the agent already issued. Use model-based screens for fuzzy intent and deterministic rules for known patterns; authorization should always be deterministic (identity, scope and an allowlist) so you can replay and prove it. Screen retrieved documents and tool results too, because injected instructions arrive there after input screening has passed. Decide what each guardrail you built does when it errors: one that fails open lets traffic through unchecked while looking healthy, so fail closed wherever a wrong pass causes harm.
If something must never happen, the control belongs in code at the action layer. A prompt lowers the odds; a hook removes the path. See 1.5 Agent SDK hooks.
Failure modes and their controls
| Failure mode | What it looks like | Primary control |
|---|---|---|
| Hallucination | Confident detail not in any source | Ground in provided documents, require citations, allow "I don't know" |
| Direct prompt injection | User tells the model to ignore its rules | Input screening, refusal rules, action-layer limits |
| Indirect prompt injection | An email or web page contains instructions | Treat tool content as data; approval for sensitive actions |
| Excess permissions | Agent can do more than the task needs | Remove tools; scoped, delegated credentials |
| Data leakage | Personal data in outputs or logs | Redaction, output checks, log policies |
| Inconsistency | Same input, different answers | Structured output, multi-trial tests |
| Token-budget exhaustion | Oversized or padded input truncates work or inflates cost | Input size limits, chunking, budget alerts |
| Runaway agents | Loops, spend or subagent trees grow | Turn, depth and budget limits |
| Model retirement | Requests to a retired model ID fail | Version register and migration tests |
| Untrusted Skill | Bundled scripts reach the network, files or credentials when run | Approved sources only, read the bundle against its stated purpose, run sandboxed with least privilege |
Human in the loop
Choose the review pattern by the cost of a wrong action.
| Situation | Pattern | Why |
|---|---|---|
| Irreversible or high-value action (payments, account changes) | Approve before acting, enforced in code | Cannot be undone |
| Reversible action with low harm | Act, then review a sample, with rollback | Keeps speed; finds systematic errors |
| Unclear or outside policy | Escalate with a case summary | Avoids guessing |
| High volume, mostly routine | Route by risk to people | Review effort where it matters |
Stakes come from reversibility and the cost of a wrong decision; confidence only estimates how likely an output is to be wrong, and only if it is calibrated. Route low-confidence decisions that are irreversible or costly; let confident, reversible, cheap ones through. Sending everything to review floods the queue until people approve without reading. Give reviewers the inputs, the output and the reason the item was flagged, or review becomes a rubber stamp, and add every override to the test set. See 5.2 Escalation and ambiguity and 5.5 Human review and confidence.
Compliance: what Anthropic states and what the architect owns
The architect maps where regulated data enters, is processed, stored and logged, chooses deployment options and contracts that cover those flows, and involves legal and compliance early. An agreement covers the provider's part, not your system. What Anthropic's own pages state:
- GDPR. Anthropic's privacy centre says its Data Processing Addendum, with Standard Contractual Clauses, is incorporated into its Commercial Terms for commercial products such as the Claude API. Customers who use Claude through a third-party platform are covered by that platform's terms instead.
- HIPAA. Anthropic describes a HIPAA-ready configuration with a Business Associate Agreement for Enterprise and API customers. Its BAA page lists which features are covered and which are not; at the time of checking, the Batch API, Files API and code execution were among those listed as not covered. Check the current list against every feature your design uses before any protected health information flows.
- FedRAMP. Anthropic's public sector FAQs state that Claude for Government is FedRAMP High authorised, and that Claude is available at FedRAMP High through Amazon Bedrock in AWS GovCloud and through Google Vertex AI with Assured Workloads. The same page notes that FedRAMP certifies the hosting platform, not the model, so the platform you deploy on decides the authorisation.
Everything else (residency, retention, access reviews, records of processing, breach procedures) is your organisation's design work. A compliant route is required, but it proves nothing on its own. Keep a control register: each obligation, the control that meets it, a named owner and an artifact a reviewer can open as evidence (a signed agreement, a configuration screen, a query that returns the log). Revalidate it on a schedule, because a logging change can quietly start copying data out of region. Keep "not used for training" and "not retained" as separate claims.
Bias, fairness and transparency
- Bias. Break results down by the groups you serve (language, region, segment) and compare outcomes, not only averages.
- Fairness. Keep protected characteristics out of decision inputs without a documented, lawful reason, and put consequential decisions about people under human review.
- Transparency. Tell users when they deal with AI, cite sources, and log enough to explain an output later.
Skew enters at four points you control: the retrieval corpus, the prompt's framing, the few-shot examples and the routing after the output. A model that passed its vendor's bias tests can still give unequal results over your corpus. Log each decision's inputs, retrieved context, output and routing so you can answer an affected user, a regulator and your own engineers. That log holds personal data, so it needs minimisation, retention limits and access control too.
Example: a control matrix for an accounts-payable agent
| Risk | Control | Enforced by | Evidence |
|---|---|---|---|
| Injected email text changes bank details | Bank-detail changes always need human approval | PreToolUse hook denies update_bank_details without an approval ID | Hook logs; weekly denial count |
| Invoice paid twice | Duplicate check on invoice number and amount | Code before create_payment | Unit tests; incident log |
| Personal data in summaries posted to Slack | Redact names and account numbers | Screen before posting | Sampled output review |
| Wrong amounts extracted | Line items must sum to the total, or a person checks | Validation code | 98% field accuracy gate |
| Agent loops on a bad PDF | Turn and spend limits | Agent SDK options | Alert when a limit is hit |
| Model ID retired | Version register, migration test | Release checklist | Regression results per upgrade |
Rules that decide exam answers
- "Must never happen" means a deterministic control. Hooks, permissions and validation in code beat stronger prompt wording, a larger model or after-the-fact audits.
- Remove a risk before you monitor it. Least privilege comes first; logging and confirmation are compensating controls.
- Third-party content is data, never instructions. Deliver it in tool results, label it untrusted, and gate the sensitive actions it could trigger.
- Review by risk, with context. Approve-before-act for irreversible actions; samples for routine ones; reviewers see the sources.
- A contract is not compliance. Check which features and platforms an agreement covers, then design the rest of the data flow to fit.
- Measure fairness by group. An average accuracy figure can hide a group the system fails.
Where it appears in the exam
Governance, Safety & Risk Management carries 14% of the CCAR-P exam. The guide's audience includes architects who lead security, legal and executive discussions in sectors such as financial services, healthcare and government. Expect agents that can act, sensitive data, review queues and regulated deployments, asking which control closes the risk.
Two sample questions
These are original Timo practice questions. They are not official exam questions.
Build exercise
- List one Claude system's failure modes using the table above, and mark each existing control as deterministic, probabilistic or detective.
- Choose a review pattern for every action the system can take, with one sentence of justification.
- Draw the flow of one regulated data type from entry to logs, noting what covers each step.
- Break your latest evaluation results down by two user groups and note any gap beyond tolerance.
Practise this topic
- Claude Certified Architect Professional practice exam: free, 20 questions, no sign-up
- Claude Certified Architect hub
- CCAR-P study guide: all topics
- Worked example: Enterprise Claude safety stack
- Same topic in another exam: Governance, Risk and Responsible Use (CCAO-F)
- Previous topic: Evaluation, Testing and Optimization
- Next topic: Stakeholder Communication and Lifecycle
Sources
- Claude Certified Architect, Professional Exam Guide, version 1.0, effective July 2026 (Anthropic), Domain 5: Governance, Safety & Risk Management
- Anthropic documentation: Mitigate jailbreaks and prompt injections
- Claude Help Center: Public Sector FAQs
- Anthropic Privacy Center: Business Associate Agreements (BAA) for Commercial Customers
- Anthropic Privacy Center: How do I view and sign your Data Processing Addendum (DPA)?
By Amotion AI