TimoBy Amotion AI

Claude Code in CI/CD: CCAR-F task statement 3.6

CCAR-F · Claude Code Configuration & Workflows (20% of the exam)

Task statement 3.6 sits in Claude Code Configuration & Workflows, 20% of the CCAR-F exam. It tests how you run Claude Code in a pipeline with nobody watching: the flags that make it non-interactive, how to get findings as JSON a script can post, and what context makes reviews and generated tests worth having.

What the official guide covers

The Claude Certified Architect Foundations exam guide (version 1.0, effective July 2026) lists this under task statement 3.6, "Integrate Claude Code into CI/CD pipelines":

Knowledge ofSkills in
The -p (--print) flag runs Claude Code non-interactively in automated pipelinesRunning Claude Code in CI with -p so it never waits for interactive input
--output-format json and --json-schema enforce structured output in CICombining them to produce machine-parseable findings that a script posts as inline PR comments
CLAUDE.md gives CI runs their project context: testing standards, fixture conventions, review criteriaPassing prior review findings when re-running after new commits, and asking only for new or still-open issues
The session that generated code reviews it less well than an independent review instanceGiving existing test files as context so generated tests do not repeat covered scenarios
Documenting testing standards, what makes a test valuable, and available fixtures in CLAUDE.md

Running Claude Code without a person

  • claude -p "prompt" runs once, prints the result to stdout and exits. It exits with 0 on success and a non-zero code on failure, so the pipeline can branch on it.
  • It reads stdin, so you can pipe in a diff or a log: gh pr diff 123 | claude -p "Review this diff".
  • Nobody is there to approve a tool call. Pre-approve what the job needs with --allowedTools, using permission rule syntax such as "Read,Bash(npm test *)". For a locked-down run, --permission-mode dontAsk denies anything that is not pre-approved.
  • --max-turns caps the number of agentic turns in print mode and exits with an error when the cap is reached.
  • --bare skips auto-discovery of hooks, skills, MCP servers and CLAUDE.md, so every machine gets the same result. That also drops your CLAUDE.md. If the job depends on standards in CLAUDE.md, either do not use --bare or pass the standards with --append-system-prompt-file. Without --bare, -p loads what an interactive session would, including hooks in the checked-out .claude/settings.json and servers in .mcp.json, with no trust prompt. Keep that in mind when the job runs on code from someone else's branch.
  • For a job in two steps, such as plan then implement, read session_id from the JSON output of the first call and pass it to claude -p --resume "$id" in the second. The second call has the first call's full context.

Structured findings with --json-schema

--output-format json returns one JSON object with the text answer in result plus metadata such as session_id. Add --json-schema with a JSON Schema, and Claude Code returns output that matches the schema in the structured_output field. An invalid schema stops the run with an error.

#!/usr/bin/env bash
# review.sh <pr-number>: writes findings.json for a later step to post as inline comments
set -euo pipefail
PR="$1"

SCHEMA='{
  "type": "object",
  "properties": {
    "findings": {
      "type": "array",
      "items": {
        "type": "object",
        "properties": {
          "file":     {"type": "string"},
          "line":     {"type": "integer"},
          "severity": {"type": "string", "enum": ["bug", "security", "missing-test"]},
          "message":  {"type": "string"}
        },
        "required": ["file", "line", "severity", "message"]
      }
    }
  },
  "required": ["findings"]
}'

gh pr diff "$PR" | claude -p \
  "Review this diff using the review criteria in CLAUDE.md.
   Findings from the previous review are in prior-findings.json.
   Report only issues that are new or still not fixed. Do not repeat resolved findings." \
  --output-format json \
  --json-schema "$SCHEMA" \
  --allowedTools "Read" \
  --max-turns 10 \
  | jq '.structured_output.findings' > findings.json

Each finding has a file, a line and a message, which is what an inline PR comment needs. Passing prior-findings.json on every re-run is how you stop the same comment appearing after each push.

Context that makes CI output useful

A CI run starts with no memory of your team's habits. CLAUDE.md is where it gets them. For where CLAUDE.md files live and how they layer, see 3.1 CLAUDE.md hierarchy. A section for CI might read:

## Review criteria
- Report bugs, security issues and missing tests. Skip style: the linter handles it.

## Testing standards
- Tests use pytest. Shared fixtures are in tests/conftest.py: db_session, fake_clock, payments_stub.
- A valuable test checks behaviour a user would notice. Do not test getters or framework code.
- Name tests test_<function>_<condition>.

For test generation, also give Claude the existing tests for the changed code, so it can see which scenarios are already covered. For writing review criteria that keep false positives down, see 4.1 Prompts with explicit criteria.

Run the review as its own claude -p call, not as a continuation of the session that wrote the code. A fresh session sees the diff and your criteria, not the reasoning that produced the change, so it is not biased towards code it just wrote.

GitHub Actions

Anthropic's GitHub Action, anthropics/claude-code-action@v1, runs Claude Code inside a workflow. Run /install-github-app in Claude Code for guided setup, or add the app, an ANTHROPIC_API_KEY secret and a workflow file yourself. With a prompt input the action runs on any event you choose; without one it waits for an @claude mention. claude_args passes CLI flags. It reads the repository's CLAUDE.md on every run, so keep that file short.

name: Claude test suggestions
on:
  pull_request:
    types: [opened, synchronize]
jobs:
  suggest-tests:
    runs-on: ubuntu-latest
    permissions:
      contents: read
      pull-requests: read
      id-token: write
    steps:
      - uses: actions/checkout@v6
        with:
          fetch-depth: 0
      - uses: anthropics/claude-code-action@v1
        with:
          anthropic_api_key: ${{ secrets.ANTHROPIC_API_KEY }}
          prompt: |
            Run git diff origin/${{ github.base_ref }}...HEAD to see this change.
            Read the existing tests for each changed file. Suggest tests only for
            behaviour those tests do not already cover. Follow the testing standards
            and fixtures in CLAUDE.md.
          claude_args: '--max-turns 8 --allowedTools "Read,Grep,Glob,Bash(git diff *)"'

In this automation mode the result goes to the workflow run log unless the prompt tells Claude to post and it has a tool that can.

Managed review, the Action, or your own script?

NeedChooseWhat to know
Review every pull request, nothing to hostCode Review (managed, Team and Enterprise, research preview)Posts inline findings tagged by severity; never approves or blocks a merge; reads CLAUDE.md and a review-only REVIEW.md; apply findings locally with /code-review --fix
Claude acts in the repository: implements from an @claude comment, runs a scheduled reportanthropics/claude-code-action@v1Tune with claude_args; grant only the tools the job needs
Another CI system, or your own posting logicclaude -p with --json-schemaYour script owns the output and the exit code
Repeat work on a schedule, an API call or a GitHub event, with no runner to maintainRoutines (cloud, research preview)Runs start from a fresh clone; schedules run at most hourly

Whichever you pick, check an unattended run by its evidence, not its summary: the exit code, the structured output and the diff. A tidy summary can hide an edit to a file nobody expected.

Which fix for which pipeline problem?

ProblemFixWhy
The job must run with nobody at the keyboardclaude -pRuns once, prints and exits
A script cannot parse the review--output-format json with --json-schema; read structured_outputSame shape on every run
Re-runs after new commits repeat old commentsPass prior findings; ask for new or unresolved issues onlyClaude checks each old finding against the new code
Generated tests repeat existing onesInclude the existing test files in contextClaude can see what is covered
Generated tests are trivial or ignore fixturesWrite standards and fixtures in CLAUDE.mdEvery run reads them
The generating session also reviews its own workSeparate review instanceNo bias from its own reasoning
The job needs an API keyStore it as a CI secret and read it by name, as in ${{ secrets.ANTHROPIC_API_KEY }}A key committed to a file stays in git history and must be rotated

Rules that decide exam answers

  • CI means -p. A pipeline script that calls the claude CLI without -p (or --print) starts an interactive session that nobody can answer.
  • Machine-readable means a schema. --output-format json alone puts free text in result; --json-schema gives validated structured_output.
  • A re-review needs the last review. Without prior findings in context, Claude cannot tell a fixed issue from a new one.
  • The reviewer must not be the writer. Use an independent instance, not the session that generated the code.
  • Standing context goes in CLAUDE.md. Testing standards, fixtures and review criteria belong there, not repeated in every pipeline prompt.
  • Unattended means pre-approved, not bypassed. Use --allowedTools or dontAsk so nothing waits for a person; keep bypass permissions for an isolated container or VM.

Where it appears in the exam

Claude Code Configuration & Workflows is a primary domain in three of the six exam scenarios: Code Generation with Claude Code, Developer Productivity with Claude, and Claude Code for Continuous Integration. This task statement is the core of the Continuous Integration scenario, where Prompt Engineering & Structured Output is the other primary domain.

Two sample questions

These are original Timo practice questions. They are not official exam questions.

Question 1

A CI job runs claude -p "Review this diff" --output-format json, and a script posts the review as inline PR comments. The script often fails because the findings inside the result text are laid out differently on each run. What should the team change?

Answer: D. With --json-schema, Claude Code returns output that matches the schema in structured_output, so the script gets the same shape every time. A still depends on Claude following an instruction. B changes how output is streamed, not its shape. C cuts findings without fixing the format.

Question 2

After a developer pushes fixes to a pull request, the automated review runs again and posts 12 comments. Nine repeat issues from the first review that the new commits already fixed. What change addresses this?

Answer: B. With the earlier findings in context, Claude can check each one against the new code and skip the resolved ones. A hides real lower-severity issues and still repeats critical ones. C limits how much work Claude does, not which findings are duplicates. D stops reviewing the new commits at all.

Build exercise

  1. In a test repository, run claude -p "Summarise the last commit" --output-format json and look at the result and session_id fields.
  2. Save review.sh from above. Pipe in a real diff and check structured_output.findings with jq.
  3. Save the findings as prior-findings.json, add a commit that fixes one issue, and run the script again. Confirm the fixed issue is not reported.
  4. Add the GitHub Actions workflow with your ANTHROPIC_API_KEY secret, open a test pull request, and read the suggestions in the workflow run log.

Practise this topic

Sources