Control evidence for AI coding agents

Agents write the code. Your controls still apply.

AutoDevOps records every instrumented AI coding agent session — tool calls, prompt-submit events, policy decisions, approvals — and turns that record into evidence for the controls your examiners already test.

  • Pre-production only
  • Prompt text only on opt-in
  • Built for your cloud · AWS validated
Session recordedsession cc-7f2a · fincard/payments-apiDemo data
  1. 09:41:02session.startedclaude-code · payments-apiSR11-7.INV-01
  2. 09:41:09intent.bundle.attachedspec PAY-2141 · sha256:a41f…SR11-7.SND-02
  3. 09:41:33tool.startedEdit src/ledger/settle.tsEUAI.ART12-02
  4. 09:41:34approval.requestedconfirm · payments-writeEUAI.ART14-03
  5. 09:42:10approval.decidedapproved · second reviewer · rationaleFFIEC.SOD-01
  6. 09:43:05hook.blockedBash curl → public hostDORA.RES-03
  7. 09:43:41model.completed1 secret redacted before emitEUAI.ART15-04
  8. 09:46:01session.finishedPR evidence digest sha256:f9b0…FFIEC.SDLC-02SOC2.CC8-03

Controls with new evidence

9 / 16

Illustrative session. Control IDs show the catalog mapping for each event type: 8 events evidence 9 of 16 mapped controls.

Mapped control-by-control

Reference mapping to evidence you produce — not a certification.

01The control gap

Every SDLC control assumes a person wrote the change.

Change management, dual control, attribution, and logging were designed around human authors. When an agent writes the change, the evidence breaks exactly where examiners look first.

Ask these five questions about the last change an agent merged in your repos.

  1. FFIEC.SOD-01

    Who authorized this change?

    Without a session record

    A pull-request approval from a reviewer who never saw what the agent did.

    With AutoDevOps

    The approval record, bound to the session and the rule that fired, with the reviewer’s rationale when the decision was made in the portal.

  2. FFIEC.SDLC-02

    Where did this code come from?

    Without a session record

    A commit under a developer’s name. The agent, model, and prompt are invisible.

    With AutoDevOps

    Session, agent, model, and commit bound into one pull-request evidence digest.

  3. SR11-7.SND-02

    Did the agent build what was asked?

    Without a session record

    A ticket and a diff, with nothing connecting the two.

    With AutoDevOps

    The spec attached to its session, with a deterministic intent-fidelity score.

  4. EUAI.ART12-02

    Can the log be trusted?

    Without a session record

    Shell history on a laptop, if anyone kept it.

    With AutoDevOps

    Signed events in an append-only ledger, exported through hash-chained evidence packages.

  5. DORA.RES-03

    What happens when review is unavailable?

    Without a session record

    The agent keeps going.

    With AutoDevOps

    The action is denied, and the block is written to the record.1

1 Approval gates and fail-closed blocking ship in the customer-cloud distribution. The public CLI records telemetry.

02How it works

From agent session to control evidence.

One loop — the same on a laptop, in CI, and in cloud workers. Every step writes to the record; the control catalog maps each kind of record to the controls it supports.

  1. 01

    Capture

    Hooks in Claude Code, Cursor, and Codex emit signed events for sessions, prompt submits, and tool calls. Agent actions only — never keystrokes or screens.

  2. 02

    Decide

    Deterministic policy resolves each sensitive action to allow, confirm, or block before it runs. Confirm waits for a named reviewer.1

  3. 03

    Record

    Decisions, approvals, and any rationale given land in an append-only ledger. Secrets are redacted before anything is emitted.

  4. 04

    Prove

    Evidence packages and pull-request digests bind the record into files an examiner verifies offline — no raw prompts or source inside.

Approval queue · FinCard demo workspace · fictional data

Allow · Confirm · Block1

Approval queue in the FinCard demo portal: a packaging command pending second-line review, a Bedrock rerun approved by a named reviewer, and a public model endpoint call blocked by the policy engine.

ConfirmEUAI.ART14-03

A high-risk command waits for a named reviewer before it runs.

ApproveFFIEC.SOD-01

The approval is bound to the session, the command, and the reviewer.

BlockDORA.RES-03

A public model call is denied, and the block is written to the record.

1 In the customer-cloud distribution, policy decides allow, confirm, or block before the action runs.

03The control catalog

16 controls. 6 frameworks. One record.

Each control names the clause it maps to, the artifact that evidences it, and a suggested reviewer role; your organization names the owner. Every row opens its full entry.

Reading the catalog

Implementation evidence
Runtime or portal behavior produces related evidence as agents work. It does not mean the control is operating in your environment.
Monitoring signal
A signal the portal computes on read from recorded sessions and approvals.
Verifier-checkable
An artifact a local verifier checks offline, without access to your systems.
Needs your sign-off
AutoDevOps produces the receipt; a named owner in your organization signs it.

A reference mapping to evidence you produce — not a certification, and not a statement that any control is met.

Open the full catalog

04Controls as code

The policy file is the control implementation.

Governance, sensitivity, provider, and budgets — versioned in the repo and reviewed like code. The same file applies locally, in CI, and in cloud workers. The comments below show which catalog control each setting supports.

Deterministic decisions
Rules decide — never an LLM’s interpretation. Prior approvals let similar actions flow.
Replayable audit
Every decision replays with its context, on a laptop or inside a cloud worker.
Budgets with reasons
Per-commit, daily, and monthly limits. The run stops with the reason on record.
Per-agent analytics
Activity, cost, approvals, and risk per developer, team, and agent.

Policy · control notes added

Same contract everywhere
yaml
# .verifier/config.yaml
governance:            # DORA.RES-03 · policy blocks
  enabled: true
  mode: enforcement
sensitivity:           # EUAI.ART14-03 · human oversight
  enabled: true
  confirm_threshold: 70
  block_threshold: 95
providers:
  bedrock:             # DORA.TPRM-02 · your AWS boundary
    region: us-east-1
budgets:
  per_commit_tokens: 5000
Laptop
CI pipeline
Cloud worker

05Proof surface

Evidence an examiner can verify offline.

Portal views compute from the ledger. The export is a hash-bound package an auditor verifies without access to your systems; the catalog maps each of its content sets to the controls it supports.

Operator view · demo data

FinCard demo

Verification portal overview for the fictional FinCard demo workspace, showing agent health and pending decisions.

A demo workspace (FinCard, a fictional bank): sessions, approvals, and audit in one surface.

Evidence package manifest · excerpt, sample values

SHA-256 bound
json
{
  "schemaVersion": "autodevops.evidence_package.v1",
  "packageId": "epkg_7f2a…c91e",
  "generatedAt": "2026-09-30T09:52:07.214Z",
  "exporter": {
    "id": "exporter.evidence-package",
    "version": "0.1.0"
  },
  "scope": { "kind": "pr", "id": "482" },
  "redactionPolicy": {
    "rawPrompts": "excluded",
    "sourceContent": "hash_only",
    "secrets": "scrubbed"
  },
  "sourceQueryIds": ["3f9c…e21a", "b7d0…4c58"],
  "includedEventIds": ["0b4e…71c2", "1d93…a0f7"],
  "contents": {
    "agentRunAuditReports": { "rowCount": 3, "sha256": "71fd…" },
    "approvals":            { "rowCount": 2, "sha256": "9a07…" },
    "policyDecisions":      { "rowCount": 7, "sha256": "e2b4…" },
    "prMergeEvidence":      { "rowCount": 1, "sha256": "c0aa…" },
    "provenanceQueries":    { "rowCount": 2, "sha256": "5b93…" },
    "sessions":             { "rowCount": 3, "sha256": "4c1e…" }
  },
  "packageHash": "a4e8…9b2f"
}

Shortened: a real export lists every included event and query ID, and every content set, empty ones included.

Content key → catalog controls

See every control mapping
  • Session and approval history

    Never raw transcripts, prompts, or source.

  • Every signal traced

    Each figure resolves to an audit record you can open.

  • Verifier included

    A local CLI rechecks hashes, chains, and redaction offline.

06Scope, stated plainly

What the record proves. What you still attest.

Examiners trust evidence that knows its limits. Every mapping in the catalog separates what AutoDevOps produces from what your organization signs.

AutoDevOps produces

  • Signed, normalized events from every instrumented agent session
  • Policy decisions, each with the rule that fired
  • Human approval decisions with the reviewer — and the rationale, which portal reviews require
  • Hash-chained exports and offline-verifiable evidence packages
  • Sign-off receipts for the thresholds your owners approve

Your organization attests

  • Control ownership and the procedures around each control
  • Model-risk and agent-trust threshold sign-off
  • Validation of the deployment inside your own cloud
  • The audit opinion — that belongs to your auditor

The regulatory mapping is a reference, not a certification. Customer-cloud validation of AutoDevOps itself is still open in every mapping.

07Capture

Collect from the agents you already run.

The public CLI installs telemetry hooks into Claude Code, Cursor, and Codex. There is no proxy in the model call path: agents emit, AutoDevOps records.

AI coding tools and their AutoDevOps capture status
AgentStatus
Claude CodeSupported today
CodexSupported today
CursorSupported today
Copilot CLIRoadmap
Other harnessesVia SDK

Claude Code, Cursor, and Codex are supported today with one install command. Copilot CLI is roadmap. Other harnesses, including Antigravity, Gemini CLI, and OpenCode, emit through the connector helper — they do not ship a hook pack. Product names and logos belong to their owners and indicate capture support, not endorsement.

Install the hooks

terminal
npm install -g @autodevops/verifier
verifier install --harness claude-code --mode telemetry

Telemetry mode records session, prompt-submit, and tool-use events. Prompt text is opt-in. Swap --harness for cursor or codex. The MCP server lets a session attach its spec.

Confirm it’s on the record

terminal
verifier harness doctor --harness claude-code

The doctor command verifies the install receipt and files; events post once your portal credentials are set. Approval gates ship in the customer-cloud enterprise distribution.

Replay the session in the portal

FinCard demo workspace · fictional data

Session replay in the FinCard demo portal: a Claude Code session on payments-risk-service with fourteen events, an approval requested and approved by Security Review, and a verified rerun.
  1. 13:04expectation.capturedDeveloper expected Bedrock-only packaging validation before merge.
  2. 13:06tool.completedClaude Code edited src/risk/review-thresholds.ts.
  3. 13:08approval.requestedSandboxExec packaging command paused as high risk.
  4. 13:11approval.approvedSecurity Review approved after command and diff inspection.
  5. 13:14result.trustedVerifier rerun passed with audit evidence attached.
Events land in session order: the attached spec, the agent’s edits, the action that paused for review, who approved it, and the result. Prompt text appears only when you opt in.

08Deployment

Built for your cloud. Owned by you.

AWS is the validated path today: Bedrock inference on your own credentials and an IAM-scoped worker deployed into your account. One-stack deployment of the portal, ingest, and storage inside your cloud is still in progress. Azure and Google Cloud adapter tracks are roadmap.

Where agents run

  • Developer laptops

    Claude Code · Cursor · Codex hooks

  • CI pipelines

    Verifier hooks on pre-commit and pre-push

Solid: validated on AWS. Dashed: one-stack deployment inside your account is still in progress.

Amazon Web Services

Validated

Bedrock model calls on your own AWS credentials; IAM-scoped workers in private subnets with VPC endpoints and no general egress.

Microsoft Azure & Google Cloud

Roadmap

Adapter code exists for Azure OpenAI, Functions, Vertex AI, and Cloud Run. Not yet customer-validated.

Boundaries

Agent actions only
No keystrokes, no screen capture.
Pre-production scope
Laptop, CI, and cloud workers. Never production traffic.
Your cloud, your record
Designed for BYOC, so prompts and source stay in the customer account.
Allow · Confirm · Block
Risky actions wait for a person when enforce mode is on.

Pick one of 16

Bring one control. Leave with its evidence.

Choose the control your examiners test first. We walk the validated AWS path with you — from agent session to approval to the evidence package that answers it.