SOC 2 audit log for AI agent commands
reflex audit exports one row per decision the gate logged: which agent ran what, where, in which environment tier, what Reflex decided and who approved it. It reads Reflex's local logs, so it is evidence for a change management control, not tamper-proof storage. A webhook can post redacted decisions off the machine.
Install: npx @ursuciprian/reflex setup starts with local rules in shadow mode, no account or key.
Plugins for Claude Code, Codex CLI and opencode: setup guide.
How do I audit AI agent commands for SOC 2 or ISO 27001?
Run reflex audit --since 90d > agent-commands.csv. It writes one row per decision the gate
logged: time, agent, session, working directory, production or not and why, the command with
secrets redacted, the decision, the rule and who approved it (a human in the approval queue,
System 2, or the agent's own prompt). --prod-only, --agent and --format json|jsonl narrow and
shape it. It only reads Reflex's local logs, so it is evidence for a change management control
(SOC 2 CC8.1, ISO 27001 Annex A 8.32), not tamper-proof storage; read-only commands are not logged.
To keep decisions outside the machine, set notify in config.json to an https webhook (Slack or
json): it posts redacted denies, asks or production decisions from a detached process and never
delays the agent.
See: GUIDE: audit log for AI agent commands.
Audit log for AI agent commands (SOC 2)
reflex audit exports one row per decision the gate logged: a document auditors can use as
evidence for change management controls such as SOC 2 CC8.1 and ISO 27001 Annex A 8.32. It shows
which agent ran what, where, in which environment tier, what Reflex decided and who approved it.
It reads the trace (trace.jsonl and its rotated files), the approval queue and the execution
feedback, writes nothing and calls nothing.
reflex audit # the last 7 days, csv on stdout
reflex audit --since 90d --format csv > agent-changes-q3.csv
reflex audit --since 24h --prod-only --format json
reflex audit --agent claude-code --format jsonl
| Column | Meaning |
|---|---|
time | When the gate decided (UTC, ISO 8601). |
agent, session | The agent (claude-code, codex, opencode, pi, hermes, shell) and its session id. |
cwd | The working directory. |
env_tier, env_reason | prod or non-prod, and the marker that made it production (cwd=/infra/envs/prod, aws_profile=production, command: prd). unknown for rows logged by an earlier version. |
command | The command, with secrets redacted when it was logged and again on export. |
decision | What the agent was told: pass, allow, ask or deny. |
judged, mode, source | What the judgment was before the mode applied (in shadow mode a Jev ask is logged, not shown), the mode, and who decided: rule, fast-lane, jev, local, judge (System 2), queue, runaway. |
rule_id, rule | The rule or reason, such as freeze, prod-destroy or team:prod. |
approved_by | Who answered, when that is known: approved in the approval queue (q-1a2b3c4d5e by alice at ...), System 2 approved (confidence 0.9), approved at the agent's prompt (it ran), rejected at the agent's prompt, or no answer recorded. The queue records the account that ran reflex queue approve. |
--format is csv (default), json or jsonl. In csv, a cell that a spreadsheet would run as a
formula (starting with =, +, - or @) starts with a quote. --since takes 7d, 12h or
30m; --prod-only keeps production rows; --agent keeps one agent.
What an auditor should know about it:
- Read-only commands (
ls,git status,kubectl get) are not logged, so they are not in the export. Everything else the gate judged is, whatever the decision. - The trace is a local file owned by the user, and rotates at 50 MB into files the export still reads. It is evidence of what the gate decided, not tamper-proof storage: an agent or a person with shell access as that user can edit it. For retention, export on a schedule, or send decisions to a system you control with the webhook below.
- An approval at the agent's own prompt is inferred: the command ran after an ask. The export says
so in the
approved_bytext.
Decision webhook
Reflex can post decisions to a webhook, for a Slack channel or a log pipeline:
{"notify": {"url": "https://hooks.slack.com/services/T000/B000/XXXX", "on": ["deny", "ask", "prod"], "format": "slack"}}
on lists what is sent: deny and ask (what the agent was told) and prod (any judged
production command, whatever the decision). The default is ["deny"]. format is json (default:
an object with event: "reflex.decision", time, agent, session, cwd, prod, prod_by, command,
decision, judged, mode, source, rule_id and reason; the cwd has your home directory as ~) or slack (a text message, with <, >
and & escaped so a command cannot mention a channel).
A webhook is data egress, so it is held to these rules:
- The URL is
https, orhttponlocalhost,127.0.0.1or[::1], with no user name or password in it. Anything else is a doctor error and nothing is sent. - The URL comes from
config.json. A team policy'snotifyis used only while you trust that exact file (reflex trust .shows its host), because a committed file must not send your decisions somewhere. - The command and the reason are redacted with the same patterns as the trace, and no environment
value is sent:
prod_bynames the marker (aws_profile,kube_context,cwd), not its value. Doctor andreflex policyshow the webhook's host, never its path, since a Slack webhook's path is its secret. - The hook never waits for it. A detached child posts each message once, with a 2 second timeout, no retries and no redirects, while the hook returns its decision. A webhook that is down or slow loses messages; it never delays or changes a decision.
reflex doctor --notify-test sends one dry-run message ("dry_run": true, no command) to each
configured webhook and reports the HTTP status. Doctor sends nothing without that flag.