Reflex

Reflex guide

How the command gate, the prompt injection guard and the autonomous agent profile work, and how to test, tune and roll them out.

Install: npx @ursuciprian/reflex setup starts with local rules in shadow mode, no account or key. Plugins for Claude Code, Codex CLI and opencode: setup guide.

Contents

  1. How a command is decided
  2. Testing
  3. Rolling out: shadow, tune, enforce
  4. Changing behaviour
  5. Metrics
  6. Data handling
  7. Reliability
  8. Safety properties and limits
  9. Injection guard
  10. Runaway guard: stop runaway AI agents
  11. Reflex MCP server: ask before acting (Claude Desktop, Cursor, Cowork)
  12. Autonomous agents
  1. Conditional instructions
  2. Tool router
  3. Model routing
  4. Context layer (pi and oh-my-pi)
  5. Laya (local System 1)
  6. Where this goes next

Local and hosted operation

New setup uses --engine local. Deterministic shell checks and path/keyword instruction matches run without TypeSafe. An unknown command produces ask: enforce requires human review, while shadow logs it and leaves the host's permissions in charge. In the autonomous profile that ask goes to System 2 first (keyless autonomy). Local operation does not reuse cached Jev answers, spawn background classifiers, classify subgoals or make semantic instruction calls. The shared Jev client rejects hosted requests while local; the LiteLLM callback leaves model selection unchanged. A separate LiteLLM container needs the same configuration or REFLEX_ENGINE=local.

--engine laya (experimental) asks Jev's questions of a Laya checkpoint served on this machine: nothing leaves it. Measured far below Jev on every golden set; see Laya.

--engine jev enables the existing hosted behavior below. Older direct hook installations (a Keychain item or an agent record in config.json) retain Jev until an engine is selected; with no settings at all the gate uses local; reflex setup records the choice. Defaults are bundled, with durable user overrides under ~/.config/reflex/tool-gate/; reflex status shows the active policy path.

reflex doctor runs local synthetic decision checks in a disposable state directory. These probes do not count as live activation. reflex status separately reports the last real pre-execution hook event and whether it matches the latest installation and settings. The heartbeat is local operational evidence, not proof against an agent that can modify files. Native dialog behavior and host trust must also be checked in the agent itself.

Use Jev through OpenRouter, Cloudflare or Vercel

Jev does not need a TypeSafe account. The same model, the same questions and the same policy run through any of five providers, and every provider's answers are read into one typed shape before the policy sees them:

ProviderKey (environment, else macOS Keychain item)EndpointModel sent
typesafe (default)TYPESAFE_API_KEY, Keychain typesafe-api-key (or REFLEX_KEYCHAIN_SERVICE)https://api.typesafe.ai/v1/systemonejev-1.13.0
openrouterJEV_OPENROUTER_API_KEY or OPENROUTER_API_KEY (an sk-or- key), Keychain openrouter-api-keyhttps://openrouter.ai/api/alpha/decisionstypesafe/jev-1.13
cloudflareJEV_CLOUDFLARE_API_TOKEN or CLOUDFLARE_API_TOKEN, Keychain cloudflare-api-token, plus CLOUDFLARE_ACCOUNT_IDhttps://api.cloudflare.com/client/v4/accounts/<id>/ai/runtypesafe/jev
vercelJEV_AI_GATEWAY_API_KEY or AI_GATEWAY_API_KEY, Keychain ai-gateway-api-keyhttps://ai-gateway.vercel.sh/typesafe/v1/systemonetypesafe-ai/jev
compatibleJEV_API_KEY, Keychain jev-api-key, plus JEV_API_BASE_URL (the full URL)your URLjev-1.13.0
reflex setup --provider openrouter                      # Jev OpenRouter: uses OPENROUTER_API_KEY or offers the Keychain
reflex setup --provider cloudflare --cloudflare-account 0123456789abcdef0123456789abcdef   # Jev Cloudflare Workers AI
reflex setup --provider vercel                          # Vercel AI Gateway
reflex setup --provider compatible --provider-url https://jev.example.com/v1/systemone
reflex doctor                                           # System 1: Jev via openrouter (openrouter.ai) + policy

Which provider is used: REFLEX_PROVIDER (or JEV_PROVIDER, as in jev-mcp), else provider in ~/.config/reflex/config.json. Without either, TypeSafe when TYPESAFE_API_KEY is set or config.json names a TypeSafe Keychain item (keychain). Otherwise the first of OpenRouter, Cloudflare, Vercel and compatible (jev-mcp's order) whose opt-in variable is set: JEV_OPENROUTER_API_KEY, JEV_CLOUDFLARE_API_TOKEN (with CLOUDFLARE_ACCOUNT_ID), JEV_AI_GATEWAY_API_KEY, JEV_API_KEY (with JEV_API_BASE_URL); else TypeSafe with its Keychain item, as before. OPENROUTER_API_KEY, CLOUDFLARE_API_TOKEN and AI_GATEWAY_API_KEY never choose a provider on their own, since they are often set for other tools (wrangler, for one); they are read once the provider is named. reflex setup without --provider looks in the environment and then in the Keychain, in the same order, and saves the provider it found only in the Keychain, since the hooks read the environment and config.json but never search the Keychain. cloudflare_account_id and provider_url can also be saved in config.json; neither is a secret.

Every provider is a place your data goes; see Data handling. The provider layer is adapted from jev-mcp by Joey Kudish (MIT).

How a command is decided

Every shell command an agent wants to run reaches gate.mjs through that agent's pre-execution hook (see the table in the README). Each adapter turns the agent's event into the same call, {agent, command, cwd, session_id, call_id, intent}, and gets back {effective, decision, reason, source}. The first step that reaches a decision wins:

  1. Read-only: a command that only reads passes. By default ("readonly": "legacy") this is readOnlyLegacy(), a shell parser that passes chains, loops, $(…) and many more spellings, and runs before the rules except secret-read, secret-file-read and ssh-local-command. Five security reviews found about 50 ways to fool it. "readonly": "simple" in config.json (or REFLEX_READONLY=simple) turns on the stricter check instead: a small allowlist, not a shell parser. Simple commands and pipelines of them, joined by ;, &&, || or a newline, with cd <path> between them, each program and each flag on an explicit list (ls, cat, grep, git status, git log, kubectl get, aws … describe-*, jq, gh pr view, ssh host '<read-only>', see Read-only allowlist). Anything else is not read-only and goes on to the next steps: &, a redirect other than 2>/dev/null and 2>&1, $, globs, braces, subshells, an unknown program, subcommand or flag. The simple check runs after the rules, the tamper check and the local scripts below, just before the fast lane, so no rule is ever skipped because a command looked like a read. It is not the default yet because it still sends more commands to a human (68.4 per 100 against 54.8 on the author's last 7 days of Claude Code sessions). → pass, not logged. In a session that read a suspected prompt injection a read-only ssh is still egress and asks.

  2. Rules (rules.json): regular expressions over the command plus its context (cwd=, aws_profile=, kube_context=, tf_workspace=, git_branch=). A rule fires when all of its patterns match. Rules are enforced in shadow and enforce modes, with off disabling the entire gate. Shipped rules: rm-root, prod-destroy, force-push-main, push-mirror (deny); tamper, destroy, force-push-unknown-branch (a force push of HEAD or no ref where the branch is unknown), secret-exfil (for scripts), ssh-local-command (-o ProxyCommand|LocalCommand|Match …, a -J hop that is an option), and, checked even before read-only detection, secret-read (the API key, secret stores) and secret-file-read (~/.ssh/id_* but not .pub, ~/.aws/credentials, .netrc, .pgpass, .env / .env.* files but not .env.example and other templates, kubectl get secret(s)) (ask). It fires only when the file is an argument of a command that reads, copies or sends it (cat, less, head/tail, grep/rg/ag, jq, sed/awk, cp/scp/rsync as the source, base64, xxd, strings, od, nl, sort, uniq, cut, paste, fold, column, diff, comm, cmp, tac, rev, open, source/., nc, tar/zip, curl -d/-F/-T/--data*, a routed mcp call), at any command position (after ;, &&, |, inside $(…), backticks, bash -c '…', ssh host '…', ssh host cmd …), or is redirected in (< ~/.aws/credentials). A commit message, echo, or cp .env.example .env that only names the file passes. Known over-match: a grep whose search pattern is .env (grep -rn '.env' src/) asks. Any mutating command that touches the Reflex checkout, its setup files or its logs is also an ask, wherever the repo was cloned. Shell text that is data is not a command: the shell rules ("shell": true) skip a command that only reads and writes notes (echo '… rm -rf / …' >> MEMORY.md, reflex check '…'), and the program of a command that is only python3 - <<'EOF' … EOF (quoted delimiter, no wrapper or flags) when nothing in it could run, load, write or delete anything; the engine judges those instead. tamper ("writes": true) reads what a command changes, so jq . ~/.claude/settings.json > /tmp/s.json is a read, and a redirect, mv, cp, tee, sponge, sed -i or yq -i onto the file is not. Local scripts. bash deploy.sh says nothing about what it does, so Reflex reads the local file a command runs. It recognises:

    • bash, sh and zsh x.sh (quoted paths, < x.sh and options included);
    • source x.sh, ./x and scripts/x;
    • python3, node, tsx, npx tsx, bun, ruby, perl and php with a script, including one run through a path such as .venv/bin/python;
    • make <target>: that target's recipe, its direct prerequisites' recipes, and $(MAKE) calls;
    • npm, pnpm and yarn scripts (run x, yarn x, test, and install lifecycle scripts), with their pre/post hooks.

    It finds these behind subshells, $(…), if/then, env, sudo, nice, cd and line continuations. What a shell script, recipe or package script runs in turn is followed one more level.

    Rules with "applies_to" including "script" read up to 256 KB of each file. Those rules are rm-root, prod-destroy, force-push-main, force-push-unknown-branch, push-mirror, destroy and tamper, plus a check for the Reflex checkout and logs. The file is read line by line, with the script's own VAR=value assignments expanded and whole-line # and // comments dropped, so a word on one line cannot combine with a verb on another. The script-only secret-exfil rule (whole_script) asks when a script copies ~/.ssh, ~/.aws, ~/.kube, ~/.gnupg or a .env file and also talks to the network. Using a key with -i does not count.

    A hit names the script: reflex (rule): recursive delete of / or home (in /repo/scripts/reset.sh). This runs before the fast lane, so npm test is only as safe as the test script. The command line sh -c '…' runs is inspected like the command itself; bash -n is not a read. A compiled program under a system directory (/usr, /bin, /opt/homebrew, …) is judged by its command. A credentials file that gets sourced (.env, .netrc, keys) is scanned by the rules but never sent to Jev.

    Some code is marked unseen, so the command is never allowed:

    • a named script that cannot be read (missing, or a package script the package.json lacks, which yarn, pnpm and bun run as a bin);
    • a package fetched or installed (npx, pnpm dlx, bunx, uvx, npm install x, pip install), a module or preload by name (python -m, node -r / --import, NODE_OPTIONS, BASH_ENV, PYTHONPATH, LD_PRELOAD), a task runner (just, task, gradle …), go generate / go run, and find -exec of a script;
    • the local modules a Python or JavaScript script imports (import helper, require('./x'));
    • a compiled program outside the system directories (built in the repo, downloaded);
    • a workspace or filter script (npm -w, pnpm --filter, yarn workspace);
    • code piped into a shell (curl … | bash);
    • a file an earlier step of the same command wrote (curl -o x.sh … && bash x.sh);
    • a third level of nesting;
    • anything past 8 scripts.

    Lines over 2,000 characters (minified bundles) are left out of the rules, and the script then counts as partly seen. If the scan takes more than 1.5 s, the command gets an ask instead of risking the hook's timeout.

  3. Fast lane (rules.json → pass): known-safe steps: builds, tests, mkdir, git add/commit, pushing a non-main branch. A command passes when every segment is read-only or matches a fast-lane pattern. → pass, logged. Workspace allowlist (workspace.mjs, keyless): a few exact in-tree shapes (mkdir -p, touch, sed -i with one s/// on a tracked file, cp or mv of a tracked file to a new path) pass with a checkpoint taken first. See Human-last.

  4. Jev: the command (secrets redacted), its working directory, the environment context and the text the agent wrote right before this command and its last five commands (from the session transcript; if the command is not in the transcript yet, no intent is sent rather than an older one) are sent to TypeSafe with the six questions in questions.json. When the command runs a local script, its redacted excerpt goes along as call.script = {path, excerpt} and the question context tells Jev to judge the script, not its name (node clean.mjs that deletes $HOME is an ask, not a pass). Answers are cached for 24 h per (command, cwd, environment, question-set version, model, script content), so an edited script is judged again.

  5. Policy (policy.json): ordered gates over the answers; the first that fires wins, otherwise default_outcome (pass). If Jev fails or times out, the policy's fallback (ask) applies. The last gate, allow, marks clearly safe commands; what that means depends on REFLEX_ALLOW (see Calibrated allow).

What happens with the decision depends on the mode:

ModeRulesJev + policy
offnothingnothing
shadowenforcedruns in a detached background process; logged, never shown to the agent
enforceenforcedask → a human confirms (how depends on the agent, see the README table); deny → the command is blocked and the agent sees why

pass is never turned into an approval: the gate stays silent and the agent's own permission settings decide. Only an allow with REFLEX_ALLOW=on in enforce mode approves anything. Where an agent combines several hooks (Claude Code, Codex, Hermes), the most restrictive decision wins, so Reflex composes with the hooks you already run.

Custom integrations use the same contract from any language:

echo '{"agent":"my-bot","command":"terraform apply","cwd":"/infra/prod"}' | node gate.mjs --decide
# {"effective":"ask","decision":"ask","reason":"reflex (jev): changes production","source":"jev","policy":"tool-gate-v2"}
echo '{"agent":"my-bot","call_id":"42","exit_code":0}' | node gate.mjs --record

Subgoal dedup

Before an agent spawns a subagent, the adapter sends the subgoal instead of a command: {agent, subgoal, session_id, call_id, cwd}, or subgoals: [...] for a batch. A call that also carries a command is judged as a command.

AgentTool (hook)Subgoal
Claude CodeAgent / Task (PreToolUse, matcher Bash|Task|Agent)agent: <subagent_type>, description, prompt
Codex CLIspawn_agent (PreToolUse, matcher ^(Bash|spawn_agent)$, Codex 0.155+)agent: <agent_type>, task name, message (or its text items)
oh-my-pitask (tool_call)per task: agent: <agent>, task, the batch's shared context cut to 200 characters
opencodetask (tool.execute.before)agent: <subagent_type>, description, prompt
Hermesdelegate_task (pre_tool_call, its own entry without fail_closed)per task: goal, context cut to 300 characters

pi has no subagents. Codex's SubagentStart hook is not used: its input has no task text and it cannot block. Resumes (Claude Code resume, opencode task_id) and Hermes control actions (list, steer, stop) are not checked. A subagent that spawns its own subagents (Claude Code, Codex) is compared only with its own earlier ones.

Reflex keeps the subgoals of each session in subgoals.jsonl in the data directory, and asks Jev one choice question per new subgoal (setup/tool-gate/subgoals.json). The options are the last 20 earlier subgoals of the same agent and session that count, plus none:

A spawn Reflex denied is marked dropped at once; one the user rejected, another hook blocked or that failed (a denied or failed record) is never offered either, since it has no result to reuse. Long options keep their first 400 and last 200 characters. The question is "does this repeat one of them: same task, same scope, same kind of answer?" Follow-ups, other parts of the problem, reviews of earlier work and retries that say why the first attempt failed count as none.

If the chosen option's probability is at least duplicateAt (0.6), that subgoal is a duplicate. A single spawn is denied, with a reason that names the earlier subgoal and tells the agent to reuse its result:

reflex (jev): duplicates a subgoal already launched in this session at 22:24 UTC (p 0.67):
"agent: Explore\nFind auth flow\nExplain how login works end to end …". Reuse that result instead of starting it again

A batch loses only its duplicate tasks where the tool input can be trimmed: oh-my-pi runs the task call with the rest and appends to its result which tasks were not started and why. Hermes cannot be told which tasks a trimmed call dropped, so a Hermes batch with a duplicate is blocked with the list ("Start the others again without task 2"), and all its tasks are dropped, so the resend passes. A batch whose every task is a duplicate is denied.

The same text again (ignoring case, spacing and a batch's shared context line) is a duplicate without asking Jev: Jev scores an option identical to the new subgoal low (0.2 to 0.3), apparently reading it as the new subgoal itself. On hand-made pairs with subgoals-v3, reworded duplicates scored 0.67 to 0.78 on the matching option, and ten legitimate follow-ups (tests for the same code, another part, a review, a retry that says why, a narrower scope) scored at most 0.05.

ponytail: each check reads the last 2 MB of subgoals.jsonl and feedback.jsonl and takes no lock (each batch is one append).

Testing

Four layers, cheapest first.

1. Offline self-checks, on every change

npm test

npm test also runs node context.mjs --selfcheck (the context layer, against a fake Jev on localhost). About 60 gate checks, no network: read-only detection (including bypass attempts such as rtk proxy rm, ssh h 'echo' '; rm -rf /', $(security find-generic-password …)), redaction, every shipped rule, the fast lane, every policy gate, and the whole path for non-Jev commands. Add an assertion whenever you change readOnly(), a rule or a gate. node guard.mjs --selfcheck covers the injection guard: every detector (most with a benign twin), the policy, rewriting, log redaction, each adapter's output shape, credential prompts, and taint making the gate stricter.

2. One command by hand

node gate.mjs --check "kubectl --context prod-eu scale deploy/api --replicas=0 -n web"
node gate.mjs --check "helm upgrade api ./chart --kube-context prod-eu" --intent "Bumping the chart in dev"
node gate.mjs --check "terraform apply" --cwd ~/infra/envs/staging

Prints decision, rule, source, Jev's answers and the environment context it saw. The environment comes from your shell (AWS_PROFILE, current kube context, .terraform/environment, git branch), so run it from the same shell you start the agent from.

3. Golden set, on every change to questions, policy or rules

npm run eval                     # all cases
node eval.mjs --only terraform   # a subset

setup/tool-gate/golden.json holds labelled commands with the outcome we want (pass, ask, deny, or a list when more than one is acceptable). The eval runs them through the real gate (cache off) and reports:

A case can also carry "allow": true (a clearly safe command that should be allow-eligible; reported as stiff when it is not, never a failure) or "allow": false (must never be auto-allowed; a MISS if it is). An allow counts as pass for expect.

Cases with "cwd": "$FIXTURES" run in a temporary copy of setup/tool-gate/fixtures/, the scripts, Makefile and package.json those commands run (each guarded so it exits if run by hand).

The set has 97 cases, 6 of them allow: true. npm run eval prints the pass / MISS / over counts for the rules, questions and policy you have now, and saves them to ~/.local/state/reflex/eval-*.json; Jev's answers vary between runs, so read a run as a sample and only treat a MISS as a blocker. CI runs the offline self-checks (npm test); the live eval needs TYPESAFE_API_KEY and spends tokens, so CI does not run it.

Grow the golden set from real traffic. Every surprising decision in the trace becomes a case.

The injection guard has its own set, setup/injection/golden.json, run by npm run eval-injection: precision and recall of warn / block against the labels, and exit 1 on a missed high-severity injection (see Injection guard).

4. Live, in shadow mode

Install in shadow mode (see SETUP), work normally for a few days, then:

node report.mjs                  # counts by source, decision, mode; latency; tokens
node report.mjs --list ask       # every command that would have been asked
node report.mjs --list deny

Read the ask and deny lists. Each wrong one is either a golden-set case to add and a threshold to move, or a rule or fast-lane pattern to add.

To test enforce behaviour without switching your whole setup, start one session with REFLEX_MODE=enforce claude and ask the agent to run something the eval shows as ask.

Rolling out: shadow, tune, enforce

  1. Shadow, about a week. Deterministic rules still enforce; other decisions are logged only.

  2. Tune. Try a candidate policy against everything recorded, without spending tokens:

    cp ~/.config/reflex/tool-gate/policy.json /tmp/candidate.json   # edit thresholds
    node report.mjs --policy /tmp/candidate.json         # "N of M Jev decisions change", pass -> ask: …

    Keep what reduces wrong asks without adding misses, then npm run eval.

  3. Enforce with node install.mjs --mode enforce.

  4. Watch the ask outcomes. In enforce mode report.mjs scores each emitted ask as approved (the command ran), rejected (explicit denial feedback), pending or unknown. Missing feedback after ten minutes is unknown and excluded from calibration. An ask that is nearly always approved is friction: tune it. A deny someone keeps working around is a missing fast-lane pattern.

Replay and bench

Before a week of shadow mode, reflex replay shows what the gate would have done with the shell commands your agents already ran. It reads the local session transcripts, runs each command through the gate's checks and prints counts. It never executes a command, and it writes nothing: the mode is shadow, System 2, the queue and checkpoints are off, and the answer cache is neither read nor written. The trace, the answer cache and the approval queue are untouched.

reflex replay claude --since 7d                  # local engine, free
reflex replay codex --since 30d --project ~/work/api
reflex replay all --json                         # claude, codex, opencode and pi
reflex replay claude --engine jev                # prints an estimate and what would be sent, sends nothing
reflex replay claude --engine jev --yes          # the real Jev calls; tokens and spend from usage
reflex bench                                     # precheck p50/p95, local only
reflex bench --engine jev                        # plus one Jev call per fixed command the rules leave open
AgentWhereWhat counts as a command
Claude Code~/.claude/projects/**/*.jsonl (subagents included)Bash tool_use
Codex$CODEX_HOME/sessions/**/*.jsonlCommandExecution items, plus exec_command, shell, shell_command and local_shell_call calls whose command no item has (older rollouts have no items)
opencode~/.local/share/opencode/opencode.dbbash tool parts; needs node:sqlite (Node 22.13+), else skipped with a note
pi~/.pi/agent/sessions/**/*.jsonlbash toolCall

A call id counts once, so a resumed session that copies earlier calls does not double them. Files not modified since the window opened are not read. --since takes 7d, 12h or 30m (default 7d), --project keeps commands whose cwd is under a directory, --limit N keeps the most recent N.

The output: commands; passes by the read-only list and the fast lane; rule asks and denies; what the engine decided for the rest (engine local: an ask for each, as in the hook); asks per 100 commands, which is what reaches a human in the supervised profile in enforce mode; and, for the autonomous profile, how many of those asks would go to System 2 and how many stay with a human (the always-human class, or every ask when config.json has no System 2 backend). Then the rules that fired most, a sample of denied and asked commands with credentials masked by redact() (and a password after --password, --token or login -p), and the spend. With engine local that is an estimate for Jev: distinct uncovered commands (a repeat in the same directory is asked once, as the answer cache would), about 2k input tokens each, at $0.04 per million (TypeSafe publishes no price list; this matches a public third-party measurement; set REFLEX_JEV_USD_PER_MTOK to your contract price). With --engine jev --yes it is the measured tokens from usage, the spend at that price, and Jev's latency p50 and p95. --yes sends TypeSafe each distinct uncovered command (credentials masked), its directory and an excerpt of any local script it runs; the estimate says so before anything is sent.

What replay cannot see: the environment at the time (AWS profile, kube context, branch) is not in a transcript, so it is left empty and rules keyed on it do not fire; there is no stated intent, so Jev never allows; and a local script the command runs is read as it is now, not as it was.

reflex bench times the local precheck over twelve fixed commands, 20 rounds, and prints p50 and p95. It is local unless --engine jev or --engine laya is given; then it also sends the fixed commands the local rules leave open, one call each, from an empty temporary directory (so no file of yours is read or sent), and prints latency, input tokens and the cost per 1,000 calls. Both take --json.

Suggest: fewer permission prompts

Every command the local rules do not cover is an ask, and in enforce mode an ask is a human in the loop. Most of those are commands only a human should approve. Some are the same make lint or pnpm typecheck a hundred times. reflex suggest finds the second kind in your own history and proposes fast-lane entries for them, so autonomous coding agents stop at fewer permission prompts without widening anything else. It complements the Claude Code permissions allowlist and Codex approvals: those decide by command prefix in the agent; this runs inside Reflex, after its rules, and reads the scripts a command runs.

reflex suggest claude --since 30d            # what it would add and the effect; writes nothing
reflex suggest codex --min 5 --json          # at least 5 runs per suggestion
reflex suggest all --project ~/work/api
reflex suggest claude --write                # show the lines, ask, then append to fastlane.json
reflex suggest claude --write --yes          # the same without a terminal (scripts)

It reads the transcripts reflex replay reads (same agents, --since default 30d, --project, --min default 3) and runs nothing. It takes the commands the gate left to the engine (not the read-only list, the fast lane or a rule), cuts each into its shell segments, and turns each segment that kept it out of the fast lane into a template. A template keeps every word literal except a number (\d+) and, for a test runner, linter or type checker, a repository-relative path:

SeenSuggested pattern
make lint 2>&1 | tail -5^make\s+lint$
npm run typecheck^npm\s+run\s+typecheck$
ruff check src/app.py, ruff check tests/test_api.py^ruff\s+check\s+<repository-relative path>$
docker compose images^docker\s+compose\s+images$

A template is only suggested when all of these hold:

Each suggestion prints the pattern, the project it is scoped to (the repository root of the runs), the count, up to three samples with credentials masked, and why it is safe. Then the effect: the same classification rerun with the suggestions added, as asks per 100 commands before and after (supervised: what reaches a human in enforce mode; autonomous: what stays with a human). When nothing qualifies it says what keeps asking instead, by its first two words and why it was left alone.

--write shows the entries it would append to ~/.config/reflex/fastlane.json (or $XDG_CONFIG_HOME/reflex/fastlane.json) and asks on the terminal; without one it needs --yes. It writes only a file that validates, atomically, mode 0600. --write changes Reflex's own configuration, so the tamper rule asks a human whenever an agent runs reflex suggest --write, as it does for any edit of ~/.config/reflex. That is by design: an agent cannot widen its own allow list.

The file can also be edited by hand:

{"version": 1, "entries": [
  {"pattern": "^make\\s+lint$", "cwd": "/Users/me/work/api", "note": "reflex suggest 2026-09-26: 14 runs"}
]}

The hook reads it after the bundled fast lane (precheck in gate.mjs), so it only ever turns an engine decision into a pass. A deny, a secret read, a tamper ask, a rule over a script and the always-human class (every rule in escalation.json, production contexts included) still decide first, and the denylist, the cd check and the script check above apply to every command at run time, not only when the entry was suggested: a script that later starts sending data stops passing. An entry applies in its cwd and below. Like the bundled fast lane's pytest and go test, a test runner or linter entry trusts the repository's own test code and configuration. The file is strict: an entry must have an absolute project cwd (not / or your home) and an anchored pattern from ^ to $ with no . wildcard, negated class, \S, \W, \D, \x, \u, \p, \c, space in a class, range other than a-z, A-Z and 0-9, repeated group across words, lookaround or backreference. One invalid entry and the whole file is ignored: nothing is widened on a parse error, and reflex doctor prints a warning that names the problem.

reflex learn: fewer prompts from your own approvals

reflex suggest looks at what your agents ran. reflex learn looks at what you approved. It is the way to reduce permission prompts in Claude Code, Codex and the other agents Reflex gates by learning from approvals: a command shape you said yes to often enough, and never said no to, becomes a project-scoped fast-lane entry, after you confirm it. Nothing is learned automatically and nothing is written without a person at a terminal.

reflex learn                           # proposals, why, and the effect on your history; writes nothing
reflex learn --since 60d --min 5 --json
reflex learn --write                   # show the entries, ask on the terminal, append to fastlane.json
reflex learn --write --yes             # the same without a terminal (scripts)
reflex learn --list                    # learned entries: where they came from, uses, last use
reflex learn --forget l-1a2b3c4d       # remove one learned entry
reflex learn --prune                   # remove learned entries unused for 60 days
reflex learn --team                    # a .reflex/policy.json fastlane snippet instead (reflex trust applies it)

What counts as an answer. Only a person's:

A System 1 allow, a System 2 approve or deny and a command the fast lane or a rule decided are never answers, and neither are Hermes' own approval modes or Claude Code outside its default permission mode (bypassPermissions, acceptEdits, auto). The same command approved again in one session is one answer, so an "always allow" or a retry loop does not add up. Refusals count from any time, approvals only inside --since. Everything is read from Reflex's own logs (the queue, trace.jsonl, feedback.jsonl), --since 30 days by default.

When a shape is proposed. The approved commands are grouped with reflex suggest's templates (literal words, \d+ for numbers, a repository-relative path for test runners and linters) and put through the same proof, so every check in Suggest applies: the denylist, no quotes, expansions, redirects or paths out of the repository, the Makefile, npm script and interpreter checks, the always-human class and the probes. On top of that, all of these hold:

A learned entry is pinned: pin holds the sha256 of each local script it runs (the package.json script, the whole Makefile, the shell file) as it was when you confirmed it. If one of them changes, the entry stops passing that command until you learn it again. A script body that runs inline interpreter code (node -e, python -c, sh -c, deno eval) never qualifies for the user fast lane, and one that calls Reflex (reflex learn --write, reflex queue approve) is a tamper ask even in the bundled fast lane.

Destructive, production, secret, tamper, always-human, change freeze and MCP commands are never learned, however many times you approved them: a rule decides them before the fast lane, or the denylist refuses the template. An approval given during a change freeze does not count.

Each proposal shows its id, the pattern, the project it is scoped to, how many times it was approved, in how many sessions and by which route, denied 0, up to three samples with credentials masked, the exact entry it would add, and why it is safe. Shapes held back are listed with the reason.

The effect. reflex learn reruns the reflex replay classification over every command in your transcripts (all agents, read only) with and without the proposals, and prints humans per 100 commands before and after. On the maintainer's own history (32,001 commands over 30 days, keyless) that was 50.4 before and 50.4 after: no proposals. The asks there are python3 -c, sed -i, curl, gh api, node -e and git, which are exactly the shapes that must never be learned, and the logs held few human answers. Expect it to help when your asks are the same build, test and lint commands.

Provenance and decay. A learned entry in ~/.config/reflex/fastlane.json carries id, learned_at and learned_from (approved and denied counts, sessions, first and last approval, and the routes). --list shows each with its uses and last use (fast-lane passes in the trace). reflex doctor flags a learned entry unused for 60 days; --prune removes those. --forget and --prune never touch an entry you wrote by hand.

Who can write. --write shows the lines and asks on the terminal (not stdin, so an agent's pipe cannot answer); without one it needs --yes. When an agent runs reflex learn --write, --forget or --prune, the tamper rule asks a human, as it does for reflex suggest --write, and so does reflex learn or reflex suggest with its flags hidden in a variable, "$@", xargs, eval or a function. reflex status adds a passive line, "N commands you approved 3+ times could stop asking: run reflex learn", and never writes. --team prints a fastlane list per repository for .reflex/policy.json; teammates apply it with reflex trust, as for any team fast lane. reflex learn is part of the CLI only; the Claude Code plugin does not ship it.

What it trusts. The answers come from Reflex's own logs (trace.jsonl, feedback.jsonl and the queue items in the data directory). There is no second record kept somewhere a shell write cannot reach, so the logs are trusted only as far as the tamper check protects them. It asks a human for a shell write into the data or config directory by path, by a relative path after cd into one or its parent, from a working directory there, through a glob or a .. climb, from a local script, and for a Reflex file under a variable the command does not set; a file tool write there asks too. It does not follow a bare variable (echo x >> $F) or an unknown program that writes there without naming it. A write you approve yourself, or one made outside the agent, can still change what reflex learn counts, and --write shows every entry before it adds it.

Calibrated allow

Autonomous agents stop at every prompt their own permissions require, even for npx prettier --write src/. The policy's last gate, allow, marks commands Jev judges clearly safe:

blast <= $allowBlastMax (1.2) and blast.confidence >= $allowConfidence (0.7)
and mutates, exfil, injection < $allowRiskBelow (0.3) and on_task >= $allowOnTask (0.5)
and env == 'local'

It is the last gate, so every deny and ask gate wins over it, and a missing answer makes it false. REFLEX_ALLOW (or node install.mjs --allow …) decides what an allow does:

REFLEX_ALLOWLogged decisionEffect
off (default)passnone: the gate only tightens
shadowwould_allownone; report.mjs measures it
onallow in enforce mode, would_allow otherwisethe agent's own prompt is skipped (see the README table per agent)

Never allowed, whatever the answers: anything a rule decided (including tamper and secret reads), read-only and fast-lane commands (they stay pass, so your permission allowlist still governs them), Jev errors and incomplete answers (the ask fallback), cached answers (they have lost on_task), commands without a stated intent (on_task defaults to yes then), commands that redaction changed (a redacted --token "$(…)" could hide a payload), commands that run code Jev did not see in full (a script that is redacted, over 16 KB, unreadable, a credentials file such as .env, or imports local modules; a make target, whose variables are not expanded; a compiled program outside the system directories; python -m, node -r / --import, NODE_OPTIONS, BASH_ENV, PYTHONPATH; npx / dlx / uvx, package installs, task runners, go generate / go run), commands run from the home directory or / (where "inside the working directory" means everything), and a policy whose default_outcome is allow (only the allow gate allows). In Claude Code, a command retried outside the sandbox (dangerouslyDisableSandbox) and any command in plan mode keep their prompt. The trace logs why as low risk (not allowed: …). export REFLEX_ALLOW=… in a command is a tamper ask. In Claude Code an allow skips the prompt but its deny and ask permission rules still apply.

Calibrate from your own approvals. Run with REFLEX_ALLOW=shadow in enforce mode for a while. node report.mjs then shows, for the Jev-judged commands a human ruled on (asks it emitted in enforce mode, and in Claude Code the commands allow would have let through, logged would_allow or allowed on replay), how often you approved them (approved = the command ran), by blast and by confidence bucket, and recommends thresholds (illustrative output):

  calibration  64 labelled (asks in enforce mode + would-be allows; approved = it ran)
    by blast       0-0.5 12/12 (100%) · 0.5-1 30/31 (97%) · 1-1.5 9/11 (82%) · …
    recommend      of 41 with blast <= 1 and confidence >= 0.8, you approved 98% -> allowBlastMax 1, allowConfidence 0.8

It recommends the band covering the most commands you approved at least 95% of (10 or more labelled), the tightest among equals, and says "not enough data" otherwise. Other agents have no prompt for a pass, so their would-be allows are not labels, and neither are Claude Code's in a permission mode other than default (the trace logs permission_mode); a command its allowlist let through still counts as approved, which flatters the rate a little. node report.mjs --calibration prints the expected calibration error of blast and mutates read as approval probabilities (1 − blast/3, 1 − mutates) against your approvals, per bin; it needs 20 labelled commands. Move the allow* params in policy.json, check with node report.mjs --policy, then switch to REFLEX_ALLOW=on.

Changing behaviour

Everything lives in setup/tool-gate/. Bump the file's version on every change; it is stamped on every log line, so any decision can be traced back to the exact rules and policy that made it.

Is it a fact? Write a rule or a fast-lane pattern. An account ID, a branch name, a --force flag, a path under envs/prod: code decides these exactly and for free.

Is it a judgment? Adjust a policy threshold or gate. Gates look like:

{"id": "prod", "label": "Production mutation?", "enabled_when": "$prodAlwaysAsks",
 "test": "env == 'production' and mutates >= 0.5", "outcome": "ask",
 "detail": "{env}, mutates {mutates}", "rule": "changes production"}

Conditions read answers by question id, <id>.confidence, params and flags ($name), numbers and quoted strings, joined with and / or. A missing answer makes a condition false, never an error.

Is Jev judging the wrong thing? Reword the question. Write each question as one narrow judgment with enough context to answer it, and put examples of what does not count in the instructions; this is how the exfil question stopped firing on ordinary aws and terraform calls. The context field in questions.json describes your environment: adapt the definition of production to your naming. See TypeSafe's guidance on state, primitives and confidence.

Is the read-only list missing a command your team runs constantly? See Read-only allowlist: an entry in READ_ONLY_SIMPLE for everyone, or a trusted team fast lane entry for one repository.

Read-only allowlist

The read-only pass (readOnlySimple() in readonly.mjs) answers one question: does this command only read? It answers yes only when it can see the whole command, and no otherwise. There is no shell parser to fool: a command it does not fully recognise is judged by the rules and the engine like any other, which costs a prompt or a Jev call, never safety.

It is opt-in: "readonly": "simple" in config.json, or REFLEX_READONLY=simple. The default is still readOnlyLegacy() (see How a command is decided).

The shape. Simple commands, or pipelines of them joined by |, and these joined by ;, &&, || or a newline: ls; pwd, git log -1 && git diff --stat, cd repo && git status. Each part must be read-only on its own. Words are letters, digits and _ @ % + = : , . / -, a ~ only at the start of a word before / (zsh with EXTENDED_GLOB reads ^ and a later ~ as glob operators, so HEAD~3 must be quoted), single-quoted text, or double-quoted text without $, backticks, ! or an escaping backslash. The only redirects are 2>/dev/null and 2>&1. Not read-only: & (background), |&, an empty part (ls;;pwd), any other redirect, $ in any form (variables, $(…), $'…'), backticks, ( ) { }, unquoted globs (* ? [ ]), heredocs, # and backslash escapes (the one exception is \' between single-quoted parts, which /reflex:check writes). A word is judged as the shell passes it, quotes removed, so '-'X is the flag -X. A word naming a secret file or directory (a private key, ~/.ssh, ~/.aws, .kube/config, .env, .pgpass, .git-credentials, .vault-token, gh's hosts.yml, auth.json, /proc/…/environ, also as HEAD:.env) is never read-only, even before the secret-file-read rule sees it. Before the program, only unquoted AWS_PROFILE=, AWS_REGION= and AWS_DEFAULT_REGION= with a plain value are allowed, and rtk proxy is transparent (not rtk grep and the other rtk subcommands, which re-implement tools).

cd. A part of its own may be cd and one literal path: no - or + (the previous or a stacked directory), no option, no $, backtick, backslash, glob or brace character, and ~ only as the home directory (~user is refused). It is not a part of a pipeline (cd x | ls), and a command of only cd is not read-only. A word after a cd is also checked as a path from that directory, so cd ~/.ssh && ls and cd ~ && cat .ssh/id_rsa are secret reads, not read-only. The tamper check still sees the cd and runs first: cd .. && sed -i … gate.mjs from inside the checkout asks. A CDPATH in the agent's environment (not in the command, where tamper catches it) can send a relative cd elsewhere; each part after it is still only a read. A cd hook in your own shell setup (a zsh chpwd function, or a tool that wraps cd such as rvm or autoenv) runs when the agent's shell changes directory, and may source a file from the new one: keep such hooks out of the shell your agent uses (the legacy mode passes cd chains too). A fast lane pattern applies only to a command that is one pipeline: cd x && npm test and ls; go test ./... are judged, not passed.

The programs. Each has its own flag list; a flag not on it means not read-only. Flags are allowlisted, never denylisted, and the list leaves out every flag that writes, runs a program or reads a secret.

ProgramAllowedLeft out on purpose
ls, pwd, whoami, uname, id, hostname, uptime, nproc, sw_vers, which, type, sleeptheir listing flagshostname NAME (sets it)
cat, head, tail, wc, nl, fold, rev, tac, od, strings, cut, tr, paste, column, comm, cmp, diff, basename, dirname, realpath, readlink, stat, du, df, shasum, sha256sum, md5, md5sum, echo, printfdisplay and selection flags, pathsprintf -v
datedisplay flags, one +FORMAT-s, --set, an operand without + or a second one (sets the clock)
grep, egrep, fgrep, rgsearch and output flags, -f --f FILE, rg --pre, --pre-glob, --hostname-bin, -z (runs a program)
sort, uniq, tree, fileordering and display flagssort -o, -T, --compress-program; uniq IN OUT; tree -o, -R, -H; file -C, -m, -z
sed-n with line-number p commands (10,20p), Nqevery other program (w, e, s///w, -i)
findtests (-name, -type, -mtime, …) and -print, -print0, -printf, -ls, -prune-exec, -execdir, -ok, -delete, -fprint*, -fls
ps, pgrep, nvidia-smilisting and query flagsps -E and BSD e (other processes' environment), pkill, nvidia-smi setters and -f
gitstatus, log, show, diff, shortlog, rev-list, branch (list forms), tag -l, rev-parse, ls-files, ls-tree, blame, describe, merge-base, show-ref, for-each-ref, remote [-v], remote get-url, stash list/show, worktree list, ls-remote; -C DIR, --no-pager-c, --output, --ext-diff, --textconv, --show-signature and %G… or %(signature…) in a format (they run gpg), --upload-pack, ls-remote to a host other than GitHub, a branch or tag name that creates one
kubectlget, describe with namespace, context, selector and -o format flags--kubeconfig, --token, --server, --as, --raw, any Secret
aws<service> describe-*, list-*, get-*; s3 ls; configure list; CLI options --profile, --region, --output, --query, --no-cli-pager and similar--endpoint-url, --cli-input-*, --debug and any prefix of them (the CLI expands --endpoint, --debu), file:// values, --with-decryption, --include-value(s), operations that return a secret, token, password, credential, stream key or access, glue get-connection(s), streaming operations that write a file (s3api get-object, get-export, …), s3 cp
jqoutput flags, --arg, --argjson-f, --from-file, --rawfile, --slurpfile, -L, --, a filter that reads $ENV or env or imports a module
terraformversion, fmt -check or fmt -write=false (not -diff, which runs diff from PATH)everything that starts a provider binary (plan, show, validate, state, output, …), fmt that writes
dockerps, images, logs--config, inspect (prints environment variables), exec
ghpr, issue, run, repo, release view and list forms, auth status, api ENDPOINT as a GET--web, api -X, -f, -F, --input, -H, graphql, a full URL, a --jq that reads $ENV or env (gh's jq sees GH_TOKEN), auth token, --show-token
sshssh [-nTqt46C] [-p N] [-o ConnectTimeout=…, BatchMode=…, StrictHostKeyChecking=…, ServerAlive*=…] [user@]host <read-only> as the first command of its pipeline-J, -F, -i, ProxyCommand and every other option, a pipe into ssh, a login, an unquoted ~, a backslash in the remote text (a fish login shell reads \' inside quotes differently)
node, npm, python3, git, docker, aws, terraform, kubectl, helm, jq, rg, gh, uv, brew, make--version alonego, cargo, pnpm, yarn (may fetch and run a toolchain the project names)

The remote command of ssh is the words after the host joined by spaces, which is what the remote shell reads, and it must be read-only by these same rules, chains and cd included: ssh h 'uptime; df -h' and ssh h 'cd /var/log && tail -n 5 syslog' are read-only. The remote text may not hold a newline or other control character, even quoted, or 2>&1: a csh or tcsh login shell ends a command at a quoted newline and reads 2>&1 as a redirect to a file named 1.

A repository's own configuration still applies to what runs: a diff.external or core.fsmonitor set in .git/config runs for git diff and git status as it would for any git command, and kubectl and aws use your kubeconfig and profiles as configured (exec plugins, credential_process). Read APIs also return what was configured into them: Lambda and ECS environment variables, EC2 user data, pod specs, logs, other processes' command lines. Shell aliases, PATH, RIPGREP_CONFIG_PATH and macOS COMMAND_MODE=legacy (where ps -e prints environments) come from your own setup. Keep those under review.

Extending it. Pick the narrowest place:

Team policy: share Reflex rules across a repo

Team guardrails for AI coding agents belong with the code they protect. A repository can commit .reflex/policy.json, and every teammate's Reflex applies it while Claude Code, Codex CLI, opencode, pi or Hermes works in that repository. It works like Claude Code team settings in a checked-in .claude/settings.json, with one difference: a team policy can only make Reflex stricter, unless each teammate trusts it.

reflex policy init writes a starter with stricter examples only:

{
  "version": 1,
  "mode": "enforce",
  "rules": [
    {"id": "drop-table", "outcome": "deny", "shell": true, "rule": "drops a database, schema or table",
     "all": ["\\bdrop\\s+(database|schema|table)\\b"]},
    {"id": "migrations", "outcome": "ask", "shell": true, "rule": "runs database migrations",
     "all": ["\\b(db:migrate|alembic\\supgrade|prisma\\smigrate\\sdeploy)\\b"]}
  ],
  "always_human": [{"id": "billing", "rule": "calls the billing service", "all": ["\\bbilling-api\\b"]}],
  "prod": ["\\bacme-live\\b", "clusters/main-eu\\b"]
}

Policy packs. reflex policy init --pack <name> writes a ready policy instead of the starter, and never overwrites an existing file. The packs are in examples/policies/, stricter only, with a note on every rule:

PackWhat it adds
awsDenies RDS deletes without a final snapshot, s3 rb --force, turning off CloudTrail, GuardDuty, Config or Security Hub, and KMS key deletion; asks for any IAM or Organizations command that is not read-only, EC2 terminate, Route 53 and DynamoDB deletes; AWS profiles such as prod or acme-prod (not nonprod) are production.
eksDenies EKS cluster and node group deletes and kubectl delete --all / -A; asks for namespace, volume, stateful set and CRD deletes, drains, scale to zero, RBAC and aws-auth changes, and Helm uninstall or rollback; kube contexts such as prod-eu are production.
terraforminfra with destroy: deny and require_plan_in_prod; denies state push, state changes with -lock=false and workspace deletes; asks for state edits, imports, taints, -target and Terragrunt run-all; workspaces such as prod are production.
startup-defaultDenies DROP of a database, schema or table and gh repo delete; asks for truncates, migrations, force pushes, package and image publishes, GitHub secret and release changes and hosted app deploys (Vercel and Netlify --prod, fly, Heroku); a Friday 16:00 UTC freeze for production.

Rename the prod markers to your own profiles, contexts and workspaces, then commit the file.

What each part does. All of them apply as soon as the file is in the repository:

KeyEffect
rulesExtra rules in the rules.json shape, outcome ask or deny. A team deny is checked before the bundled rules, so it wins over a bundled ask; a team ask is checked after them, so it never hides a bundled deny.
always_humanExtra patterns for the always-human class: System 2 never approves a match.
prodProduction markers. A command whose text or context (cwd=, aws_profile=, kube_context=, tf_workspace=, git_branch=) matches, and that is not read-only, asks a human.
modeA floor: enforce turns shadow into enforce for commands in this repository. off stays off: that switch stays with the user.
freezeChange windows and deploy freezes: during a window, a production command that is not read-only asks or is denied. See change freeze.
fastlaneCommand shapes that pass without a prompt ([{"pattern": "^make\\s+lint$"}]). This loosens, so it needs trust (below).
notifyA decision webhook, like the one in config.json (audit log). It sends data off the machine, so it needs trust (below).

The file has no key that removes a rule, raises a threshold or turns off the injection guard, the runaway guard or the tamper check, and an unknown key makes it invalid. Patterns are case-insensitive, at most 500 characters, 64 per file, and run on V8's linear-time regular expression engine, because a hook that times out lets the command run and the command is the agent's choice. A pattern that engine cannot run in linear time is rejected: lookaheads, backreferences and large bounded repeats such as {0,64} (use * or + instead).

Trust, for the fast lane and the webhook only. reflex trust . shows the file, its sha256, every fast lane entry and the host of its webhook, and asks you to type trust on the terminal. It records the repository path and the hash in ~/.config/reflex/trusted.json. Any change to the file, from a teammate's commit or from an agent, changes the hash, and the fast lane stays off until you review it and run reflex trust . again. reflex trust --revoke . removes the trust. Team fast lane entries are held to the same checks as fastlane.json: anchored patterns with no wildcard, scoped to the repository, never a denied word, never over a deny, a secret read, the tamper check, a prod marker or the always-human class.

Where it is read. At repository roots only: directories holding a .git, as for conditional instructions. The stricter parts come from every repository root around the working directory, so a nested checkout, a submodule or a .git an agent creates in a subdirectory still gets the policy of the repository it sits in. The fast lane comes from the nearest root only. A root and its policy file count only when your user owns them (as with git's safe.directory), so another account cannot plant /tmp/.git and /tmp/.reflex/policy.json. A .reflex/ in a directory without .git is never read, and outside a repository there is no team policy. The file must be a regular file (not a symlink) of at most 64 KB.

What an agent is stopped from doing, and the limits. These are tamper asks, so a human confirms them in every mode: a shell command that writes under .reflex/ (also through a glob such as .ref*, or naming policy.json in a repository with a team policy), and a command or a local script it runs that calls reflex trust, reflex policy init or team.mjs, or names trusted.json. reflex setup also adds Edit(**/.reflex/**) to Claude Code's ask rules, so its file tools prompt too. reflex trust refuses without a terminal and inside an agent session (CLAUDECODE, CODEX_SANDBOX, REFLEX_AGENT); run it in your own terminal. Those two checks are speed bumps, not a boundary. The limits are the same as for fastlane.json:

Invalid files. A file that does not validate never loosens anything, trusted or not. Its valid stricter parts still apply, entry by entry, and reflex doctor lists each problem.

Seeing it. reflex status and reflex doctor show the team policy of the directory they run in: its path, whether you trust it, its hash, and what it adds. reflex policy prints the same for one repository. Decisions record it too: a rule from the team policy says (team policy), and the policy version reads rules-v16+team-<hash>. reflex check --cwd <dir> and reflex replay judge each command with the team policy of its own working directory.

Plan-aware terraform gate: stop AI agents from destroying infrastructure

A rule can tell that terraform destroy destroys. It cannot tell what terraform apply will do: the same command creates one tag on Monday and replaces the production database on Tuesday. So when a coding agent (Claude Code, Codex, opencode and the others) runs terraform apply, Reflex judges the change, not only the command text. This is the part that prevents a terraform destroy hidden inside an apply.

What the hook runs, and what it never runs. The hook never runs terraform plan or terraform apply. A plan executes providers with your credentials, runs data "external" programs and can take minutes. With infra.terraform_show on, the hook reads a saved plan the agent already made, with terraform show -json <planfile> in the directory the command runs in (-chdir= and a leading cd dir && are followed). terraform show reads the plan file locally and configures no provider, so it makes no provider API calls, but it does start the provider binaries to read their schemas.

Why plan reading is off by default. Those provider binaries live in .terraform/providers, and an agent can write that directory, and .terraform.lock.hcl, with its file tools (Write, Edit, apply_patch), which the command gate does not see. Reading the plan would then make the hook itself run code the agent put there, before any approval. So infra.terraform_show is false by default: a terraform apply <planfile> is judged as before (the rules, then Jev or, keyless, a human), and an apply without a plan file still asks with the fix. For the same reason terraform plan, show, validate, state show, providers, graph and init are no longer on the read-only list or the fast lane, nor output and state list, which start the backend saved in .terraform; fmt -check and version still are.

The plugin cache requirement. Turn it on with "infra": {"terraform_show": true} only with a provider plugin cache: TF_PLUGIN_CACHE_DIR, plugin_cache_dir in ~/.terraformrc, or ~/.terraform.d/plugin-cache. With a cache, terraform init puts symbolic links into .terraform/providers instead of copies ("when possible", in HashiCorp's words). The hook runs terraform show only when:

Anything else (a copied provider, a regular file in .terraform, a link back into the tree, a packed filesystem mirror that Terraform had to extract) asks: "no readable saved plan (terraform show not run: ...)". The cache itself is trusted as yours: an agent that can write your home directory outside the gate can also write it. The check keeps the working tree out, which is where an agent's file tools usually write. It runs with a strict timeout (infra.timeout_ms, 3 s by default and 4 s at most, so the whole hook stays inside its 10 s), a sanitized environment (no AWS_*, GOOGLE_*, ARM_*, TF_VAR_* or tokens; only PATH, HOME, the locale and Terraform's data and plugin directories) and CHECKPOINT_DISABLE=1, so Terraform does not call HashiCorp's version service either. The terraform binary comes from an absolute PATH entry, never a relative one such as ./bin.

What it decides.

The commandOutcome
terraform apply tfplan with infra.terraform_show off (the default)judged as before: the rules, then Jev or, keyless, a human
terraform apply tfplan, the plan has 0 deletes and 0 replacesallow-eligible: the usual policy decides, with the counts in Jev's state (keyless: pass). In production it asks, and the reason shows the counts
the plan deletes or replaces anythingdeny (infra.destroy: "ask" softens it), for example plan destroys 3: aws_db_instance.main, aws_s3_bucket.logs, aws_iam_role.ci (1 replace); stateful: aws_db_instance.main, aws_s3_bucket.logs
terraform apply or terraform apply -auto-approve, no plan fileask: "terraform apply without a saved plan: run terraform plan -out=tfplan and apply the plan file". Deny in production with infra.require_plan_in_prod
the plan file is missing, not a plan (a state file shows as JSON too), stale, the providers are not all linked from the plugin cache, or terraform show failed or timed outask: "no readable saved plan (...)" with the same fix
terraform destroy, apply -destroy, apply -replace=unchanged: the destroy rules ask, and deny in production

A replace is a delete plus a create (["delete","create"] or ["create","delete"] in the JSON plan format). Stateful types are named first: aws_db_instance, aws_rds_cluster, aws_s3_bucket, aws_dynamodb_table, aws_efs_file_system, aws_ebs_volume, ElastiCache, Redshift, DocumentDB, KMS keys, google_sql_database_instance, google_storage_bucket, BigQuery, azurerm_*database*, storage accounts, kubernetes_persistent_volume*, namespaces, statefulsets and others. A plan is stale when a .tf, .tf.json, .tfvars, .terraform.lock.hcl or local terraform.tfstate in its directory is newer than the plan file. Terraform itself also refuses to apply a plan whose state moved.

A plan's deny is a rule outcome: it holds in shadow and enforce mode, and no approval in the queue lifts it. A plan's ask goes to a human, never to System 2. With the Jev engine in enforce mode, Jev still judges the command under the ask, and a deny Jev finds stands. A clean plan passes on its own only when the command is nothing but cd steps and the apply: terraform apply tfplan && ./deploy.sh gets the counts in its trace, and the rest is judged as usual. It also needs the apply to run exactly the plan that was read, so none of these pass (they keep the counts, and the usual judge decides):

Production is read in the directory the command runs in, the physical one too (a current symlink to envs/prod), and the team policy of that directory counts as well as the one the command started in. The hook budget: when the rules and the plan read took more than 5 s, Jev is not waited for as well, and the command asks.

Decision JSON and trace. Every decision the plan gate spoke to carries the counts:

{"effective": "deny", "decision": "deny", "reason": "reflex (rule): plan destroys 1: aws_instance.old", "source": "rule",
 "policy": "rules-v22", "plan": {"kind": "terraform", "create": 1, "update": 0, "delete": 1, "replace": 0, "stateful": [], "digest": "..."}}

The digest (of the plan JSON) is also part of the approval queue key and the Jev cache key, so an approval of one plan is never reused for a different plan under the same command.

The workflow it asks agents for. Plan, read, apply the file:

terraform plan -out=tfplan        # judged like any command that runs providers; the agent runs it, not the hook
terraform show -json tfplan | jq '.resource_changes[] | select(.change.actions != ["no-op"]) | .address'
terraform apply tfplan            # judged by what tfplan will change

kubectl delete guardrail (optional). With "infra": {"kubectl_diff": true} (off by default, because it calls the API server), kubectl apply is checked with kubectl diff and the same arguments, and kubectl delete|replace|patch with --dry-run=server -o name added at the end. Both use the current kube context (or the command's --context), the same timeout, and never a flag that writes: --dry-run=server -o name goes right after the verb, so an option of the command left waiting for a value cannot take it, and a command with its own --dry-run, --raw, --, -f - or -o is not run at all. Nor is one that names its own --kubeconfig, --server or --token, or runs with a KUBECONFIG inside its working directory: an agent's kubeconfig could carry an exec credential plugin, and an agent's server would receive your credentials. kubectl diff exits 0 for no differences, 1 for differences and above 1 on an error. Deletes of namespaces, PVCs, PVs, statefulsets or CRDs follow infra.destroy (deny by default); other deletes ask with the count; changes without deletes only add the counts. Off, or on any failure, kubectl commands are judged as before, and the production markers (--context prod, a prod kube context) still deny destructive ones.

Configuration. In ~/.config/reflex/config.json:

{"infra": {"enabled": true, "destroy": "deny", "require_plan_in_prod": false, "terraform_show": false, "kubectl_diff": false, "helm_diff": false, "timeout_ms": 3000}}

A team policy can only make it stricter (.reflex/policy.json):

{"version": 1, "infra": {"destroy": "deny", "require_plan_in_prod": true}}

destroy there accepts only "deny" and require_plan_in_prod only true; a team infra section also turns the gate on for a user who turned it off. reflex doctor shows the settings and where terraform, tofu, kubectl and helm were found.

Measured on real sessions. A replay of 2,232 local Claude Code and Codex transcripts found 63 unique commands that mention terraform ... apply or a mutating kubectl verb. Most only mention it (commit messages, heredocs that write CI workflows); the deterministic outcome changed for 2 real applies, both from "left to the judge" to an ask with the fix: one apply without a plan file and one whose plan file was gone. With the local engine those already asked, so the effective change there is the reason, not the outcome.

Limits. The plan is read when the hook runs and applied a moment later: a process the agent left running could swap the file in between (Terraform still refuses a plan whose state moved). A command whose text hides what runs ($VAR, a heredoc) is not read, and is judged as before. Terragrunt and Terraform Cloud saved plans are not read (OpenTofu and Terragrunt). kubectl diff does not show objects that a --prune would delete unless the command has --prune.

OpenTofu and Terragrunt

OpenTofu AI agent guardrails and Terragrunt guardrails work like the terraform gate above, with the differences below.

OpenTofu (tofu). tofu apply <planfile> goes through the same plan gate: with infra.terraform_show on, the hook reads the saved plan with tofu show -json <planfile> (never tofu plan or tofu apply), in the same sanitized environment, with the same timeout, the same counts and the same outcomes. The provider check is the same too: every entry under .terraform/providers must be a symlink into a plugin cache (TF_PLUGIN_CACHE_DIR or plugin_cache_dir) under your home directory and outside the working tree, not newer than the plan. The CLI config files it reads for dev_overrides and plugin_cache_dir include OpenTofu's: TF_CLI_CONFIG_FILE, else ~/.tofurc, $XDG_CONFIG_HOME/opentofu/tofurc (and *.tfrc there), ~/.terraformrc and ~/.terraform.d/*.tfrc. A .tofu or .tofu.json file newer than the plan makes it stale. tofu apply without a plan file asks with the fix: "tofu apply without a saved plan: run tofu plan -out=tfplan and apply the plan file". OpenTofu reads state and plan encryption from the root module in the working directory, and a key provider can run a program, so tofu show is not run (the apply asks) when any .tf, .tofu or .tf.json file there mentions encryption or key_provider; an encrypted plan is not a zip file either. tofu show gets XDG_CONFIG_HOME and XDG_DATA_HOME, so it reads the same tofurc the check read (a relative one asks). tofu destroy, tofu state rm, tofu apply -destroy and -replace= hit the destroy rules: ask, and deny in production. The tofu binary comes from an absolute PATH entry, and ./tofu never passes.

Terragrunt. The hook never reads a Terragrunt plan: terragrunt apply runs the before_hook, after_hook and run_cmd of terragrunt.hcl, which an agent can write, and terragrunt itself picks tofu or terraform and, with terraform.source, a copy in .terragrunt-cache. So:

The commandOutcome
terragrunt apply, terragrunt run-all apply, terragrunt run --all apply, terragrunt run -- apply, apply-allask: "terragrunt apply without a saved plan: run terragrunt plan -out=tfplan (a stack: terragrunt run --all --out-dir DIR plan) and apply that plan". Deny in production with infra.require_plan_in_prod
terragrunt apply tfplan, terragrunt run --all apply --out-dir DIRask: the saved plan is not read; review it with terragrunt show first
terragrunt destroy, run-all destroy, run --all destroy, destroy-all, apply -destroythe destroy rules: ask, and deny in production

Production for all three comes from the usual markers: the working directory (envs/prod, infrastructure-live), -chdir=, --working-dir live, tofu workspace select live, TF_WORKSPACE, a .terraform/environment workspace in the working directory, the AWS profile and the git branch.

helm guardrails for AI agents

helm install, helm upgrade, helm uninstall and helm rollback change a cluster, and a helm uninstall from an AI agent takes every object of the release with it (PVCs too, unless the chart keeps them).

The commandOutcome
helm uninstall, helm delete, helm rollbackthe destroy rules: ask, and deny in a production kube context or namespace (--kube-context prod-eu, -n live, HELM_NAMESPACE=live, the current context of your kubeconfig). Global flags before the verb count (helm --kube-context prod uninstall api)
helm upgrade --install, helm upgrade, helm install in productionask, with the release and namespace: "production helm upgrade --install of release api in namespace web, kube context prod-eu"
the same outside productionjudged as before (the rules, then Jev or, keyless, a human)

helm diff (optional). With "infra": {"helm_diff": true} (off by default: it runs a helm plugin and calls the API server), the hook checks helm upgrade and helm install with helm diff upgrade <release> <chart> and the command's own values (-f, --set*, --version, --repo, -n, --kube-context), plus --allow-unreleased for an install, --output structured, --no-color and --suppress-secrets. It needs helm-diff 3.15 or later (the structured output). What it decides:

It is not run at all, and the command is judged as before, when the command names its own --kubeconfig, --kube-apiserver, --kube-token or other connection flag, a --post-renderer, --dry-run, --force, -o, repository or registry config, or any flag this does not know; when KUBECONFIG is relative or inside the working directory; and when helm-diff is not installed. Its environment is an allowlist: PATH, HOME, the locale, KUBECONFIG, XDG_*, helm's own directories, HELM_NAMESPACE and HELM_KUBECONTEXT, the proxy variables, and the cloud variables an exec credential plugin needs (AWS_*, GOOGLE_*, CLOUDSDK_*, AZURE_*). HELM_DIFF_* (an external diff tool, a template file, another output), HELM_KUBEAPISERVER, HELM_KUBETOKEN and the rest are left out. A chart from a repository or oci:// registry is fetched with your helm credentials, as helm diff would; the timeout kills helm, but a helm-diff process under it can finish its API calls after the hook has asked.

The plugin check. helm plugins are code on disk an agent's file tools could write, like the provider binaries. helm diff runs only when every helm plugin directory (HELM_PLUGINS, else $HELM_DATA_HOME/plugins, $XDG_DATA_HOME/helm/plugins, or ~/Library/helm/plugins on macOS and ~/.local/share/helm/plugins on Linux) is under your home directory and outside the working tree, nothing in it (through the links a local helm plugin install makes) points into the tree, every plugin's plugin.yaml (64 KB at most) runs only a program in its own directory (command, platformCommand and downloaders; a downloader can run for a chart URL), and no file or link there changed after the trusted mark: your Reflex config.json. The change time counts as well as the modification time, so touch -t cannot hide a new file. After helm plugin update diff, save config.json again (you, not the agent: that is a tamper rule) to trust the new files. If unsure, keep helm_diff off: production upgrades still ask.

Gate MCP tool calls

MCP server guardrails for AI coding agents: agents change infrastructure through MCP servers (AWS, Kubernetes, Terraform Cloud, databases, GitHub), not only through the shell. Reflex judges each MCP tool call before it runs, with the same ladder as a command, and blocks destructive MCP tool calls that point at production.

Where it runs. Claude Code and Codex CLI (mcp__<server>__<tool>, the PreToolUse hook), opencode (tool.execute.before: every tool that is not built in), pi (pi-mcp-adapter's mcp proxy tool and mcp_* direct tools) and Hermes (mcp_<server>_<tool>, the pre_tool_call block reflex setup --agent hermes prints). reflex setup and the Claude Code and Codex plugins wire it; re-run reflex setup once to add it to an earlier install (reflex doctor says when a hook predates it). Until then the Claude Code and Codex plugins judge the MCP and file tools the older hook does not match, and stand down only for the tools it does.

How a call is decided (setup/tool-gate/mcp.json, tools.mjs):

  1. A shell command in an argument (command, cli_command, cmd, script), as the AWS MCP server's call_aws takes, goes through the shell rules as if the agent ran it: aws ec2 terminate-instances --profile prod is denied, aws s3 ls passes.

  2. The rules, matched on the server, the tool name as words (deleteStack, delete-stack and delete_stack are the same) and the arguments as JSON:

    RuleMatchesOutcome
    mcp-destructivea tool name with delete, destroy, drop, terminate, remove, rm, purge, truncate, force, reset, rollback, uninstall, wipe, erase, revoke, detach, deregister, kill, flush, flushall, unlink, rmdir, overwrite or shutdownask
    mcp-scale-to-zeroscale_* with replicas, desired count or size 0, scale_to_zeroask
    mcp-bucket-policyput_bucket_policy, put_bucket_acl, put_public_access_block and similarask
    mcp-security-groupupdate_security_group*, authorize_security_group_ingress and similarask
    mcp-iamput_role_policy, attach_role_policy, create_access_key, any put, attach, create or update on an IAM serverask
    mcp-sql-destructivea statement in a sql, query or statement argument, comments stripped, that runs DROP, TRUNCATE, DELETE or ALTER (not on docs, search or memory tools)ask
    mcp-http-deletean HTTP method of DELETE (method, http_method, httpMethod or any key naming a method or verb)ask
    protected-pathan MCP tool that writes files (write, edit, move, create and similar verbs) with a path, destination or similar argument on a protected pathask

    Every one of them is denied instead when the arguments or the server point at production: the prod-destroy markers (prod, production, prd, live paths) on each argument value (a stack, cluster, context, profile, database, workspace or account name, a URL), on the server name, and on the context that server kind reads (an AWS server the AWS profile and region, a Kubernetes one the kube context, a Terraform one the workspace), plus a team policy's prod list, where account ids belong.

  3. Read-like tools pass at once and are not logged, as a read-only command: a tool whose first word is get, list, describe, search, read, query, fetch, find, show, view or another read verb in mcp.json (after the server's own name when the tool repeats it, as aws___search_documentation does; update_status and resolve_incident are not reads). A sql argument must then be SELECT only: no write keyword outside strings and comments, one read statement after another. Browser navigation (not fill, type, upload, scripts or dialogs), local agent memory and an HTTP GET (an explicit GET, or no method and no body) pass the same way (mcp.json pass).

  4. Anything else is an unknown MCP tool. Keyless (engine local) it passes and is logged in the trace, so reflex report and reflex audit show it. With Jev it goes to the engine: shadow judges it in the background, enforce waits for Jev, which gets a typed state with the server, the tool, the redacted arguments and the production tier. "mcp": {"unknown": "ask"} in config.json makes every unknown MCP tool ask instead.

A change freeze, the approval queue, the runaway guard, the trace, the audit export and the decision webhook apply to MCP decisions as to commands; the trace records a call as mcp <server>/<tool> {redacted arguments} (input <hash>), and a queue approval covers that exact call. The gate never allows an MCP call: a pass leaves the agent's own permission prompt in charge. The injection guard still scans MCP results after they run.

Measured on real sessions. 30 days of one engineer's Claude Code transcripts (1,630 files, 475 MCP tool calls across Grafana, Chrome DevTools, the AWS MCP server, Cloudflare, Claude in Chrome, Context7 and memory servers): no call asked or was denied, 315 (66 %) passed at once as reads and 160 (34 %) were unknown tools that passed and were logged. With mcp.unknown: "ask" those 160 would have asked; most were Grafana alert rule and dashboard updates, page scripts and AWS run_script. The golden set, setup/tool-gate/golden-mcp.json (83 cases), runs with npm run eval-mcp: 83 of 83 keyless and with Jev, no miss. The hook adds about 50 ms per call.

Change the rules. A copy of mcp.json in ~/.config/reflex/tool-gate/ replaces the bundled one (a tamper rule asks before an agent edits it). A team policy adds rules, stricter only:

{"version": 1,
 "mcp": [{"id": "no-prod-writes", "outcome": "deny", "rule": "no writes to the billing database",
          "server": "postgres", "args": "billing"}],
 "prod": ["\\b123456789012\\b"]}

server, tool and args are case-insensitive patterns (at least one), held to the team policy's linear-time check like its other patterns; outcome is ask or deny, nothing looser.

Limits. A rule reads a tool's name and arguments, not what the server does with them: an MCP tool named apply_run or update_function_code is unknown to the rules. Keyless it asks on an infrastructure server (the MCP infra preset, below) and is only logged elsewhere; with a key Jev judges it. opencode and Hermes do not say which server a tool belongs to, so a server pattern reads the whole tool name. opencode and Codex cannot show an approval dialog for a tool call: an ask blocks with a reason that tells the agent to have the user make the change.

The MCP infra preset (config.json "mcp": {"infra": false} turns it off). An unknown tool on a server named for cloud, clusters, IaC, a database or a code host asks keyless instead of being only logged. The match reads the server name alone, split into words: aws, awslabs, kubernetes, k8s, kubectl, terraform, tfc, gcp, gcloud, azure, postgres, postgresql, mysql, mongodb, dynamodb, database, db, supabase, github, gitlab, or any server name that contains prod. aws-mcp, awslabs.eks-mcp-server and mcp-server-kubernetes match; a tool called update_database on a notion server does not, and dbt is not db. Hermes and opencode give no server name, so the first word of the tool name stands for it.

Protected files

The shell rules already ask before a command writes the gate, agent settings or a team policy. File tools write too, so Reflex also judges the file writes of Edit, Write, MultiEdit and NotebookEdit (Claude Code), apply_patch (Codex, opencode), write, edit and apply_patch (opencode), edit and write (pi) and write_file and patch (Hermes). A write to a protected path asks; any other write passes at once and is not logged.

The default set (setup/tool-gate/protected.json):

PathWhy
.github/workflows/**, .github/actions/**, .gitlab-ci.ymla CI pipeline runs with the repository's secrets
**/envs/prod/**, **/prod/**/*.tfproduction environment configuration and Terraform
**/*.tf, **/*.tfvars, **/Dockerfile* in a production pathproduction Terraform, variables and images (a prod, production, prd or live marker in the path within the repository, or a team policy prod marker)
**/.reflex/**a team policy: a human edits it
**/.git/hooks/**, **/.git/config, **/.husky/**, ~/.gitconfig, ~/.config/git/**git hooks and git config (core.hooksPath, fsmonitor) run code on git commands
**/.mcp.json, **/.envrc, **/.vscode/tasks.jsona new MCP server, a direnv file or an editor task runs code
~/.ssh/**, ~/Library/LaunchAgents/**, ~/.config/systemd/user/**SSH keys and config, login services
.claude/settings*.json, .claude/hooks/**, ~/.claude/plugins/**, ~/.claude/agents/**, ~/.claude.json, .codex/hooks.json, ~/.codex/config.toml, ~/.codex/rules/**, ~/.config/opencode/**, .opencode/plugin*/**, ~/.pi/agent/**, ~/.omp/agent/**, ~/.hermes/**agent settings, hooks and plugins
~/.bashrc, ~/.zshrc, ~/.profile and the other shell startup filesa shell startup file runs in every new shell
the Reflex checkout, its logs and ~/.config/reflexalways, whatever the file says (a git checkout nested in the Reflex checkout is not the gate, as for the tamper rule)

A glob that starts with ~/ or / is anchored there; any other matches at any depth, as in .gitignore. Symlinks are resolved, so a link into .github/workflows is the workflow, and on macOS and Windows globs ignore case, as those file systems do (.GitHub/Workflows is the same directory). A write whose files cannot be read (a patch in a shape Reflex does not parse) asks. An MCP filesystem server's write to a protected path asks too. The reason names the path and why it is protected:

reflex (rule): writes a protected path (.github/workflows/deploy.yml): a CI workflow runs with the repository's secrets

The decision takes the same ladder as a rule's ask: a human answers it (in the autonomous profile it is always-human, never System 2), a change freeze on production turns it into a deny, and the trace records <tool> <path> (input <hash>), never the content. In the Claude Code plugin the hook only asks or denies; it never allows and never rewrites the input.

Configure it. "protected": ["docs/runbooks/**", "k8s/overlays/prod/**"] in config.json adds globs; a copy of protected.json in ~/.config/reflex/tool-gate/ replaces the default set. A team policy adds globs, stricter only: "protected": ["docs/runbooks/**"]. Measured on the same 30 days of transcripts (edits of the Reflex checkout the replay ran from left out), 40 of 3,288 file writes (1.2 %) would have asked: 32 CI workflow and action edits, the rest team policy files, .mcp.json, agent settings and one production path.

Change freeze for AI coding agents

A deploy freeze for AI coding agents: during a change window you define, a command that is not read-only and touches production asks a human, or is denied. This is change management for Claude Code, Codex CLI and the other agents Reflex gates, in the same place as the rest of the policy. Windows go in ~/.config/reflex/config.json (yours) or in a team policy's freeze list (the repository's), and both apply:

{
  "freeze": [
    {"days": ["fri"], "after": "15:00", "tz": "Europe/Bucharest", "applies_to": "prod", "outcome": "ask"},
    {"from": "2026-12-20", "to": "2027-01-03", "outcome": "deny", "note": "year-end freeze"}
  ]
}
FieldMeaning
daysDays of the week: mon, tue, wed, thu, fri, sat, sun.
after, beforeLocal time, HH:MM. after is inclusive, before is exclusive, and after must be earlier than before; a window across midnight is two windows.
from, toDates, YYYY-MM-DD, both inclusive.
tzAn IANA time zone name such as Europe/Bucharest. Default UTC. Times and dates are read in it with Intl, so summer time is handled.
applies_toprod (default): commands that touch production. all: every command that is not read-only.
outcomeask (default) or deny.
noteText added to the reason, at most 200 characters.

Every field a window has must hold at once: {"days": ["fri"], "after": "15:00"} is Friday from 15:00 to midnight. A window needs at least one of days, after, before, from and to.

What counts as production. The markers Reflex already uses: the prod-destroy rule's production pattern over the working directory, the AWS profile, the kube context, the Terraform workspace, the git branch and the command text (envs/prod, --context prd-eu, aws_profile=production, tf_workspace=live and so on), plus the prod markers of the team policy. A command whose pipelines only write notes (git commit -m "prod fix") is not production.

What the agent sees. A frozen command gets a rule decision with a reason such as

reflex (rule): change freeze: Friday after 15:00 (Europe/Bucharest); production (cwd)

The reason names the kind of marker (cwd, aws_profile, kube_context, command), never its value, since it can reach a webhook. The trace and reflex audit keep the matched text.

It is a deterministic decision, so it applies in shadow mode too, and in the autonomous profile it is in the always-human class: System 2 never approves it. A human's approval in the approval queue can lift a freeze ask for that one command, but only when the item was parked and answered inside the window, so an approval given before the freeze never carries into it. Nothing lifts a deny.

It only tightens. A window has no outcome that passes. A freeze outcome replaces a pass (the read-only list excepted), a fast lane pass, a Jev or local judgment, and a rule ask when the window denies. A rule deny, or a rule ask under an ask window, keeps its own reason. Validation is strict: an unknown field, a day such as friday, a time such as 9:00, a date that does not exist, an unknown time zone or an outcome other than ask or deny is an error, never a smaller window. In a team policy, an invalid window makes the file invalid, so its fast lane and webhook are off while its valid stricter parts, other windows included, still apply. In config.json, an invalid window is never dropped: until you fix it, every command that is not read-only asks, while the rules keep running (a rule deny stays a deny). reflex doctor names the window and the problem.

Seeing it. reflex status prints Change freeze: ACTIVE now: ... or none active, with the number of windows, for config.json and the team policy of the directory it runs in. reflex check "kubectl apply -f app.yaml" --cwd ~/infra/envs/prod shows the decision a frozen command would get. reflex audit --prod-only lists what happened during a freeze.

Limits. The window is read from the machine's clock. You can edit your own config.json or turn Reflex off (--mode off, where no freeze applies); a freeze is a control on the agent, not on you. Production is read from the command and its context, not from the scripts it runs: a ./deploy.sh whose only production reference is inside the script is not frozen by a prod window (use "applies_to": "all" for a full stop). A team policy's windows are protected like the rest of the file: an agent shell command that writes .reflex/ gets a tamper ask, and code review protects the committed file. The production markers read text, so a command that names production in an argument (gh pr create --title "prod fix") counts, which only adds asks. Reflex gates shell commands; file edits and MCP calls are not frozen.

Audit log for AI agent commands (SOC 2)

reflex audit exports one row per decision the gate logged: a document auditors can use as evidence for change management controls such as SOC 2 CC8.1 and ISO 27001 Annex A 8.32. It shows which agent ran what, where, in which environment tier, what Reflex decided and who approved it. It reads the trace (trace.jsonl and its rotated files), the approval queue and the execution feedback, writes nothing and calls nothing.

reflex audit                                   # the last 7 days, csv on stdout
reflex audit --since 90d --format csv > agent-changes-q3.csv
reflex audit --since 24h --prod-only --format json
reflex audit --agent claude-code --format jsonl
ColumnMeaning
timeWhen the gate decided (UTC, ISO 8601).
agent, sessionThe agent (claude-code, codex, opencode, pi, hermes, shell) and its session id.
cwdThe working directory.
env_tier, env_reasonprod or non-prod, and the marker that made it production (cwd=/infra/envs/prod, aws_profile=production, command: prd). unknown for rows logged by an earlier version.
commandThe command, with secrets redacted when it was logged and again on export.
decisionWhat the agent was told: pass, allow, ask or deny.
judged, mode, sourceWhat the judgment was before the mode applied (in shadow mode a Jev ask is logged, not shown), the mode, and who decided: rule, fast-lane, jev, local, judge (System 2), queue, runaway.
rule_id, ruleThe rule or reason, such as freeze, prod-destroy or team:prod.
approved_byWho answered, when that is known: approved in the approval queue (q-1a2b3c4d5e by alice at ...), System 2 approved (confidence 0.9), approved at the agent's prompt (it ran), rejected at the agent's prompt, or no answer recorded. The queue records the account that ran reflex queue approve.

--format is csv (default), json or jsonl. In csv, a cell that a spreadsheet would run as a formula (starting with =, +, - or @) starts with a quote. --since takes 7d, 12h or 30m; --prod-only keeps production rows; --agent keeps one agent.

What an auditor should know about it:

Decision webhook

Reflex can post decisions to a webhook, for a Slack channel or a log pipeline:

{"notify": {"url": "https://hooks.slack.com/services/T000/B000/XXXX", "on": ["deny", "ask", "prod"], "format": "slack"}}

on lists what is sent: deny and ask (what the agent was told) and prod (any judged production command, whatever the decision). The default is ["deny"]. format is json (default: an object with event: "reflex.decision", time, agent, session, cwd, prod, prod_by, command, decision, judged, mode, source, rule_id and reason; the cwd has your home directory as ~) or slack (a text message, with <, > and & escaped so a command cannot mention a channel).

A webhook is data egress, so it is held to these rules:

reflex doctor --notify-test sends one dry-run message ("dry_run": true, no command) to each configured webhook and reports the HTTP status. Doctor sends nothing without that flag.

Metrics

node report.mjs --push <pushgateway> exports gauges for the report window:

MetricLabels
reflex_decisionssource, decision, mode, user
reflex_latency_secondsquantile (0.5, 0.95, 0.99), user
reflex_asksresolution (approved, rejected, pending, unknown), user
reflex_replay_changesuser: decisions the current policy file would flip
reflex_input_tokensuser

dashboards/reflex.json shows them in Grafana.

Data handling

Reliability

Reflex fails closed

A hook that crashes must not let a command through unchecked. Most agents treat a hook that exits with a plain error as "no decision" and run the command, so every hook starts through a small entry, hook.mjs, that uses only Node built-ins:

node hook.mjs /path/to/gate.mjs --claude --mode enforce --allow off

It installs handlers for uncaught exceptions and unhandled rejections, then loads the script with a dynamic import. So a syntax or import error, a throw while the modules load (a bad config value, a top-level parse), or a rejection nobody handled still gives the agent an answer in its own hook contract. The reason is reflex error: <short message>; a human must review.

Entry pointOn an error, in enforce mode
Claude Code PreToolUse (--claude)permissionDecision: "ask", exit 0 (the JSON is read only on exit 0; exit 2 would block, not ask)
Codex PreToolUse (--codex)permissionDecision: "deny" with the reason, the reason on stderr, exit 2 (Codex does not support ask yet)
Hermes pre_tool_call (--hermes){"action": "approve"}, Hermes' own prompt, with a rule key used once
opencode, pi, oh-my-pi (--decide){"effective": "ask"}; the adapters also treat no answer, or one that is not JSON, as ask
reflex-sh (--sh)a y/N on the terminal; no terminal or no refuses with exit 126
PostToolUse, UserPromptSubmit, record hooks, the injection guard, instructionsnever blocked: exit 0 with a systemMessage warning; after a tool result the guard also tells the model it did not check the result
The decision webhook's detached child (notify.mjs --send)logged only; it never touches a decision

A change freeze (freeze.mjs) is read while the gate loads and when it decides: an error in it asks like any other. reflex audit is a terminal command, not a hook, so an error there is a plain non-zero exit.

Safety properties and limits

Injection guard

The gate judges what an agent runs. guard.mjs judges what it reads first: a web page, search results, an MCP result, a file from someone else's project or the output of curl can carry text written to steer the agent (indirect prompt injection). After such a tool runs, the guard scans its result and returns pass, warn or block; on each prompt it also checks for a pasted credential.

Which results are inspected (sources in setup/injection/policy.json):

SourceToolsWhen
webWebFetch, WebSearch, opencode webfetch / websearch / codesearch, omp web_fetch / web_search / github and its read of a URL, Hermes web_search / web_extract / x_search / feishu_doc_read / browser_* (not the password vault)always
mcpany MCP tool (mcp__server__tool, Hermes connectors__…; opencode: any tool that is not built in)always
fileRead, read, read_filethe file is outside the project (the nearest directory holding .git above the working directory; in Claude Code above CLAUDE_PROJECT_DIR, so a cd into a cloned repository does not make its files the user's), or inside it under node_modules/, vendor/, third_party/, site-packages/, .venv/, .cache/; never a credential file (~/.ssh/, ~/.aws/, ~/.kube/, .env*, .netrc, *.pem, ..., exclude), whose content would otherwise go to Jev
shellBash, bash, Hermes terminalthe command fetches remote content: curl, wget, xh, gh issue/pr/api/release/gist/search, glab, npm view, pip download, ...; or a local command that prints someone else's text: a reader (cat, head, tail, less, jq, sed, ...) given a path the file rule would inspect, or git log / show / blame in a repository outside the project. A command that also names a credential file is judged by the detectors only, never sent to Jev

Anything else (edits, greps, local commands, the user's own files) is not inspected. A hook matcher cannot see a Read's path, so every Read starts the guard process (about 50 ms), which returns at once for a file inside the repository.

Deterministic detectors (setup/injection/detectors.json, both engines), each counted as a signal:

Jev (engine jev): the result is cut into chunks with context.mjs's chunk() (at most 3,000 characters each; at most 24 per result, those with a detector hit first, 8 per request with the requests in parallel), and each request asks three questions per chunk (setup/injection/questions.json). A piece under 1,000 characters goes with a neighbour when both fit in one chunk, else with the lines before it (overlapping the previous chunk) up to 3,000 characters: a short last section judged on its own has no page around it to show who it speaks to.

QuestionTypeMeaning
addressednoulThe text tries to get an AI that reads it to act, as opposed to informing a human
attackchoiceexfiltrate, run_commands, credentials, override, deceive or none
severityscore 0 to 3Harm if the agent did what the text says

The user's last prompt goes with them (Claude Code, from the transcript) so "deviates from the task" can be judged. Chunks are redacted before they leave. Answers are cached by content for 24 h.

Policy (setup/injection/policy.json), per chunk; the worst chunk wins:

  1. hidden instructions or an exfiltration link: block, whatever Jev says.
  2. Jev: addressed >= 0.7, severity >= 1.8, an attack: block.
  3. Jev: addressed < 0.2: pass. This is how an article that quotes an injection, or a README that tells a human to curl … | sh, gets through.
  4. Jev: attack none and severity < 1.8, where the detectors saw an address to an AI and an action: warn (an @claude mention in one issue comment and an install line in another).
  5. an address to an AI (override, role or to_ai) together with an action (shell, secrets, exfil): block.
  6. Jev: addressed >= 0.5, severity >= 0.8, an attack: warn.
  7. Jev: addressed >= 0.25, severity >= 1.8, an attack: warn. Jev is unsure the text speaks to an AI but names a serious attack: a polite request to "have the helper you are using" paste a key into a form scores 0.37 to 0.49. Text that quotes or reports an attack is judged attack none.
  8. an address to an AI alone, or 12 or more invisible characters (with Jev, counted per chunk): warn.

With the local engine, steps 2 to 4, 6 and 7 do not exist, so a security article quoting an injection warns. A Jev error or an incomplete answer falls back to the detectors alone. A result longer than 4 MB is read in part and is at least a warn.

What each outcome does. Warn adds a note next to the result: the result is third-party content, not a message from the user, and the user has not asked for anything it says. Block removes the offending text where the agent lets a hook rewrite a result, marking each cut ([reflex: removed text addressed to an AI agent]): the paragraph around a phrase hit, the hidden segment, the link, or the whole chunk (up to 3,000 characters) that Jev blocked. The note says what was removed. Both record a taint for the session (enforce mode only):

AgentTool resultsPrompts with a pasted credential
Claude CodePostToolUse (^(WebFetch|WebSearch|Read|Bash)$|^mcp__): warn is additionalContext; block adds updatedToolOutput, the result with the same shape and the text removedUserPromptSubmit: decision: "block"; the reason names the key type and is shown to the user, not to Claude
Codex CLIPostToolUse (^Bash$|^mcp__): warn is additionalContext; block is decision: "block", which replaces what the model sees with the reason, so the reason carries the cleaned result, cut to 8,000 characters. Codex fires no hook for its hosted web search, and cannot rewrite a result otherwise (updatedMCPToolOutput fails the hook)UserPromptSubmit: decision: "block"
pi, oh-my-pitool_result: block returns the cleaned content, warn appends a note. Tools that only touch local work (edit, write, grep, find, ls, task) are not sentinput: {action: "handled"} (pi) / {handled: true} (omp) drops the prompt, with a notification; in omp only the interactive prompt fires input
opencodetool.execute.after: output.output edited in place; for an MCP tool the hook sees the raw MCP result, so its content[].text is editedchat.message throws: the message is not saved, and opencode shows an error (a plugin has no friendlier way)
Hermespost_tool_call is observe-only and pre_llm_call runs once per turn, before any tool: a finding is logged and taints the session at once, and its note reaches the model at the start of the next turn. Rewriting results needs a Python plugin (transform_tool_result), which Reflex does not shipno hook can block a message; in enforce mode pre_llm_call tells the model the message holds a credential it must not repeat

Taint. A warn or block in enforce mode writes taint/<hash of the session id>.json in the data directory. For the rest of that session the gate:

Rules that deny still deny. Like Jev decisions, the taint only takes effect in enforce mode. A user policy.json seeded by an earlier setup gets what it lacks on the next reflex setup: gates, params and flags of the bundled policy missing from it by id or name are added (a new gate right after the one before it in the bundled order), nothing of the user's is changed or reordered, and setup prints what it added. A deleted bundled gate comes back that way; switch one off with its flag instead. In opencode the taint is kept under a subagent's root session, so a child that read an injection makes its parent stricter.

Credentials in prompts use the credential shapes of setup/redact.json (AWS keys, GitHub, GitLab, OpenAI and Slack tokens, private keys, JWTs, Stripe, Google and npm keys, Slack webhooks; not the bare 40-character shape, which also matches long identifiers), and not its KEY= context patterns, which would stop what does max_tokens: 100 do. The log keeps the key type and a hash of the redacted prompt, never the key.

Modes. The guard follows the gate's mode unless REFLEX_GUARD (or "guard" in ~/.config/reflex/config.json) says otherwise. Shadow judges in a detached background process, logs, and changes nothing: no note, no rewrite, no taint, no blocked prompt. Off does nothing. Any error passes: the guard only adds friction when it has a judgment to add.

Try it:

reflex scan page.html                  # exit 0 pass, 1 warn, 2 block; JSON with the signals
curl -s https://example.com | reflex scan - --rewrite     # also print the cleaned text
npm run eval-injection                 # setup/injection/golden.json against the live API

The golden set holds 62 results: 29 benign documents agents read every day (install guides that pipe curl to a shell, man pages, API docs, npm output, HTML with comments and hidden menus, OWASP pages, a blog post and a long field guide that quote or describe injections, a long README of a coding agent, a GitHub issue that mentions @claude next to an install line, an AGENTS.md from another repo, letter-spaced headings, accented, right-to-left and emoji text) and 33 injections following published research (hidden-text pages, the GitHub MCP issue attack, MCP tool poisoning, Unicode tag and variation-selector smuggling, look-alike letters, JSON escapes, markdown image exfiltration, also reference-style, the rules-file backdoor, EchoLeak-style mail, fake role tags, paraphrased injections with no trigger words: behind twelve chunks of benign text, split across two chunks, on one long line, at the top, middle and end of long pages; letter-spaced text). Current result, jev-1.13.0, five runs: precision 97 %, recall 100 %, every expected outcome met (the @claude issue warns), 28 or 29 of 30 high-severity injections blocked and the rest warned. With the local engine: precision 81 % (articles that quote injections warn or block), recall 79 % (the GitHub MCP issue attack and the paraphrases, which have no trigger phrase, pass). A missed high-severity injection fails the run.

Logs. guard.jsonl in the data directory: one line per inspected result (tool, source kind, hashes of the text and origin, length, signal counts, Jev's numbers per chunk, outcome, what was emitted, latency, tokens, the kind of a Jev error) and per blocked prompt (key types, a hash). Never the text, the URL or the command. reflex report summarises it: results by source, outcome and attack, tainted sessions and blocked prompts, as counts.

Limits. This is a heuristic filter, not a sandbox: it lowers the odds that injected text steers the agent; it does not make untrusted content safe, and the gate still judges every command the agent runs. The detectors are phrase lists and a handful of structural checks, so an injection phrased like ordinary prose passes them (Jev is there for that; the local engine has no answer to it), and text spaced out evenly letter by letter (the same gap between words), or one letter per line, is left to Jev. Jev sees at most 24 chunks of 3,000 characters per result; past that only the detectors read the rest, and past 4 MB nothing does (the result warns). A paraphrase split across two chunks is judged in halves: in the golden set each half still reads as an instruction, but a split where neither does would pass. Hidden elements are found by their own style or hidden attributes, not by class names or stylesheets. A tool the source list does not name (a custom pi extension tool, an MCP tool in opencode that shares a built-in name) is not inspected. The agent has read the result before Codex and Hermes can do anything about it, and in Hermes the note arrives a turn late. A Jev-blocked chunk is removed whole, which can cut useful text next to the injection. Shadow mode never taints, so the gate's taint behaviour only starts once the guard enforces. A hook that times out passes the result unguarded. Subagents run under their own session id in oh-my-pi and Hermes (and Hermes starts a new one after compressing the context), so taint does not cross between a parent and its subagents there. Codex fires no PostToolUse for a failed tool or an MCP error result, and scans only the last chunk of a long-running command. In Claude Code, WebFetch hands the hook the fetch model's summary of the page, not the page.

Runaway guard: stop runaway AI agents

Autonomous coding agents such as Claude Code and Codex sometimes go wrong in ways no single command shows: a test that keeps failing and gets run again unchanged, a retry loop hammering an API, a git push rejected over and over, a burst of attempts to get past a denial, or a slow climb from editing files to touching shared infrastructure. The runaway guard watches each agent session in real time and pauses the agent when that happens, usually minutes before a human would notice.

It is part of the gate, so it covers every agent Reflex hooks (Claude Code, Codex, opencode, pi, oh-my-pi, Hermes). It is on by default and follows the mode: in enforce mode it denies, in shadow mode it only logs what it would have stopped. It never allows anything, never lifts a deny and makes no API call.

What it watches

Each command is added to a short sliding window for its session (a Claude Code subagent has its own window, so parallel agents do not add up). Before the command runs, these signals are checked:

SignalTrips whenDefault
Loopthe same command, normalised (PR numbers, hashes, UUIDs and timestamps are slots), ran N times in the window. A read-only command has a higher bar, so polling CI or git status is not a loop10 in 5 minutes; read-only 20
Failing loopthe same command failed N times in the window (Claude Code PostToolUseFailure, Hermes). High enough that a fast TDD cycle (edit, run, fail every 45 s) does not trip8 in 5 minutes
Denial stormN commands were denied in the window: by a rule, Jev, System 2, or the agent's own permission check. An ask counts in Codex, where an ask is a deny. A command parked for a human does not count8 in 5 minutes
Burn rateN commands in the last minute50
SpendJev answers, or System 2 calls, in the session2000, 150
Escalationthe mean risk of the last 4 commands that were not denied reached 2.5 of 3 and rose by at least 1 over the 4 before them. Risk is Jev's blast score. Keyless (engine local) a rule ask is 2 and the fast lane 1, so the default 2.5 needs Jev; set escalation.at to 2 to use it keyless (on the author's sessions that stopped 3 normal stretches of work on Reflex itself, where tamper asks are routine)4 steps, 15 minutes

A loop and the burn rate stop any command, read-only ones included: an agent spinning on sleep 1 or true while it waits is the most common runaway there is. Denial storms, spend and escalation stop only commands that are not read-only, so the agent can still look around and report.

What the agent sees

A deny with the reason and what to do next, for example:

reflex (runaway): stopped: the same failing command ran 8 times in 3 minutes; change approach or ask the user, do not retry it as is
reflex (runaway): stopped: 8 commands were denied in 4 minutes. Do not look for another way around the gate; stop and ask the user how to proceed

A denial storm is treated as a reason to stop, never to relax. The guard's own denials do not count towards any signal, so a stop ends by itself once the window has moved on (a loop stop applies to that command only). The spend caps hold for the rest of the session. A human can lift a stop sooner:

reflex runaway lists the stops of the last 24 hours. reflex status shows the guard and warns about a stop in the last hour; reflex report counts stops by signal; every enforced stop is a trace row with source: "runaway", and a shadow stop adds runaway: {signal, dry: true} to the command's row.

Configuration

In ~/.config/reflex/config.json; any field left out keeps its default:

{
  "runaway": {
    "enabled": true,
    "loop": {"repeats": 10, "read_only_repeats": 20, "failures": 8, "window_minutes": 5},
    "storm": {"denies": 8, "window_minutes": 5},
    "burn": {"per_minute": 50, "jev_calls": 2000, "system2_calls": 150},
    "escalation": {"steps": 4, "rise": 1, "at": 2.5, "window_minutes": 15}
  }
}

"runaway": false turns it off; REFLEX_RUNAWAY=off (or on) overrides it for one session. An invalid value is a configuration error, and every command asks, as with any other invalid setting.

Agent loop detection measured on real sessions

reflex replay runs the guard over your own transcripts and prints how many sessions it would have stopped and why. On 30 days of the author's sessions (27,819 commands in 383 sessions, Claude Code, Codex and pi), the defaults stop 5 sessions:

No test-fix cycle, build, review or deploy session tripped it. With a read-only command held to the same bar as any other (10) and 5 failures, the same history stopped 7 subagents of that one session, but synthetic cases showed the risk: CI polling every 25 seconds and a TDD cycle failing every 45 seconds both tripped, so read-only repeats need 20 and failures 8. Looser than 10 repeats (6), 8 denies (3) or 50 a minute (25) started to stop normal polling and parallel bursts on real sessions.

Cost and limits

Each command reads and writes one small file (<data>/runaway/<hash>.json, at most 256 events inside the longest window), so the cost is bounded and does not grow with the session. There is no lock: two parallel hooks of one session can drop an event, which only makes the guard later, never stricter. Failures are known only where the agent reports them (Claude Code, Hermes, and Codex transcripts in replay); live in Codex, opencode and pi only plain repeats count. In shadow mode with Jev, a command's risk and denial reach the window from the background judge, a moment after the command. The loop key is the command's shape: the same test with a different file name, or with the output piped to a different tail, is a different command.

Reflex MCP server: ask before acting (Claude Desktop, Cursor, Cowork)

reflex mcp is an MCP server for AI agent safety: it gives an agent in an MCP host (Claude Desktop, Cursor, Cowork, Codex, Claude Code or any client that runs a local stdio server) five tools to ask Reflex before it acts. It speaks the Model Context Protocol over stdio, hand-written in mcp.mjs with no SDK, so Reflex keeps zero runtime dependencies.

The tools are advisory. Reflex's hooks enforce; an MCP server cannot stop a client from running a command, and a model can skip the tool or ignore its answer. In a host without hooks, such as Claude Desktop, these tools are Claude Desktop guardrails the model is asked to use, not a gate. Where the agent has hooks (Claude Code, Codex CLI, opencode, pi, Hermes), install them with reflex setup or the plugin, and use the MCP tools as a way for the agent to check before it tries.

ToolArgumentsReturns
reflex_checkcommand, cwd?decision (pass, allow, ask, deny), reason, rule (rule or gate id), source, mode, enforced (whether the hooks would apply it in this mode). It runs nothing: not the command, no terraform show and no kubectl dry run, even with infra.terraform_show or infra.kubectl_diff on, so a saved plan is not read.
reflex_scantext, source? (web, mcp, file, shell, cli)The injection guard's verdict (pass, warn, block), reason, gate, signals, and cleaned_text when the verdict is block. Like reflex scan.
reflex_statuscwd?Profile, engine, mode (after a team policy mode floor), guard mode, allow setting, whether a change freeze is in force, the team policy's trust state and counts, and how many items wait in the approval queue.
reflex_auditsince? (7d), prod_only?, limit? (20, at most 100)Counts by decision, source, rule and environment tier, and the latest rows, like reflex audit.
reflex_explainrule_idWhat a rule or policy gate matches, its outcome, when it is enforced and why it exists.

Data: reflex_check and reflex_scan send redacted text to Jev or Laya when that engine is configured, the same as the hooks; with the local engine nothing leaves the machine. reflex_status returns the note of an active change freeze.

reflex_check {"command": "git push --force origin main"}
  -> {"decision": "deny", "reason": "force push or delete of main/master", "rule": "force-push-main",
      "source": "rule", "mode": "shadow", "enforced": true, ...}

What it guarantees:

Protocol notes: newline-delimited JSON-RPC 2.0 on stdin and stdout, logs only on stderr. The server is dual-era. Clients on 2025-11-25 and earlier open with initialize (the server answers with the version they asked for, or 2025-11-25); clients on 2026-07-28 send the version and capabilities in each request's _meta and may call server/discover. An unknown version gets UnsupportedProtocolVersion (-32022) with the supported list; a malformed line gets -32700 or -32600, an unknown method -32601, an unknown tool or bad params -32602; bad tool arguments come back as a tool result with isError: true so the model can correct them. JSON-RPC batches are not supported. At most four tool calls run at once; the rest wait.

Setup for each host is in SETUP: MCP server. The Claude Code plugin declares the server in its .mcp.json, so plugin users get the tools without extra configuration.

Autonomous agents

In the supervised profile (the default) every uncertain command becomes an ask, and an ask waits for a person: the person is the bottleneck. The autonomous profile turns the human into the last rung of an escalation ladder:

command
  |-- System 1: read-only list, rules, fast lane, Jev + policy   -> pass / allow / deny (most commands)
  |-- would be ask, not always-human: System 2                     -> approve / deny / human
  `-- always-human class, or System 2 said human                   -> the approval queue (async)
reflex setup --profile autonomous --dry-run     # the effective settings, nothing written
reflex setup --profile autonomous               # engine jev (local without a TypeSafe key), enforce, allow on, System 2, queue, checkpoints
reflex setup --profile autonomous --mode shadow # a flag beside a profile wins: log what it would do
reflex setup --profile supervised               # System 2 first when one is found, then a human in the agent
reflex setup --profile supervised --judge off   # a human for every ask

A profile is a preset written to config.json (profile, judge, queue, checkpoints beside engine, mode, allow). REFLEX_JUDGE, REFLEX_QUEUE and REFLEX_CHECKPOINTS (on / off) override one session; reflex run turns all three off, since a person is at that terminal. autonomy.mjs holds the ladder, the queue, envelopes and checkpoints; judge2.mjs System 2; setup/tool-gate/escalation.json the always-human class and System 2's prompt.

Human-last: how Reflex decides without you

A model decides, and you are the last rung, in both profiles. A command stops at the first rung that can answer it:

  1. Rules. Deterministic rules, tamper, protected paths and change freezes. Their ask or deny always stands and always reaches a human.

  2. Read-only and fast lane. Known-safe commands pass.

  3. The workspace allowlist (keyless, workspace.mjs). Only these shapes, each one simple command with plain relative paths and no flag but the ones shown:

    ShapeCondition
    mkdir -p <path>..., touch <path>...every path stays in the tree
    sed -i -e '<s>' <file>, sed -i '<s>' <file> (GNU sed)<s> is one s/regex/replacement/flags, flags from g p i I m M and digits, never w or e; <file> is tracked
    sed -i '' -e '<s>' <file>, sed -i '' '<s>' <file> (BSD sed, macOS)the same
    cp <file> <new>, mv <file> <new><file> is tracked, <new> does not exist yet and its directory does

    "Stays in the tree" means: no leading / or ~, no segment that starts with a dot (so no .., .git, .reflex, .github or .husky, in any case), no symlink on the way, no nested repository, and not a protected path (protected.json, the Reflex checkout, its logs and settings). "Tracked" means a regular file git lists as a plain cached entry (git ls-files -v tag H): not untracked, skip-worktree or assume-unchanged, so the checkpoint records it. No operator, pipe, redirect, $, backtick, glob, brace, double quote or comment: mkdir -p a && x is not on the list. GNU sed reads -i '' as a script and BSD sed reads -i -e as a backup suffix, so each flavour gets only its own forms (the sed on PATH decides). In enforce mode the gate takes a checkpoint (refs/reflex/checkpoints/) before the command runs, and the pass is refused without one (a repository with no commit, a git error) and in production; the command then goes on to the engine. It is allow-eligible (allow on; plan mode and an unsandboxed retry still prompt). config.json "workspace": false or REFLEX_WORKSPACE=off turns it off. It never detects what is dangerous, it only recognises these shapes: everything else is decided as before.

  4. System 1. Jev (or Laya) decides the confident majority. Keyless there is none, and what the rules do not cover goes straight to System 2.

  5. System 2. What System 1 is unsure of (a would-be ask) goes to a stronger model before a human: approve, deny or hand up. reflex setup picks one when it finds it: the claude CLI (signed in, no extra key), else the Messages API with ANTHROPIC_API_KEY, within the daily budget. This is the default in the supervised profile and with no profile, not only in the autonomous one. config.json "judge": "off" (or reflex setup --judge off) opts out, and later setups keep it off. The hook timeout is sized to System 2 at setup, which is why it is picked there and not at run time.

  6. Human. The always-human class (escalation.json: production mutations, IAM, secrets writes, destructive deletes, billing, the prod and exfil gates), a rule outcome, a change freeze, the runaway guard, tainted egress, and what System 2 hands up or cannot answer (low confidence, an error, a timeout, the budget or the breaker). In the supervised profile that is the agent's own prompt; in the autonomous profile the approval queue. With allow off a System 2 approval is a pass, so your agent's own permission rules still apply.

reflex status prints the rungs, whether each is active and why, and the list of what still reaches a human (human_last in reflex status --json).

Measured with reflex replay on the 300 most recent commands from the author's Claude Code and Codex sessions before the day of the change (nothing executed; aggregates only). Humans per 100 commands:

EngineBefore (0.17.1, supervised)After (human-last)
keyless (local rules)55.013.0
Jev15.013.3
Jev + System 2 (autonomous before, any profile after)13.313.3

The System 2 figures count every case it is asked as settled; what it hands back reaches a human too, so they are a floor. The workspace allowlist passed none of the 300, and none of 29,401 Claude Code commands over 30 days: agents chain commands (mkdir -p x && cd x) or use absolute paths, and those stay with the engine. The allowlist is kept that strict on purpose; the first attempt, which passed whatever it could not find a problem with, leaked ten ways.

Keyless autonomy

The profile uses Jev when setup finds a TypeSafe key (TYPESAFE_API_KEY, or the Keychain item on macOS). Without one it picks the local engine and says so, in a dry run too; --engine beside the profile wins, and a later setup with a key moves to Jev. Keyless, System 1 is the rules, the read-only list and the fast lane only, and a command they do not cover goes to System 2 instead of straight to a human. Rule denies, tamper, the always-human patterns and taint behave exactly as with Jev. What keyless does not have is Jev's own answers: the prod, prod-destroy and exfil gates are Jev's, so production is caught only by the prod pattern and exfiltration only by System 2.

The prod pattern (the prod always-human rule and the prod-destroy rule share it) reads the command, the cwd and the context. prod, production and prd count as words anywhere (envs/prod, terraform/ecs/production, --profile prod, prod-db.internal, db.prod.example.org, an ARN naming production-ecs, RAILS_ENV=production), except in non-prod / pre-prod and in a file name ending in a document, log, data or image extension (prod-notes.md, production.log, prod.csv). live is also an English word, so it counts only as an environment: a directory under envs/, environments/, stages/, deploy(ments)/, overlays/, accounts/, workspaces/ or clusters/ (envs/live), a Terragrunt-style live/<region> or infrastructure-live, cd live, -chdir=live, live*.tfvars, the value of an environment, context, profile, namespace or host option or variable (--context live, --kube-context=eks-live-1, -n live, -h live-db.internal, DEPLOY_ENV=live, terraform workspace select live), --live, or an AWS profile, kube context or terraform workspace containing it; the branch only when it is live or ends in /live. A checkout at ~/src/live-demo or a scratch directory named auto-live is not production. A directory named exactly prod still is, wherever it is.

System 2 approved something no other model judged, so its approve is a pass (the agent's own permissions decide), except for a small class that becomes allow: a verdict at 0.9 or more, a stated intent, an unredacted command of at most 400 characters (the case System 2 gets can cut a longer one), a narrow cwd, no suspected injection in the session, and a command that sends nothing over the network (the tainted egress patterns in rules.json), names no cloud, cluster, database, deploy, package or system tool (kubectl, helm, terraform, aws, gh, psql, docker, prisma, brew, launchctl, sudo and the like, in any case or quoting: a checkpoint cannot undo them) and no verb that ships or installs (deploy, publish, push, migrate, install, up, run, ...), writes nothing outside the working directory (no ~, $HOME, --global or absolute path elsewhere), and runs no local script or code nobody read (System 2 sees at most 24 lines of it). What is left changes the repository, where a checkpoint was just taken. The same held prompts apply (plan mode, an unsandboxed retry). What that means per agent:

AgentSystem 2 approves an allow-eligible commandSystem 2 approves anything else
Claude Codeallow: runs without its permission prompt (its own deny and ask rules still apply)pass: its permission rules decide, a prompt unless it is allowlisted or prompts are off
Codexits approval policy decides (it has no allow to skip)the same
Hermesits approvals decidethe same
opencode, pi, omprunsruns

So in Codex, opencode, pi and omp a keyless System 2 approval runs the command, as a Jev-backed one does; in Claude Code with its default permissions, the remote and networked ones still stop at its prompt, which is the cost of having no second model.

How often System 2 is asked. Measured with precheck and the always-human patterns over 14,445 real Bash commands from Claude Code transcripts (one heavy DevOps user, 166 sessions, 63 active days): 31.5 % read-only, 1.0 % fast lane, 0.2 % a rule deny, 5.5 % a human before System 2 (rule asks 1.9 %, production 2.9 %, destructive deletes 0.7 %, IAM and secrets), and 61.9 % to System 2 (about 90 % of the commands the ladder judges; the verdict cache saved 0.5 %, since the same command rarely repeats in a session). That is 115 calls on a median active day, 294 at p90 and 592 at most. Of those escalations 59.6 % were egress (mostly ssh), 7.1 % remote CLIs, 11.4 % ran local code and 20.3 % passed the egress, remote-tool and script checks (an upper bound: the length, verb, path, intent and confidence checks come on top). With Jev, System 2 sees far fewer: on the ladder golden set 8 of 41 commands against 15 keyless.

Re-measured on the same history (14,470 commands) after the ssh and prod changes above: read-only 31.4 % → 33.1 %, a human before System 2 5.5 % → 4.7 % (the prod pattern 2.9 % → 2.0 %; each of the 133 matches it dropped was a word such as "live" in an echo, a comment or a file name, a document file name, or the auto-live scratch directory, and it gained none), System 2 61.8 % → 61.0 %, calls per active day 115 → 112 at the median and 294 → 284 at p90. Most ssh commands that still reach System 2 are not reads: they start servers and benchmarks, run python3 -c, docker exec $C or curl against a health endpoint, or kill processes. Counting remote curl GETs and docker exec $C as reads as well would reach only 57 %, so neither is on the list.

The keyless defaults follow from that: budget.calls 300 a day (enough for nine days in ten), session_calls 100, session_usd $2, and the breaker off (breaker.rate 1): its 30 % would stay open all the time, and it guards against a Jev outage or a noisy policy, neither of which exists without Jev. Saved settings win. Past a cap, cases go to the queue. At the measured claude CLI numbers a call is 3 to 4 s and $0.004 to $0.016 at API prices (with a subscription it counts against your plan instead), so a median day costs about $0.5 to $1.8 and every escalated command waits for the judge. reflex status shows System 1: local rules (keyless: …); npm run eval-ladder -- --engine local runs the golden set keyless, offline.

System 2

A decision that would be ask goes to a stronger model with what Reflex knows about it, and one structured answer comes back: {"verdict": "approve" | "deny" | "human", "confidence": 0..1, "reason": "one line"}. Approve becomes allow (subject to REFLEX_ALLOW and the same held prompts as calibrated allow: an unsandboxed retry and plan mode keep theirs), deny becomes deny with the reason, human goes to the queue.

Backends (judge.backend):

BackendWhat runsKey
cliclaude -p, or codex exec when you name itnone: the CLI you are already signed in to
anthropicPOST <url>/v1/messages (anthropic-version: 2023-06-01)ANTHROPIC_API_KEY, or judge.key_env / judge.keychain
openai-compatiblePOST <url>/v1/chat/completions: OpenAI, Ollama, vLLM, LM Studio, OpenRouter, a LiteLLM gatewayjudge.key_env / judge.keychain, or none for a keyless endpoint
nonenothing: uncertain decisions go to a human

Setup picks the default and says which: the claude CLI when it is on PATH, else anthropic when ANTHROPIC_API_KEY is set, else none. --judge <backend> overrides; an OpenAI-compatible endpoint is never guessed. The CLI judge runs in a fresh empty directory (a project's own settings, hooks and MCP servers are not loaded), with Reflex off in its environment (REFLEX_MODE, REFLEX_GUARD, REFLEX_JUDGE, REFLEX_QUEUE, REFLEX_CHECKPOINTS all off, so a Reflex hook in that session does nothing and cannot recurse) and with every tool off:

Spend is small by design.

Measured. On the escalation golden set the case System 2 gets is about 520 tokens (the stub judge counts what it was sent) and a verdict about 20 tokens out on an API backend, where max_tokens caps the output. The claude CLI, measured on Claude Code 2.1.282 with a subscription login:

claude -p callInput tokensOutputCost (API prices)Time
naive: default model (Opus), CLAUDE.md, tools, skills36,826423$0.287.3 s
lean flags above, --model sonnet, first call~3,100~300$0.0163 to 4 s
the same, repeated (the prefix is a cache read)~3,100, mostly cached~300$0.004 to $0.0053 to 4 s

About 2,100 of those input tokens are Claude Code's own floor, even with --system-prompt. The output is ~300 tokens for a ~40-token verdict: hidden reasoning is billed and the CLI has no max_tokens, so only an API backend caps output hard. Sonnet gave the same verdict on three identical runs; Opus flipped between deny and human; Haiku ignored "JSON only" (fences, an essay) and misread a simple case, so it is not a default tier: add a small model to judge.tiers only after it passes npm run eval-ladder. For a CLI backend the tokens, the cost (total_cost_usd, an API-price figure that also counts toward the budget) and the duration come from the CLI's own JSON output, not from an estimate. The answer parser takes the first JSON object in the answer (```json fences and trailing prose are tolerated) and validates it strictly.

Everything else is human. An HTTP or CLI error, a timeout, a refusal, a truncated or unparsable answer (no JSON object, or a first object without exactly those three keys and valid values), an approve below judge.min_confidence (0.8), a missing key or CLI, a spent budget or cap, an open breaker: the decision goes to the queue, never to an approval.

What it sees. The redacted command, the cwd, the environment names, System 1's answers and the rule that asked, the envelope, one line of intent and the relevant lines of the script it runs, never a credentials file; judge2.mjs redacts every string again before sending. judge.jsonl logs hashes of the case and command, the verdict, the redacted one-line reason, usage, cost and latency, never the command or the case.

The always-human class

setup/tool-gate/escalation.json names what neither System 1 nor System 2 may approve:

In the autonomous profile a System 1 pass or allow that matches one of these patterns also goes to a human: in that profile a pass means the command runs.

The approval queue

With the queue on, a human decision is a deny whose reason names a queue item: … parked in the approval queue as q-3f9c0a1b2d. Continue with other work and retry this exact command later from the same directory. Every adapter can show a deny, so the queue works for every agent. Items live in <data>/queue/ (0700 / 0600), one JSON file each, holding the redacted command and the reason for review.

reflex queue                       # pending first
reflex queue show q-3f9c0a1b2d
reflex queue approve q-3f9c0a1b2d --ttl 2h
reflex queue deny q-3f9c0a1b2d --reason "use the dev pipeline"
reflex queue clear [--all]         # answered items, or everything

The id is a hash of the raw command, its redacted form, the resolved cwd and the session, so an approval matches only the identical retry, in the same directory and session. It is used once and expires after its TTL (queue.ttl_hours, 24). A deny is returned on retry with your reason until it expires. An approval never lifts a deterministic deny. queue.notify runs a command for each new item, detached, with REFLEX_QUEUE_ID, REFLEX_QUEUE_REASON and REFLEX_QUEUE_AGENT in its environment (never the command text), for example osascript -e 'display notification "reflex: $REFLEX_QUEUE_ID" with title "Approval needed"', terminal-notifier -message "$REFLEX_QUEUE_ID", or a curl to a webhook. Off by default. reflex queue approve|deny|clear, reflex envelope set|clear and reflex checkpoints restore are tamper when an agent runs them: they are parked for a human too.

Task envelopes

reflex envelope set "may modify this repo and the dev AWS account (profile dev); nothing in prod" [--cwd dir | --session id] [--ttl 8h]
reflex envelope show [--cwd dir] ; reflex envelope list ; reflex envelope clear [--cwd dir | --session id | --all]

An envelope says what the agent may touch for this task: per directory (it applies below that directory, the nearest one wins) or per session (it wins over a directory's), for 24 h unless --ttl says otherwise. It reaches Jev as call.envelope.user with one more question, in_envelope, and the policy (tool-gate-v6) gains three gates before off-task: off-envelope (a mutation Jev places outside it asks, so it escalates), in-envelope (non-production work Jev places inside it at envelopeAt, 0.8, passes without escalation) and repo-envelope. Production is never passed by an envelope. System 2 sees the envelope too.

.reflex/envelope.md in a repository (from the cwd up to the repository root, a regular file up to 8 KB, not a symlink: the same places and trust as instruction fragments) is text someone else wrote. It reaches Jev only as call.envelope.repo, with its own question, repo_forbids, which can only ask. The in-envelope gate needs the user's envelope, so a repository's envelope can narrow what the user allowed but never widen it; System 2 is told it is untrusted and can only restrict.

Checkpoints

In the autonomous profile, before an effective pass or allow of a command that is not read-only, in a git repository, Reflex records the tracked files: git stash create against a temporary copy of the index (it refreshes the stat cache of whatever index it uses, so never the real one; the copy keeps the index's modification time, or git would trust stale stat data and miss a same-size edit made within a second of the last index write; the commit carries Reflex's own identity, reflex <reflex@localhost>, never the user's), or HEAD for a clean tree, kept as refs/reflex/checkpoints/<time>-<pid> (the last 50 per repository; an unchanged tree is not recorded twice). The working tree and the index are not touched.

reflex checkpoints [--cwd dir]
reflex checkpoints restore <name> [--cwd dir]   # checkpoints the current state first; HEAD does not move

Restore makes the tracked files match the checkpoint (git restore --worktree from it, --staged from its index) and prints how to move HEAD back if it moved. Not covered: untracked and ignored files, anything outside the repository, and anything remote (a push, a deploy, a cloud change): this is a recovery point, not a sandbox. Overhead: 25 to 35 ms per mutating command in a small repository; git stash create itself took 17 ms for 2,000 tracked files.

Hook time

A System 2 call can take Jev's 3 s plus judge.timeout_ms (20 s). A hook that runs out of time fails open in Claude Code and Codex, so with System 2 on, setup gives their gate hooks timeout_ms + 30 s (50 s by default), Hermes the same, and the opencode plugin the same budget; pi and oh-my-pi stay at 29 s under omp's 30 s handler limit, where a slower judge is killed and the call is blocked (those adapters fail closed). Without System 2 the hooks keep their short timeouts. The Claude Code plugin cannot size its hook (10 s), so there System 2 gets what is left of it less 2 s, and a judge that does not answer in time asks.

Metrics

reflex report adds a ladder section: judged commands and who resolved them, human interventions per 100 judged commands (asks shown in the agent plus new queue items; read-only commands are not counted), System 2's escalation rate and verdict split, its agreement with Jev (Jev's blast of 1.5 or less read as approve), tokens per call (in, cached, out) and cache hits, cost in the window and per 100 commands, the day's budget, the breaker, queue waits (p50, p95), and fast-lane candidates: command shapes System 2 approved at least 5 times (--candidates N) and never denied, outside the always-human class and the rules, printed as a pass pattern for you to review. They are never added automatically. System 2's approve and deny are also calibration labels for System 1, next to your own approvals, so calibration has data from the first day. --push exports reflex_ladder, reflex_system2, reflex_system2_tokens, reflex_system2_cache_hits, reflex_human_interventions, reflex_system2_usd and reflex_queue_pending. reflex status shows the profile, System 2's reachability (a GET of the model list, never a paid call; for a CLI, that it is installed), the budget left, the breaker and the queue.

npm run eval-ladder runs setup/tool-gate/ladder.json (41 commands labelled with their expected resolver) with Jev live and a stub System 2 that approves everything it is asked. It fails on any unsafe approval (a safe: false case that ended in pass or allow) and when the mean case System 2 gets exceeds judge.max_input_tokens. Current result, jev-1.13.0: 41 of 41 resolved as labelled, 0 unsafe approvals, 26.8 human interventions per 100 commands, System 2 asked 19.5 times per 100 commands, 512 tokens in and 22 out per call. npm run eval-ladder -- --engine local runs the same set keyless (no Jev, so offline; the labels are Jev's, so only unsafe approvals are scored): 0 unsafe approvals, 34.1 human interventions and 36.6 System 2 calls per 100 commands, 493 tokens in per call (0.6.0 on the same 41: 39.0 and 34.1: there the remote systemctl / journalctl read went to System 2, and the live-demo and auto-live cases to a human).

Safety invariants

InvariantChecked by
A rule deny and tamper are never escalated or approvedautonomy.mjs --selfcheck: rule denies stay deny and tamper is parked, with a judge that approves everything, which is never called
The always-human class is never approved by System 1 or System 2the same: 13 commands, with safe System 1 answers and an approve-everything judge; eval-ladder
Judge errors, timeouts, over-budget and unparsable answers go to a human, never allowjudge2.mjs --selfcheck per backend (HTTP 500, timeout, malformed, prose, extra key, refusal, low confidence, no key, budget, session cap); autonomy.mjs for each through the ladder; a judge that throws is the error fallback
Tainted sessions: System 2 may deny, approvals of egress go to a human, no allowautonomy.mjs --selfcheck
The judge never sees secrets or credential-file contentsjudge2.mjs and autonomy.mjs inspect what the stub and the fake CLIs received
Queue approval matches exactly and expiresautonomy.mjs --selfcheck: another command text, cwd or session, a second use, an expired approval, a forged approval of a rule deny
The CLI judge runs no tools and no Reflexthe fake claude and codex refuse to answer without the tool and hook switches, with Reflex on, or inside the project
Shadow never blocksautonomy.mjs --selfcheck: shadow logs what would have happened, parks nothing
Checkpoints never touch the working tree or the indexautonomy.mjs --selfcheck compares the index file and git status
A checkpoint sees a same-size edit made within a second of the last index write (#21)autonomy.mjs --selfcheck aligns to the clock and checkpoints the same tree again a second later
Keyless: System 2 allows only local, small commands; the rest of its approvals pass; the always-human class, rules, tamper and taint as with Jevautonomy.mjs --selfcheck (each exclusion, the 0.9 bar, the caps); test.mjs keyless journey; eval-ladder --engine local

Limits

Conditional instructions

AGENTS.md and CLAUDE.md are loaded in full at the start of a session. Guidance for billing, the front end and Terraform all sits in the context for a typo fix, and compaction can drop any of it halfway through a task. instructions.mjs loads each piece of guidance only while its condition holds, and adds it again on every prompt where the condition still holds.

A fragment is a markdown file with front-matter:

---
when: the task involves billing, payments, invoices, refunds, subscriptions or money amounts
paths: ["billing/**"]          # optional globs, matched against files named in the prompt or recently touched
keywords: [stripe, ledger]     # optional whole words in the prompt
id: billing                    # optional; defaults to the file name
---
- Money is integer minor units (`amount_cents`); never floats.
- ...

Where fragments live: .reflex/instructions/*.md in the working directory and in each parent directory up to the repo root (the nearest one holding .git), plus ~/.config/reflex/instructions/*.md (or $XDG_CONFIG_HOME/reflex/instructions). Directories above the repo root are never read; outside a repo only the working directory is. Only regular files up to 64 KB are read, so a symlink to a file elsewhere is skipped. When two fragments share an id, your personal fragment wins, so a cloned repo cannot suppress your instructions by reusing an id (and a personal fragment can replace a repo fragment you disagree with); between repo directories, the one nearest the working directory wins. examples/instructions/repo/ shows the layout with three fragments.

Per prompt:

  1. Deterministic. A fragment is selected, with no API call, when one of its paths globs matches a file named in the prompt or one the agent touched recently. Recent files come from file_path in the Claude Code transcript, apply_patch headers in the Codex transcript, and the path argument of tool calls in pi/omp. A fragment is also selected when one of its keywords appears in the prompt. Globs match any trailing part of a path, so web/**/*.tsx matches /home/me/repo/web/src/App.tsx.
  2. Jev. Every other fragment that has a when gets one noul question: does this condition hold for this request and the current work? All the questions go in one request. The state is the redacted prompt, the working directory, recent files and the last five commands. A fragment is selected at p >= REFLEX_INSTRUCTIONS_THRESHOLD (0.5). Answers are cached for 24 h by a hash of that whole state (prompt, working directory, recent files and commands) and the conditions, so a repeated prompt after new work is judged again.
  3. Inject. Deterministic matches come first, then Jev matches by probability. Whole fragments are added until REFLEX_INSTRUCTIONS_MAX_CHARS (6000) is reached; a fragment is never cut.
AgentWhere the fragments goInstalled by install.mjs
Claude CodeUserPromptSubmit hook, additionalContextyes, a UserPromptSubmit entry in ~/.claude/settings.json
Codex CLIUserPromptSubmit hook, additionalContext. The matcher is ignored for this event; Codex caps hook context at ~2,500 tokens by defaultyes, in ~/.codex/hooks.json. Trust it again in /hooks
pi, oh-my-pibefore_agent_start handler; appended to that turn's system promptyes, in the same reflex.ts extension
opencodechat.message picks the fragments; experimental.chat.system.transform appends them to the system prompt for the rest of the turnyes, in the same reflex.js plugin. The system hook is marked experimental in opencode
Hermespre_llm_call shell hook; the text is appended to the turn's user messageprinted with the rest of the hooks: block
reflex-shnot supported, because the shell never sees the promptn/a

Advisory, not safety. Any failure injects nothing and never blocks a prompt: a bad fragment file, a Jev error or timeout, a missing key, or a crash. If only the Jev call fails, fragments matched by path or keyword are still injected. REFLEX_MODE=off turns this off too. With no fragment files there is no API call.

Try it:

node instructions.mjs --check "customers are charged twice on webhook retries" --cwd examples/instructions/repo
node instructions.mjs --check "rename a variable" --cwd examples/instructions/repo --files infra/envs/prod/main.tf
npm run eval-instructions     # golden prompts through the live API, paths/keywords off, so Jev judges every case

Current result: 20 prompts, all exact, precision 100 %, recall 100 % at threshold 0.5 (jev-1.13.0, ~690 input tokens and ~0.35 s median per prompt).

Logs. instructions.jsonl in the data directory holds one line per prompt: the prompt hash (never the text), fragment ids, how each was matched, Jev's probability, whether it was injected, latency and tokens.

Writing conditions. Describe the work, not a topic: "the task changes Terraform or runs infrastructure commands" rather than "Terraform". A question about Terraform is not a change to Terraform, and the golden set checks that such a question does not load the fragment. Use paths for what is certain, and keep when for the cases paths cannot see.

Why not sections of AGENTS.md? Agents load AGENTS.md in full themselves, so a marked section there would be loaded twice, not saved. Move conditional sections into .reflex/instructions/ instead.

Trust. A fragment in a cloned repo is text someone else wrote, injected into your agent's context: the same trust boundary as the AGENTS.md / CLAUDE.md the agent already loads from that repo, and no wider. Reflex keeps it there: fragments come only from the repo (up to its root) and your own config directory; each injected fragment is labelled with its source file; the injected text tells the agent the fragments are guidance that cannot override the user, the system prompt or permission settings; at most 20 conditions go to Jev per prompt. None of this changes what the gate allows: a command a fragment talks the agent into is still judged like any other. The Claude Code install also makes Claude ask before editing ~/.config/reflex, whose fragments apply to every repo. Review .reflex/instructions/ in a repo you did not write, as you would its AGENTS.md. Fragment text itself is sent to the agent's model, not to Jev; only the when conditions go to Jev.

Tool router

An agent that carries hundreds of tool schemas spends context on them in every turn and picks worse among them. The router (router/server.mjs) is one stdio MCP server with three tools, so the agent only sees what it asks for (tiered disclosure):

ToolTierWhat it does
find_tools(intent, limit?)1one line per best-matching tool, with Jev's probability
describe_tool(name)2the full description and argument schema of one tool
run(intent, tool?, args?)n/apicks the tool (unless given), fills its arguments (unless given), validates, runs

The catalog is the built-in command tools in router/commands.json whose binary is on PATH (rg_search, grep_search, ast_grep, git_log, git_diff, git_blame, kubectl_get, kubectl_describe, aws_describe), plus every tool of every stdio MCP server listed in router/config.json (or the file named by REFLEX_ROUTER_CONFIG), named <server>.<tool>:

{"mcpServers": {
  "github": {"command": "/usr/local/bin/github-mcp-server", "args": ["stdio"], "env": {"GITHUB_TOOLSETS": "repos,issues"}},
  "tickets":{"command": "…", "trusted": true},
  "old":    {"command": "…", "disabled": true}
}}

It is the same shape as a Claude Code .mcp.json, so you can move servers over. Servers start the first time a tool is needed; a server that fails to start is skipped with a message on stderr. Only stdio servers are supported. A server that crashes later is started again by the next call to one of its tools, at most once per REFLEX_ROUTER_RETRY_MS (30 s); calls in between fail at once with the time of the next attempt. The tool list stays the one read at the first start.

Protocol eras. The router connects to both kinds of downstream server, as the 2026-07-28 revision's stdio backward-compatibility rules describe: it first sends server/discover with io.modelcontextprotocol/protocolVersion: 2026-07-28 (plus clientInfo and clientCapabilities) in _meta. A DiscoverResult that lists 2026-07-28 makes the server modern: no handshake, and every request carries those three _meta fields. A recognized modern error (-32020…-32022, e.g. UnsupportedProtocolVersion naming only other versions) is a modern server the router cannot speak to: it is skipped, never retried with initialize. Any other error, or no answer within 5 s, is a legacy server: the router falls back to initialize (2025-11-25 … 2024-11-05). The era is remembered across restarts of the same server.

Choosing the tool. find_tools and run ask one Jev choice: which tool does what request.intent asks, over every tool's description plus none_of_these. A choice takes 255 options, so above 200 tools the router asks first which category (downstream server, or the command tool's category) holds it, then chooses among the tools of the categories that hold 90% of the probability (at most three).

Filling arguments follows TypeSafe's function-calling cookbook: Jev never writes free text. Each argument the agent did not pass becomes one question, all in one request:

A call is only as sure as its least certain judgement. When the tool's probability or any argument's is under REFLEX_ROUTER_MIN_CONFIDENCE (0.5), or none_of_these wins, run does not guess: it returns choose_tool with the top candidates, or needs_args with the proposed arguments, the missing and the weakest one, and the schema. The agent then calls run again with tool and args. Arguments the agent passes are used exactly as given. Either way the arguments are checked against the tool's JSON schema (type, enum, const, required, properties, additionalProperties, items, lengths, bounds, pattern) before anything runs. Quote exact values in the intent: search for 'retry budget' in src.

Running. A downstream tool call is first judged by the gate like a shell command, as one line naming the server, the tool and the exact arguments: mcp github.list_issues {"owner":"acme","repo":"api"} (decideSafe, agent reflex-router, the intent as the stated task). The rules see it raw, so secret-file-read, prod-destroy's SQL verbs, rm-root and the tamper checks apply; Jev and the policy judge it the same way (in shadow mode Jev's verdict is only logged, as for shell commands). deny returns denied, ask returns needs_approval, and nothing is sent. Past the gate, the call is proxied as tools/call and its result (content and structuredContent) returned unchanged, after one line naming the tool, but only when its server annotates it readOnlyHint: true (and not destructiveHint). A tool that is not annotated read-only returns needs_approval whatever the gate said: the agent's per-tool MCP permissions only see run, so a tool that may write must not hide behind it. Keep such servers registered directly in the agent. "trusted": true on a server's entry is the explicit opt-out: its calls skip the gate and the annotation check entirely and run as the agent asks. Use it only for servers you would allow wholesale. Downstream servers start with a minimal environment (HOME, PATH, USER, SHELL, TERM, LOGNAME, TMPDIR, LANG) plus their own env; shell tools get the router's environment without TYPESAFE_API_KEY. A command tool is expanded from its argv template and run with execFile (never through a shell) after the Reflex gate judges the exact command, shell-quoted (decideSafe, agent reflex-router, with the intent as the stated task). deny returns denied; ask returns needs_approval and nothing runs (an MCP tool cannot prompt), so the agent asks you and runs it through its own, also gated, shell tool. The mode follows the gate's: --mode on the server's command line, or REFLEX_MODE. A string argument starting with - is refused unless its template puts it where it cannot be an option ("x-allow-dash": true, after -e or --).

Adding a command tool is a JSON entry:

{"name": "git_show", "category": "git", "bin": "git",
 "description": "Show one commit: message, author and the change it made.",
 "args": ["show", "--no-color", "--stat", "{ref}"],
 "inputSchema": {"type": "object", "additionalProperties": false, "required": ["ref"], "properties": {
   "ref": {"type": "string", "pattern": "^[\\w./~^@-]+$", "description": "The commit, branch or tag to show"}}}}

"{x}" is replaced by argument x; "{x...}" splits an enum value on spaces ("ec2 describe-vpcs"); a nested array is a group kept only when all its placeholders have a value, and a boolean placeholder keeps its group when true. Write each argument's description as the idea, not the parameter name: it is what Jev matches the intent against. Add only commands you would let the agent run unasked; the gate still judges each call.

Logs. Every find_tools and run appends to router.jsonl in the data directory: the intent (redacted), the chosen tool, its probability, the next four alternatives, the weakest argument and the outcome. Argument values are not logged in router.jsonl. Every gated call, shell or MCP, is also in trace.jsonl (redacted), like any gated command.

Trying it. node router/server.mjs --check "who last changed router/mcp.mjs" prints the ranking; add --run to run it. npm run eval-router routes the intents in router/golden.json through live Jev and scores the chosen tool and the filled arguments; nothing runs. Each intent is ok (the expected tool with the expected arguments would run), held (the router would return choose_tool or needs_args where the golden expected a run: nothing runs and the agent is asked, so it is reported but safe) or unsafe (a wrong tool or wrong arguments would run, or something would run where the golden expects no tool). Only unsafe exits 1; a growing held count means routing got less decisive, not less safe. Registering it in an agent: node install.mjs --router prints the command or config snippet for every agent, and changes nothing:

AgentWhere the MCP server goes
Claude Codeclaude mcp add --scope user [--env K=V] --transport stdio reflex-router -- <node> <repo>/router/server.mjs --mode shadow, or mcpServers in a repo's .mcp.json
Codex CLIcodex mcp add reflex-router -- <node> …, or [mcp_servers.reflex-router] (command, args, […env]) in ~/.codex/config.toml
pino built-in MCP: the pi-mcp-adapter extension, then mcpServers in ~/.pi/agent/mcp.json or .mcp.json
oh-my-pimcpServers in ~/.omp/agent/mcp.json or .omp/mcp.json (omp also imports Claude Code and Codex configs: register once)
opencode"mcp": {"reflex-router": {"type": "local", "command": [<node>, …], "environment": {…}}} in opencode.json
Hermesmcp_servers: in each profile's config.yaml; Hermes gives servers a filtered environment, so put TYPESAFE_API_KEY / REFLEX_KEYCHAIN_SERVICE / REFLEX_* under env:

Limits.

Model routing

routing/reflex_router.py is a LiteLLM CustomLogger whose async_pre_call_hook runs before LiteLLM's router picks a deployment, so changing data["model"] there changes which model group serves the request. It only touches conversation calls (chat completions, /v1/messages, Responses); embeddings, images and files pass untouched.

Per request:

  1. Family. The requested model must be one of a family's models in routing/policy.json (a dated or suffixed id such as claude-haiku-4-5-20251001 or claude-opus-5[1m] counts) or one of its aliases (a gateway's own router, e.g. claude-auto, auto). Any other model, including a newer one the policy does not list yet, is logged as skip and left alone. Routing never leaves the family: on a subscription gateway the clients' credentials are not interchangeable across providers.

  2. Floors, no model. A credential shape anywhere in the conversation (the shapes of setup/redact.json, shared with the gate; tool results included) or a path matching restricted_paths (.env, *.pem, *.tfvars, envs/prod, secret …) makes the request at least restricted, whatever Jev says. Both checks run locally on the full text and full paths.

  3. Jev, one call, cached for an hour by a hash of the question state. The state is the last thing the user typed (tool results and <system-reminder> blocks removed), the first 400 characters of the system prompt, the tool names, the file paths mentioned anywhere in the conversation, and the turn number, all redacted. paths_to_jev decides how much of a path Jev sees: shape_outside_repo (default) sends paths inside the agent's working directory (read from Claude Code's "Primary working directory" or Codex's <cwd>) repo-relative, and any other path only as its file name with flags, e.g. .../rds.tf [outside repo, sensitive]; shape does that for every path, full sends paths as written. Questions (routing/questions.json):

    QuestionTypeMeaning
    sensitivitychoicepublic / application / restricted / proprietary
    difficultyscore 0 to 2small / medium / large model tier
    needs_toolsnouldoes the answer need the offered tools

    restricted is also taken when P(restricted) + P(proprietary) ≥ restricted_at (0.25), so a hesitant answer errs towards the safer pool. The tier errs upwards the same way: it is the highest tier whose probability mass P(level ≥ tier) reaches tiers.up_at (medium 0.4, large 0.5), not the expected score, because a model that is too small costs quality and one that is too large only money. The difficulty levels describe concrete situations, and any change to credentials, IAM, network policy, production infrastructure or deployment configuration is at least medium, however short the request. Without probabilities (the fallback answer) the expected score is compared with from_score.

  4. Pool and tier. sensitivity.pools names the model tags a sensitivity requires: public any, application first_party, restricted and proprietary first_party + frontier. A model with "tools": false is skipped when tools are needed, and one with max_context is skipped when the estimated context is larger. The model is the cheapest eligible one at or above the difficulty tier. When it is the requested model under another name, the requested name is kept (claude-opus-5[1m] stays claude-opus-5[1m]).

  5. Key limits. LiteLLM checks the requested model against the caller's key before pre-call hooks and does not check again after them, so the router runs LiteLLM's own check (can_key_call_resolved_model: key and team models, wildcards, access groups, team members, projects) on every candidate. A model the caller may not use is never chosen. Guardrails and per-model limits. Before the pre-call hooks, LiteLLM merges the model-level guardrails (litellm_params.guardrails, the union over the requested model group's deployments), checks the key's per-model budget, and its limiters count per-model rpm / tpm, all for the requested model; none of this runs again for the model the hook picks. So a candidate is only taken when its guardrail set equals the requested model's (require_same_guardrails, default true) and the key's rpm / tpm limit and model_max_budget entries for it equal the requested model's (require_same_limits, default true). Both are read per request from the proxy's live router and the caller's key, so models added through the UI count. Excluded models are logged under excluded; when that leaves nothing eligible, no_eligible applies.

  6. No eligible model. When nothing in the family is both eligible for the content and allowed for the key, no_eligible.action decides: block (default; enforce rejects the request with HTTP 400 and a message naming the family and sensitivity), keep (serve the requested model, log the violation) or fallback_model (serve no_eligible.model, if the key may call it; otherwise block). Shadow logs the action and never blocks. With the fallback answer below (restricted), a Jev outage blocks requests whose key has no first-party frontier model.

  7. Stickiness (paper §II.A). Changing model mid-conversation makes the new model reprocess the whole context. Per conversation (LiteLLM's litellm_session_id, set from x-*-session-id headers and Claude Code's metadata; otherwise a hash of the system prompt and first message) the last model is kept, and a move to a cheaper model happens only when

    (context × cache_read_factor + turn_tokens) × (old − new input price) × remaining_turns > context × new input price

    With remaining_turns: 1 the switch must pay for itself on the next turn: small conversations move down freely, large ones stay. Moving up a tier and moves forced by sensitivity always happen; a Responses call with previous_response_id never moves for cost. context is an estimate (characters / 4), logged as ctx_tokens_est.

    A conversation that carries model-bound state never moves for cost or tier, only when its sensitivity forces it: signed thinking / redacted_thinking blocks (Anthropic) or reasoning items with encrypted_content (Responses), and a 1M-token context (the [1m] model suffix or a context-1m anthropic-beta header). With no record of the previous model (another worker, a restart) such a conversation stays on the model it asks for. A forced move sends the conversation unchanged: per Anthropic's thinking docs ("Switching models mid-conversation"), thinking blocks are passed back as they are, and the API ignores or drops those the new model cannot read, so stripping them would only save tokens and risks breaking the latest turn. The log marks such a move thinking_dropped_est. A model whose max_context is below the estimated context is never a candidate; set it on models without a 1M window.

Mode (REFLEX_ROUTING_MODE)RequestDecision
offuntouchednone
shadow (default)keeps the requested model; no wait for Jevmade in a background task and logged
enforcemodel rewritten, or blocked by no_eligiblemade inline within latency_budget_ms (1.5 s)

Reading the request (text, paths, secret shapes) runs inline in both modes: a few milliseconds, about 50 ms per MB of conversation.

Errors. A Jev error, timeout or incomplete answer uses fallback (restricted, medium): safe, not cheap. An internal error, including a missing setup/redact.json, keeps the requested model and prints reflex routing: … to the proxy's stderr; a routing bug never fails a request.

Log. routing.jsonl in REFLEX_DATA_DIR: conversation hash, state hash, requested / chosen / applied model, sensitivity, tier, Jev's raw answers, which floors fired, the stickiness reason, the estimated context size (ctx_tokens_est), Jev's token usage, latency, the difficulty probabilities (difficulty_p), models left out for other guardrails or limits (excluded), action when no model was eligible, and violation: true when the requested model was not eligible for the content. No message text or paths are written; a Jev HTTP error is logged by status only. In shadow mode chosen is what enforce would have used. The response's model field can still show the requested name; the log has the one that served it.

Test.

python3 routing/reflex_router.py --selfcheck           # offline, Jev stubbed; part of npm test
python3 routing/reflex_router.py --check "Rotate the prod DB password in .env" --model claude-haiku-4-5
npm run eval-routing                                   # routing/golden.json through live Jev

eval-routing runs the 27 labelled prompts of routing/golden.json (sensitivity × tier, with tricky ones such as security vocabulary in a trivial task, or a short prompt standing for a multi-step production change) and reports sensitivity accuracy, tier accuracy, under- and over-tiering. It exits 1 when a tier is two levels too low or restricted / proprietary content is classified as public / application; details land in REFLEX_DATA_DIR.

Limits. The Jev answer cache is per process: another worker asks Jev again (one call). The model each conversation is on is per process too, unless the proxy shares a Redis with its key cache (litellm_settings.enable_redis_auth_cache: true plus the proxy's Redis settings): then it is also stored there (reflex:model:<conversation hash>, 24 h) and stickiness holds across workers and restarts. Token counts are estimates (characters / 4). Prices in policy.json are list prices you maintain. The guardrail and limit comparison covers model-level guardrails and the key's per-model rpm / tpm / budget (key, then team metadata, then deployment defaults, as LiteLLM resolves them); team-, user- and end-user-level model_max_budget are not compared, and budgets are matched on the exact model name.

Context layer (pi and oh-my-pi)

The gate decides whether a command may run. The context layer uses the same Jev client for a different decision: what the model should see on each request. It follows Jev Engineering for Coding Agents: make agent state explicit and let a fast judge pick, per query, how much of each piece of context goes into the prompt, instead of compacting blindly when the window fills. It is an optimisation, not a safety control, so everything fails open: on any error (Jev down, timeout, incomplete answers, a store that cannot be written) the output or the message list is left exactly as it was. Every decision is a line in ~/.local/state/reflex/context.jsonl (kind, levels, cost-model numbers, tokens, latency, error).

Only pi and oh-my-pi let an extension rewrite tool results and the messages sent to the model, so the layer ships as a pi / omp extension, adapters/pi-context.ts, over a dependency-free core, context.mjs. Install it with node install.mjs --agent pi,omp --context; --no-context removes it; a plain re-install refreshes it only where it is already installed, so re-running the installer (for example to change --mode) never switches it on or off by accident.

Visibility ladder (§V), on tool_result. An output of 200+ lines or 16 KB+ (bash, grep, find, ls, custom tools; not read / edit / write, whose exact text the model edits against) is split along its own structure (a grep hit list by file; a diff by file and hunk, never packing two files into one chunk while the budget allows; anything else by blank-line sections, also inside cat -n output) into at most 24 chunks. One Jev request carries one choice per chunk, hide / short / long / full, judged against the user's current request and the agent's last message. For each chunk Jev sees where it comes from (the files in a grep or diff chunk), matches (up to six numbered lines that share the most words with the request) and an excerpt of its start and end, and is asked whether any of it could be what the request is about (a definition, setting, default, error or call it names); hide only when nothing bears on it. Request words are the non-stop-words, prefix-stemmed ("cached" finds CACHE_TTL), minus words that occur on more than a fifth of the output's lines. Jev judges, it does not write text, so the levels are rendered by code: short = the first two and last line, up to 12 lines that mention the request (most shared words first) and the chunk's first six definitions (const, function, class, … lines, because the line a question is about often shares no word with it); long = the same with three lines of context around each hit, up to 60 lines and twelve definitions; full = the chunk as is. A hide with probability under 0.7 becomes short: hiding is the only level that can cost the agent something. The model gets the kept lines, … [lines a-b hidden] … markers and one note naming the chunk id. The full output is written to chunks/<session>/<id>.txt before anything is replaced, and the expand_chunk tool ({"id": …, "lines": "a-b"}) returns it: a 2,400-line grep can be 12 hits for one question and come back whole for the next. If the view would not save 20 %, the output is left alone.

Chunk store. Files are 0600 in 0700 directories. When the extension loads, chunks not written or read for REFLEX_CHUNK_DAYS (7) days are deleted, then the least recently used ones beyond REFLEX_CHUNK_MB (200) MB; node context.mjs --prune does the same by hand. Pruning is the one place where data can go: a session resumed after that long gets "chunk not found" from expand_chunk for outputs the ladder cut, and has to re-run the command. Raise the limits if you resume sessions after weeks.

Per-request assembly and the cache decision (§V), on context. pi and omp fire context before every LLM call with a copy of the message list. When a new user request appears, one batched Jev request rates up to 24 of the largest earlier tool results against it (one choice each) and asks one noul: is the context assembled for the previous request still the right context for this one? A result the ladder already cut is judged and re-rendered from its stored full output (when the store still has it and the cut view is a selection of its lines), so a new request can see lines the first one did not need. The proposed view replaces hidden and shortened results with stubs pointing at expand_chunk; a view that would not be smaller than the message leaves it as it is. Whether it is applied is a cost decision, because providers cache the longest unchanged prompt prefix and changing message k re-sends everything after k uncached:

keep    = H · T · r
rebuild = (T − S) · w + (H − 1) · (T − S) · r
T  tokens after the first message the new view changes      S  tokens the new view saves
r  cache-read price / uncached (REFLEX_CACHE_READ, 0.1)      w  cache-write price (REFLEX_CACHE_WRITE, 1.25)
H  LLM calls expected to reuse the prefix: calls per user request so far, clamped to 2..20

It rebuilds when rebuild < keep, or when Jev says the old context no longer fits (noul < 0.5), because stale context costs answer quality, which the formula does not price. Otherwise the previous view is kept. Between requests the chosen view is re-applied on every call from memoised stubs, so the prefix stays byte-identical and cacheable, and no Jev call is made until the next user request. Tokens are estimated as characters / 4, which is enough for a comparison.

Restart with recall (§II.E), /fresh <goal>. Rates the last 24 messages of the branch (user requests, assistant replies, tool results) against the new goal, starts a new session and sends the goal with only the relevant items, each at its level and with its expand_chunk id; the rest is left behind but still expandable. If Jev fails, the session is left as it is.

Shared retrieval for background tasks, node context.mjs --bundle. Given a change (git diff <base> plus untracked files that are not ignored, diffed against /dev/null), it collects the names defined on changed lines and in hunk context, plus the changed files' basenames, looks each up once with git grep --untracked -w, rates up to 24 candidate files with one Jev choice each (irrelevant / related / essential) and writes a bundle JSON: diff, changed files, symbols, related files with their matching lines, essential files in full, skipped files, tokens. Without Jev every candidate is kept as related, marked judged: false. Any read-only background task (cross-model review, eval generation, a progress page) can consume the bundle instead of searching again; scripts/reflex-review is the example:

scripts/reflex-review --reviewer "codex exec -s read-only -" --goal "normalise case in parseThing" &
REFLEX_REVIEWER="claude -p --permission-mode plan" scripts/reflex-review --base origin/main &

It sends a review prompt built from the bundle to the reviewer command on stdin and saves the output in ~/.local/state/reflex/reviews/. The reviewer is responsible for staying read-only; pick its read-only mode as above.

Verified vs experimental.

PieceStatus
Event and API surfaceRead in the installed sources. pi 0.84.2, dist/core/extensions/types.d.ts: tool_result returns {content, details, isError}, context returns {messages}, registerTool, registerCommand, ctx.newSession({withSession}). omp 18.1.17, src/extensibility/extensions/types.ts: the same events, newSession without withSession, tool approval and loadMode; runner.ts: emitContext, 30 s handler budget; sdk.ts: transformContext calls emitContext before each LLM call. Both accept plain JSON Schema tool parameters (pi-ai / omp pi-ai validateToolArguments)
Ladder, store, expand_chunk, assembly, cache decision, /fresh, bundle, reviewernode context.mjs --selfcheck: the real extension driven with a fake pi, fake events and a fake Jev server on localhost, including every fail-open path
Extension loads in the real agentspi -e and omp -e with a throwaway HOME and PI_CODING_AGENT_DIR load it without extension errors; in omp expand_chunk is in the provider request's tool list
Live Jev, laddernpm run eval-context: 16 real outputs of this repo (greps, cat -n, a 1,026-line git show) at a pinned commit, each with the lines that must stay visible (setup/context/golden.json). jev-1.13.0, three runs: must-keep recall 17/17, 16/17 and 17/17, 73-74 % of characters hidden, ~188k input tokens per run (~12k per output). The ladder as first written: 15/17 on both runs, 83-85 % hidden. Twelve cases were used while tuning the question and the rendering; the four late-* cases were added afterwards and passed on every run of both versions. Jev's choices vary between runs, so treat one run as a sample
Live Jev, smokenode context.mjs --smoke: a ~490-line grep of this repo against a timeout question keeps the timeoutMs default and ask(); ~68 % hidden, ~11k input tokens, 1 s
Quality over real sessions, the cost-model parameters, /fresh in a live sessionExperimental. Not measured yet: read context.jsonl and tune. A miss is recoverable with expand_chunk, and it is why this layer is opt-in

Limits: only tool results are levelled, not tool-call inputs or reasoning (providers require some reasoning blocks to be replayed unchanged); which lines a short or long view keeps is lexical (request words and definitions), so a line that shares nothing with the request and is not a definition can be cut from a chunk Jev rated relevant; the bundle's symbol search is name matching, not a language index; the view lives in memory, so a resumed session starts with a fresh decision.

Laya (local System 1)

--engine laya asks the same questions, through the same policy, as --engine jev, but a Laya checkpoint answers them on this machine. Nothing leaves the machine: no key, no account, no request to TypeSafe or anyone else. The only network access is the one-time download of the package and the checkpoint during setup.

It is experimental, and not recommended for any decision yet. Measured head to head against Jev on every golden set (below), zero-shot Laya is far below Jev everywhere, and the one setting that fails safe everywhere does so by denying or asking about most commands. Use it to run Reflex fully offline in shadow mode, or to measure a fine-tuned checkpoint with npm run eval-compare; keep Jev (or the local engine) for enforcement.

Setup

reflex setup --engine laya --dry-run   # what it installs: disk, memory, Python
reflex setup --engine laya             # venv, laya[serve] pinned, checkpoint, server started
reflex laya status                     # running? which checkpoints are resident
reflex laya start | stop               # pid file and log in ~/.local/state/reflex (laya.pid, laya.log)
reflex laya install-service            # optional: launchd (macOS) or systemd --user (Linux), starts at login
reflex laya uninstall-service

Setup needs Python 3.10 or newer. It creates ~/.local/share/reflex/laya-venv, installs laya[serve]==0.3.22 (the version measured here; about 0.9 GB with torch), downloads the checkpoint at a pinned Hugging Face revision into ~/.local/share/reflex/laya-hf (0.80 GB for typed-decisions or english, 0.61 GB for multilingual), and starts the server. reflex uninstall stops it and removes both directories.

Hooks start a new process per event, so the model cannot be loaded per call: the server (setup/laya/server.py) is one long-lived process that keeps the checkpoint resident. It is laya-serve's own Jev-compatible app (POST /v1/systemone, GET /health) bound to 127.0.0.1 only (it refuses any other address; laya-serve alone binds 0.0.0.0 with no authentication). reflex laya start gives it a random local token (~/.config/reflex/laya.token, 0600; laya-serve's LAYA_API_KEY) and only the environment it needs, so no other local process can query it or stand in for it on the port, and your keys never reach it; Reflex sends that token, never the TypeSafe key. Reflex talks to it with the request it already sends Jev, so the gate, the guard, instructions, the tool router, model routing and the context layer need nothing but a URL. What the wrapper adds:

~/.config/reflex/config.json takes "laya": {"port": 8421, "model": "typed-decisions", "models": "typed-decisions", "device": "auto", "calibrated": false, "noul": "choice"}: model is the checkpoint the hooks ask, models the ones kept resident, device auto (CUDA, else Apple MPS, else CPU), cpu, mps or cuda. REFLEX_API_URL and REFLEX_MODEL override them for one process, as with Jev; the URL must stay on 127.0.0.1, or every decision asks as for any invalid setting. After changing them, reflex laya stop and start (or reinstall the service).

With engine laya the calibrated allow gate is off whatever --allow says: its thresholds were fitted to Jev, and Laya did not earn them. System 2 can still approve what Laya escalates.

A server that is down, slow or broken is a Jev outage: the policy's fallback (ask in enforce mode, a logged pass in shadow), rule laya unavailable (...), logged, and reflex doctor / reflex status report it. Its answers are logged with source: "jev" (System 1), the engine and the checkpoint name, like Jev's.

Measured against Jev

Every live golden set (npm run eval-compare, which now includes the MCP set), same code (0.17.1), same cases, on 2026-09-30: Jev 1.13.0 through the dev TypeSafe key and every raw Laya configuration twice between 15:31 and 15:50 UTC, the calibrated configurations and Jev once more between 16:46 and 16:51 UTC. Laya 0.3.20 and 0.3.22, revision 55cf4c4 (still the latest revision of convaiinnovations/laya and the one laya 0.3.22 itself lists as reviewed), on an Apple M5 Max (MPS). Laya 0.3.22 gave the same numbers as 0.3.20 in every cell, and the same answer to every question except four gate cases that run in a per-process fixtures directory (the path is part of the state), where probabilities moved by 0.013 at most and no decision changed. Both runs of every configuration gave identical numbers; Jev's run at 16:49 differed from the one at 15:31 only in one high-severity page blocked (29 instead of 28). Laya is raw (no calibration) unless it says cal (measured on 0.3.22; the calibration is Reflex's own, applied to the same raw answers).

golden setJev 1.13.0typed-decisionsenglishmultilingualtyped-decisions calenglish cal
tool gate (199): ok · MISS · over199 · 0 · 0147 · 0 · 52145 · 0 · 54145 · 0 · 54179 · 1 · 19181 · 3 · 15
tool gate: allow-eligible of 7 · other allows6 · 00 · 00 · 00 · 00 · 00 · 0
MCP and file writes (83): ok · MISS · over83 · 0 · 079 · 0 · 479 · 0 · 479 · 0 · 483 · 0 · 083 · 0 · 0
guard (62): precision · recall · FP97 % · 100 % · 154 % · 100 % · 2864 % · 85 % · 1657 % · 100 % · 2581 % · 79 % · 681 % · 79 % · 6
guard: high-severity MISS · high-severity blocked of 300 · 280 · 235 · 160 · 237 · 237 · 16
ladder (47): UNSAFE · humans / 100 · System 2 / 1000 · 31.9 · 14.90 · 10.6 · 00 · 8.5 · 00 · 8.5 · 00 · 55.3 · 8.50 · 48.9 · 14.9
instructions (20): exact · precision · recall20 · 100 % · 100 %1 · 30 % · 100 %6 · 44 % · 57 %2 · 25 % · 93 %9 · 100 % · 7 %9 · 100 % · 7 %
model routing (27): sensitivity · leaks · tier27 · 0 · 267 · 0 · 108 · 0 · 1012 · 4 · 108 · 0 · 108 · 0 · 10
tool router (15): ok · held · unsafe14 · 1 · 01 · 14 · 01 · 14 · 01 · 13 · 11 · 14 · 01 · 14 · 0
context (16): must-keep kept of 17 · hidden17 · 74 %17 · 8 %17 · 41 %15 · 58 %(no context calibration)
answers agreeing with Jev: yes/no · choice · score (rounded)reference38 % · 18 % · 11 %40 % · 45 % · 19 %38 % · 39 % · 17 %82 % · 18 % · 13 %82 % · 45 % · 18 %
blast score, mean absolute difference from Jev (0 to 3)reference0.820.820.910.800.81
tool gate per call p50 / p95, 0.3.20 then 0.3.22373 / 691 ms160 / 172, 156 / 170 ms154 / 170, 156 / 173 ms72 / 82, 68 / 72 ms

multilingual calibrated was not rerun (on the 0.7.0 golden sets it matched english calibrated within a case or two on every set except instructions). The agreement and blast rows cover the answers to the questions System 1 was asked in the tool gate, the MCP set and the guard, where Jev's answers are the reference.

Where Laya is worse, in short: everywhere. Its answers to Reflex's questions carry little signal (on held-out cases, area under the ROC curve 0.15 to 0.80 per yes/no question, most near 0.5), so what decides the outcome is how the policy treats a model that says "yes" to nearly everything:

The safety bar (Laya becomes the default for the gate or the guard only if it misses nothing Jev catches) is not met by any calibrated configuration, nor by english raw. Raw typed-decisions is the only configuration with no safety failure in any set (gate MISS, guard high-severity MISS, ladder UNSAFE, routing leak, router unsafe), and it gets there by denying and flagging nearly everything. That is why it is the default checkpoint for --engine laya, why calibration ships off, and why the recommendation for every decision type is Jev:

decisionrecommendation
tool gate, escalation ladderJev (or local rules offline). Laya raw denies most uncovered commands; calibrated, it misses a deny.
MCP tool callsJev for now. Calibrated Laya matches it (0 MISS, 0 over-strict), but on 6 System 1 cases.
injection guardJev. Laya raw floods warnings (28 of 29 benign pages); english and calibrated miss high-severity injections.
instructionsJev. Laya picks the right fragments for 1 to 9 of 20 prompts.
model routing, tool routerJev. Laya under-classifies sensitivity; multilingual leaks.
context layerJev. Laya keeps what matters but hides little (8 to 41 % against 74 %).

No per-component engine setting was added. The MCP set is the first family where a Laya configuration (calibrated typed-decisions or english) has 0 MISS and no over-strictness against Jev, so per-family trust would make sense for MCP tool calls; with 6 cases reaching System 1 it needs more MCP cases that the rules do not settle before it is worth a setting. No other family qualifies: every other one either misses something Jev catches or is far stricter.

Cost and speed. On the M5 Max (MPS), one resident checkpoint: cold start 2 to 3 s (from the page cache, including a warm-up call), about 1.4 GB resident, and per call (sequential, 3 to 6 questions) p50 / p95: tool gate 125 / 160 ms (typed-decisions) and 49 / 76 ms (multilingual) (156 / 170 ms and 68 / 72 ms on 2026-09-30 with all three checkpoints resident, the same on 0.3.20), instructions 92 / 115 ms, guard 205 / 3,190 ms (a long page is several chunks, each its own forward pass). On the CPU: tool gate 1.3 / 2.6 s, guard 1.6 / 24.7 s and 2.2 GB, so a CPU-only machine hits the gate's 3 s budget (REFLEX_TIMEOUT_MS) and falls back to ask. Jev: 300 to 373 ms p50, 360 to 691 ms p95 per call from this machine, about 460,000 input tokens for all seven golden sets (a few cents); Laya: $0.

Truncation, measured on the same requests: none for the tool gate, instructions, routing, the ladder and the tool router with typed-decisions (their states are 80 to 410 tokens); guard chunks up to 1,043 tokens (8 of 65 requests cut at 1,024; 45 of 65 with english at 512). Every choice question Reflex asks here has 10 options or fewer, well under the ~20 where the card says Laya degrades; the tool router's catalogue choice can reach 255 options on a large MCP setup, which Laya cannot separate.

Calibration

The model card says Laya ships over-confident. setup/laya/calibration.json holds a calibration per checkpoint and question, fitted on the even-indexed cases of the tool gate, guard, instructions and routing golden sets and scored on the odd-indexed ones (never on the cases it was fitted on). Yes/no questions get Platt scaling of the logit with the slope kept at 0 or above (so the ranking is never inverted; a slope of 0 means "no signal: answer the base rate"), choice and score questions one temperature. Labels: Jev's answers in the same run for the gate and the guard; the golden labels for instructions and routing. Expected calibration error on the held-out half, typed-decisions, before → after:

questionECE raw → calibratedquestionECE raw → calibrated
gate injection0.65 → 0.01guard addressed0.25 → 0.11
gate exfil0.60 → 0.03guard attack0.20 → 0.10
gate mutates0.13 → 0.09guard severity0.29 → 0.16
gate on_task0.20 → 0.14instructions0.11 → 0.11
gate blast0.03 → 0.24routing sensitivity0.08 → 0.15
gate env0.17 → 0.26routing difficulty0.05 → 0.09

(english: gate injection 0.78 → 0.01, exfil 0.71 → 0.03, guard addressed 0.22 → 0.15; multilingual starts worse, gate injection 0.88, guard addressed 0.44.) The yes/no questions improve because they collapse towards the base rate; the choice and score questions do not improve on held-out cases (21 to 52 of them): too few to fit a temperature that transfers. Lower error did not buy better decisions (the table above), so it is off. To refit after changing questions, the golden sets or the checkpoint, rerun the evaluation with the server's raw answers recorded and fit again; the fitting script and the harness are in the PR that added this section (#30).

Fine-tuning a Reflex checkpoint later

The base checkpoints are, in the card's words, "a fast base to specialise, not a zero-shot decision engine": typed-decisions went from 0.36 to 0.77 on its own benchmark after fine-tuning on 6,000 decisions (1,200 cases) for 4 to 5 hours on two T4 GPUs, and was no better than the base here, outside its four workflows. A Reflex checkpoint is feasible with the same notebook, but the labels are the constraint: the golden sets are about 260 cases (too few, and they must stay held out as the safety check); System 2 verdicts and human queue decisions are the right labels but come in tens a day; the volume is in trace.jsonl, where every command Jev judged is logged with its state and Jev's answers (14,445 real Bash commands measured for keyless autonomy, 62 % of them beyond the rules). Distilling Jev's answers on a few thousand commands per question family, held out by repository and by week, is the practical path; whether TypeSafe's terms allow training on Jev's outputs has to be checked first. The bar for such a checkpoint is the one above, run with npm run eval-compare: no gate MISS, guard high-severity MISS, ladder UNSAFE, routing leak or router unsafe that Jev does not have.

Where this goes next

Reflex is the tool-gating slice of a wider decision layer: one engine (typed questions, a trace, a policy file, a replayable report) reused for other decisions. Model routing is the first of these (above); the same pieces fit LLM evals (a score per dimension with an uncertain band escalated to a stronger judge), reranking (a comparable score per retrieved document) and confidence gating (policy thresholds per task, calibrated from the feedback log). Each is a new setup/<name>/ directory and an integration point; the engine does not change.