Reflex

Reflex vs Claude Code permissions and auto mode

Claude Code's permission rules match tools and command prefixes, and its permission modes decide how much it asks. Reflex is a hook that judges each shell command by what it does and where it points. By default it only emits ask or deny and leaves pass to your permission settings, so your allowlist keeps working. It gates shell commands, not file edits or MCP calls.

Install: npx @ursuciprian/reflex setup starts with local rules in shadow mode, no account or key. Plugins for Claude Code, Codex CLI and opencode: setup guide.

How is Reflex different from Claude Code permission prompts and allowlists?

Claude Code's permission rules match tools and command prefixes; Reflex judges each shell command by what it does and where it points. It reads the local script, make target or package script a command runs, and knows the AWS profile, kube context, Terraform workspace and git branch. By default Reflex only emits ask or deny and leaves pass to your permission settings, so your allowlist keeps working. Claude Code's own rules cover file edits, web fetches and MCP tools, which Reflex does not gate.

See: compared with other AI coding agent guardrails.

Can I use Reflex with --dangerously-skip-permissions?

Yes: with prompts turned off, Reflex's rule denies still apply. In the autonomous profile, a command that needs a human is parked in the approval queue and returned to the agent as a deny with a queue id, so the run continues while the command waits for you. Only shell commands (and subagent spawns) are gated, so file edits and MCP calls run unchecked in that mode.

See: comparison table, GUIDE: the approval queue.

Do I still need a devcontainer or a sandbox if I use Reflex?

Keep one if you have one: a sandbox limits what any command can reach, Reflex decides per command, and the two work together. A container still holds whatever cloud credentials, kube configs and SSH keys are mounted into it, and those reach production from inside; Reflex judges the commands that use them and scans what the agent reads. Keep IAM, network controls and least-privilege credentials as well.

See: comparison table, GUIDE: safety properties and limits.

Compared with other AI coding agent guardrails

Reflex is a hook, not a sandbox. It decides per command, using what the command is and where it points; a sandbox limits what any command can reach. The two work together.

ApproachWhat it doesWhere it is better than ReflexWhat Reflex adds
Claude Code permission prompts and allowlists (allow / ask / deny rules)Prefix and pattern rules per tool, a prompt for everything elseBuilt in, no latency, covers file edits, web fetches and MCP tools, which Reflex does not gateJudges commands the rules do not list, reads the scripts they run, knows the AWS profile and kube context. Reflex only tightens by default, so your rules keep working.
--dangerously-skip-permissions / YOLO modeNo prompts at allFastest, no interruptionsRule denies still apply, and in the autonomous profile the approval queue parks what needs a human (returned to the agent as a deny, so the run continues)
Container or devcontainer sandboxIsolates the filesystem, processes and optionally the networkHard OS-level containment of local damage, whatever the commandMounted cloud credentials, kube configs and SSH keys still reach production from inside a container. Reflex judges those commands, and scans what the agent reads.
Codex sandbox modes (read-only, workspace-write, danger-full-access) and approval policiesOS sandbox for the commands Codex runs, with network off by default in workspace-writeEnforced by the OS; no pattern can be bypassed by an unusual shell constructContext-aware blocking inside workspace-write or danger-full-access (cloud profile, kube context, the scripts a command runs). It cannot approve anything: Codex's approval policy still decides. Codex hooks cannot show a prompt, so a Reflex ask blocks and the human runs the command with reflex run.
abideEnforces your AGENTS.md / project rules on each edit and on the turn's diff, using JevChecks code the agent writes against your conventions, which Reflex does not doComplementary: abide checks edits after they happen; Reflex gates shell commands and tool results before execution. Both can run on the same agent.
Generic LLM-as-judge hooksSend each command to a general LLM for a verdictAny model, free-form reasoning, simple to writeRules, the read-only list and the fast lane settle about half of commands (51 % on one engineer's week of Claude Code in the replay above) with no API call; the rest cost one typed Jev request (about 1k tokens); a stronger model is asked only on escalation, with budgets, caps and a cache; a policy file makes decisions replayable and tunable.

For infra work, the difference by capability:

CapabilityBuilt-in agent permissions (Claude Code, Codex)Container or OS sandboxReflex
Plan-aware terraformA prefix rule can ask on terraform apply; the plan is not readNot in scopeAsks for a saved plan; with infra.terraform_show and a plugin cache, denies a plan that deletes or replaces
Production contextRules match the command text onlyLimits what a command can reach, not which account or cluster it targetsAWS profile, kube context, Terraform workspace, git branch and prod paths in every decision
Change freezeNot built inNot built inTime or date windows that ask or deny production commands, from config.json or the team policy
Audit exportEach agent's own logs and telemetry, in its own formatNot in scopereflex audit: one csv, json or jsonl row per decision for SOC 2 and ISO 27001 evidence, plus a webhook
Cross-agent policyEach agent's own settings files, in that agent's formatPer container imageOne committed .reflex/policy.json applied in Claude Code, Codex CLI, opencode, pi and Hermes

Keep IAM, network controls and least-privilege credentials, and use Reflex for the decisions a sandbox cannot make.