Reflex vs Claude Code permissions and auto mode
Claude Code's permission rules match tools and command prefixes, and its permission modes decide how much it asks. Reflex is a hook that judges each shell command by what it does and where it points. By default it only emits ask or deny and leaves pass to your permission settings, so your allowlist keeps working. It gates shell commands, not file edits or MCP calls.
Install: npx @ursuciprian/reflex setup starts with local rules in shadow mode, no account or key.
Plugins for Claude Code, Codex CLI and opencode: setup guide.
How is Reflex different from Claude Code permission prompts and allowlists?
Claude Code's permission rules match tools and command prefixes; Reflex judges each shell command
by what it does and where it points. It reads the local script, make target or package script a
command runs, and knows the AWS profile, kube context, Terraform workspace and git branch. By
default Reflex only emits ask or deny and leaves pass to your permission settings, so your
allowlist keeps working. Claude Code's own rules cover file edits, web fetches and MCP tools, which
Reflex does not gate.
See: compared with other AI coding agent guardrails.
Can I use Reflex with --dangerously-skip-permissions?
Yes: with prompts turned off, Reflex's rule denies still apply. In the autonomous profile, a command that needs a human is parked in the approval queue and returned to the agent as a deny with a queue id, so the run continues while the command waits for you. Only shell commands (and subagent spawns) are gated, so file edits and MCP calls run unchecked in that mode.
See: comparison table, GUIDE: the approval queue.
Do I still need a devcontainer or a sandbox if I use Reflex?
Keep one if you have one: a sandbox limits what any command can reach, Reflex decides per command, and the two work together. A container still holds whatever cloud credentials, kube configs and SSH keys are mounted into it, and those reach production from inside; Reflex judges the commands that use them and scans what the agent reads. Keep IAM, network controls and least-privilege credentials as well.
See: comparison table, GUIDE: safety properties and limits.
Compared with other AI coding agent guardrails
Reflex is a hook, not a sandbox. It decides per command, using what the command is and where it points; a sandbox limits what any command can reach. The two work together.
| Approach | What it does | Where it is better than Reflex | What Reflex adds |
|---|---|---|---|
Claude Code permission prompts and allowlists (allow / ask / deny rules) | Prefix and pattern rules per tool, a prompt for everything else | Built in, no latency, covers file edits, web fetches and MCP tools, which Reflex does not gate | Judges commands the rules do not list, reads the scripts they run, knows the AWS profile and kube context. Reflex only tightens by default, so your rules keep working. |
--dangerously-skip-permissions / YOLO mode | No prompts at all | Fastest, no interruptions | Rule denies still apply, and in the autonomous profile the approval queue parks what needs a human (returned to the agent as a deny, so the run continues) |
| Container or devcontainer sandbox | Isolates the filesystem, processes and optionally the network | Hard OS-level containment of local damage, whatever the command | Mounted cloud credentials, kube configs and SSH keys still reach production from inside a container. Reflex judges those commands, and scans what the agent reads. |
Codex sandbox modes (read-only, workspace-write, danger-full-access) and approval policies | OS sandbox for the commands Codex runs, with network off by default in workspace-write | Enforced by the OS; no pattern can be bypassed by an unusual shell construct | Context-aware blocking inside workspace-write or danger-full-access (cloud profile, kube context, the scripts a command runs). It cannot approve anything: Codex's approval policy still decides. Codex hooks cannot show a prompt, so a Reflex ask blocks and the human runs the command with reflex run. |
| abide | Enforces your AGENTS.md / project rules on each edit and on the turn's diff, using Jev | Checks code the agent writes against your conventions, which Reflex does not do | Complementary: abide checks edits after they happen; Reflex gates shell commands and tool results before execution. Both can run on the same agent. |
| Generic LLM-as-judge hooks | Send each command to a general LLM for a verdict | Any model, free-form reasoning, simple to write | Rules, the read-only list and the fast lane settle about half of commands (51 % on one engineer's week of Claude Code in the replay above) with no API call; the rest cost one typed Jev request (about 1k tokens); a stronger model is asked only on escalation, with budgets, caps and a cache; a policy file makes decisions replayable and tunable. |
For infra work, the difference by capability:
| Capability | Built-in agent permissions (Claude Code, Codex) | Container or OS sandbox | Reflex |
|---|---|---|---|
| Plan-aware terraform | A prefix rule can ask on terraform apply; the plan is not read | Not in scope | Asks for a saved plan; with infra.terraform_show and a plugin cache, denies a plan that deletes or replaces |
| Production context | Rules match the command text only | Limits what a command can reach, not which account or cluster it targets | AWS profile, kube context, Terraform workspace, git branch and prod paths in every decision |
| Change freeze | Not built in | Not built in | Time or date windows that ask or deny production commands, from config.json or the team policy |
| Audit export | Each agent's own logs and telemetry, in its own format | Not in scope | reflex audit: one csv, json or jsonl row per decision for SOC 2 and ISO 27001 evidence, plus a webhook |
| Cross-agent policy | Each agent's own settings files, in that agent's format | Per container image | One committed .reflex/policy.json applied in Claude Code, Codex CLI, opencode, pi and Hermes |
Keep IAM, network controls and least-privilege credentials, and use Reflex for the decisions a sandbox cannot make.