Claude Code guardrails

Claude Code guardrails and a spend cap, in CI and headless

Auto mode already screens destructive commands in an interactive session, and this page says so before it says anything else. What it does not do is cap spend, and it is not there when the session did not start in auto mode: claude -p, the Agent SDK and the cloud-provider deployments start in the permissive default mode. Hooks are loaded in all of them. Two commands arm 44 deterministic rules plus a dollar cap that refuses the next tool call once the day is over budget, with no package, no account and no server.

Last updated

What auto mode already does, and where it does not run

Start with the part that is not ours. On Pro, Max and Team plans the built-in starting permission mode for an interactive session is auto mode, and Anthropic's own documentation describes what it does:

Claude Code also runs git status itself before a command that would discard uncommitted work, such as git reset --hard or rm -rf, and shows the classifier whether staged, modified, or untracked work is present.
Claude Code never lets a permissions.allow rule or a PreToolUse hook that returns "allow" approve an rm or rmdir command that targets a critical path.

The 2.1.183 changelog, dated 20 June 2026, says the same thing about git: destructive git commands (git reset --hard, git checkout -- ., git clean -fd, git stash drop) are blocked when you did not ask to discard local work, and terraform destroy, pulumi destroy and cdk destroy are blocked unless you asked for the specific stack. If an interactive session is the only place you run an agent, you already have most of what a destructive-command guard sells, for free, and this page is not going to pretend otherwise.

Two things are still open, and they are what the rest of this page is about.

One. The classifier is auto mode's, and plenty of runs do not start in auto mode. Anthropic lists claude -p and the Agent SDK under the permissive default mode, along with Amazon Bedrock, Google Cloud's Agent Platform, Microsoft Foundry and Claude Platform on AWS. The headless documentation puts it more bluntly: "For -p, the built-in starting permission mode is Manual on every plan." The documented CI story is a static allowlist, --allowedTools "Bash(npm test)", with no target resolution and no record. Hooks, meanwhile, are read in all of those: "Without --bare, a -p session runs the hooks in a project's .claude/settings.json", and the Agent SDK runs settings-file hooks by default through setting_sources. A rule you wrote is what is left deciding there.

Two. No classifier is a cost control. A -p run reports total_cost_usd when it finishes. That is the invoice, discovered afterwards.

Sources, read 16 September 2026: code.claude.com/docs/en/permission-modes, code.claude.com/docs/en/headless, code.claude.com/docs/en/agent-sdk/hooks, code.claude.com/docs/en/hooks. A starting mode is a starting point; anyone can switch modes mid-session.

A day cap, enforced at the tool boundary

The PreToolUse payload carries transcript_path, and the transcript carries each assistant message's model and token usage. So the hook can price the session it is sitting inside. Provenrail reads the transcript from where it last stopped, prices every new message, adds the delta to a ledger shared across hook processes, and refuses the next tool call once the day is over budget.

pr guard budget 25          # or /guard-budget 25 with no CLI installed
pr guard status             # what is armed, and what has been spent today

There is no default cap. A cap you did not choose is a claim about your money, and the CLI refuses to invent one. What a refusal looks like, verbatim from the engine:

denied  budget.day

Provenrail guardrail budget.day: estimated day spend is $25.5000, over the
$25.0000 cap, so this tool call is refused. Stop here and tell the person what
happened rather than working around it. Raise or remove the cap in
.provenrail.json to continue. Spend is estimated at API list price, and
notional on Pro and Max flat-rate plans.
Read this before you trust the number. The figure is estimated at API list price, and on Pro and Max flat-rate plans it is notional, because there is no per-token charge to cap. The host writes the transcript asynchronously, so the cap stops the agent within a turn rather than at the exact dollar. A model with no verified rate is reported as unpriced rather than counted as zero, and an unreadable transcript says so once a day instead of reading as nothing spent. A cap that quietly fails to bind is worse than no cap.

Crossing the warning line (80% of the cap by default) journals a warning and lets the call through, so the agent gets a chance to wind down rather than stopping mid-edit.

What actually goes wrong

Most guardrail tools, including an earlier version of this one, screen the verb. That is the wrong half of the command. Nobody has ever lost work to the letters rm -rf; they lost it to what came after them. rm -rf .next comes back in twelve seconds. rm -rf ~/ is the end of a laptop. The two are indistinguishable to a pattern, because the difference is not in the text, it is in where the text points.

So this guard resolves the target. A delete inside your repository, under a temp directory, in a package cache or in a build directory runs without a word. A delete of a home directory, a whole disk, a two-segment system path, or a variable that becomes / when it is unset is refused. A delete that merely leaves the project asks.

The same idea runs the git rules. git reset --hard on a tree with nothing uncommitted and nothing unpushed destroys nothing at all, so it is allowed in silence. The rule runs git status --porcelain, checks for commits your remote has not seen, and asks only when there is work here that exists nowhere else.

We measured both halves against one frozen corpus: 36,977 Bash commands pulled from 1,247 real agent sessions, replayed offline with the directory each one actually ran in. The rules that shipped in 0.3.1 interrupted 923 of them and refused 305, and reading the samples was the whole argument for this release: a build directory being cleaned, a coverage file being deleted, a Tailwind class read as a SQL TRUNCATE. The rules in 0.4.0 interrupt 222 and refuse 16, and fifteen of the sixteen refusals are this repository's own adversarial test fixtures, files that contain rm -rf / as a test payload. On 40 commands taken from public data-loss reports, 0.3.1 allowed all 40 and 0.4.0 allows none. You can run the same measurement over your own transcripts: python tools/measure_guard.py, which makes no network call and prints counts and samples to your terminal. Widening the hook to every tool rather than a named list of seven brings the corpus to 56,761 calls of every kind, where 0.4.0 interrupts 273 and refuses 21.

Install

Two commands inside Claude Code, and that is the whole setup:

/plugin marketplace add pofky/provenrail
/plugin install provenrail-guard@provenrail

No package to install, no account, no server, nothing to configure. The plugin carries its own dependency-free engine, so 45 rules are armed on the next tool call your agent makes. Installing a plugin called "guard" is the opt-in; a guard that waits for a config file protects nobody.

No signup, no API key, no network call in the decision path. The verdict is computed locally before anything is sent anywhere, so a guardrail never waits on a server and never fails because one is down.

/guard-status   what is armed, and what it has actually stopped
/guard-card     a summary of what it stopped, safe to paste anywhere
/guard-rules    every rule, including the packs that are off by default

/guard-status leads with what the guard has stopped rather than with what it is configured to watch, because a guardrail that has never fired and one that is silently broken look identical from the outside.

/guard-card prints the same history in a form you can post. What makes it postable is what it leaves out: every command is reduced to its verb and its flags and every operand is dropped, so git reset --hard origin/main becomes git reset --hard and an environment assignment carrying an API key becomes export. The directory is a truncated hash rather than a name. We checked the reduction against all 36,977 commands in the corpus: no path, hostname, URL or key survives it.

Turning it up, or off

Write a .provenrail.json at your repo root. It wins over the defaults completely:

{"policy": {"use": ["git-worktree", "destructive", "database", "cloud",
                    "secrets", "production", "access", "money"]}}

An empty list arms nothing and says so once a day, so a disarmed guard can never quietly look like a working one. You can name a single rule id instead of a whole pack, and add your own rules with a regex.

Making the record into evidence

The plugin keeps a local history of every block and every prompt. That is enough for you to see what happened and not enough to show anybody else, because you could have written it yourself. Installing the CLI upgrades the same file, in the same directory, with no migration:

uv tool install provenrail   # or: pip install provenrail
pr guard receipt             # a signed, hash-chained export of what was blocked
pr verify guard-receipt.json

Prefer not to install a plugin at all? pr guard install writes the same hooks into the project's .claude/settings.json directly, preserving any hooks already there.

Headless and CI

This is the case the plugin install does not cover on its own, and the one the classifier is absent from. pr guard install is the whole setup: Anthropic documents that "without --bare, a -p session runs the hooks in a project's .claude/settings.json", and the Agent SDK loads the same settings hooks by default. Commit that file and every pipeline run decides against the same rules as your laptop.

uv tool install provenrail
pr guard install
pr guard budget 25
claude -p "run the migration" --allowedTools Bash --output-format json

--bare skips settings discovery, and Anthropic says it "will become the default for -p in a future release". When that lands, pass the settings file explicitly with --settings rather than relying on discovery.

What is blocked

The pack that has no equivalent anywhere else: fan-out. Claude Code refuses a subagent at depth 3 of 3, which limits how deep the tree goes. Nothing limits how wide it goes, and width is what the public reports are about: issue #68619 (open) records 1.2 million tokens in about 30 minutes with CLAUDE_CODE_FORK_SUBAGENT=0 ignored. Provenrail counts the spawns and asks you before the twenty-first. The cap comes from a measured distribution of 915 real sessions, not from a round number: it reaches 1.6% of all sessions and still catches every runaway in that corpus. Run pr report --fanout to see your own.

Eight rule packs are armed by default with no configuration (fan-out, git-worktree, destructive, database, cloud, secrets, production, access), which is 45 rules. 55 rules across eleven packs are available. Every row below is asserted in the test suite against the default install, so this table cannot drift away from what the code does:

PackBlocksExample
git-worktreeThrowing away uncommitted or unpushed work (asks, and only when there is work to lose)git reset --hard, git checkout -- ., git restore ., git clean -fd, git stash drop, git branch -D, git worktree remove --force
git-worktreeRemoving the way backgit reflog expire, git filter-branch, git filter-repo
git-worktreeDeleting a remote branch, or a destination directory (asks)git push --delete, git push origin :main, rsync --delete, Remove-Item -Recurse
destructiveA recursive delete of something that cannot be regeneratedrm -rf ~/, rm -rf /usr/local, rm -rf $UNSET/, rm --no-preserve-root
destructiveA delete that leaves the project (asks)rm -rf ../other-checkout
destructiveRaw disk writes and irreversible infrastructure teardowndd of=/dev/sda, terraform destroy, kubectl delete namespace, git push --force
destructiveSchema and data loss in SQLDROP TABLE, DROP DATABASE, TRUNCATE, DELETE with no WHERE
databaseDropping and recreating a database in one command (asks)prisma migrate reset, supabase db reset, rails db:drop, artisan migrate:fresh, manage.py flush, alembic downgrade base, dropdb, FLUSHALL
cloudDeleting managed storage, a database or a project (asks)aws rds delete-db-instance, aws s3 rb --force, gcloud sql instances delete, az group delete, fly volumes destroy, wrangler d1 delete, heroku pg:reset
cloudTearing down infrastructure or the state that describes it (asks)pulumi destroy, cdk destroy, terraform state rm, helm uninstall, kubectl delete pvc, docker compose down -v
secretsAPI keys, tokens and private keys appearing in a tool callsk-proj-, sk-ant-, ghp_, xox*-, AKIA, AIza, a JWT, a BEGIN PRIVATE KEY block
secretsReading this machine's own credentials (asks)~/.ssh/id_*, ~/.aws/credentials, ~/.kube/config, .netrc, .npmrc, a keychain file. Public keys are excluded
accessEditing the guardrail's own configuration (asks)a write to .provenrail.json or .claude/settings.json
destructiveAn MCP tool whose method deletes or destroys (asks)mcp__railway__deleteVolume, mcp__aws__delete_bucket
secretsReading a .env file (asks, does not deny)any path ending .env or .env.local. .env.example and the other template spellings are excluded
accessPermission changes that expose files, or weaken authenticationchmod 777, chmod a+rwx, disabling MFA
productionA connection string pointed at production (asks)postgres://...db.prod...
productionA deployment (asks)wrangler deploy, vercel deploy, kubectl --context production

Rules are declared in a .provenrail.json file in the repo, so the policy is reviewable in a pull request like any other code. Nothing is "suspicious by default": if a rule did not declare it, it is not flagged.

Why some things ask instead of deny

This is the part most guardrail tools get wrong. A tool that hard-blocks everything risky gets uninstalled by lunchtime, because half of what it blocks is the work.

Provenrail rules carry an effect. deny stops the call and tells the agent which rule fired. require_oversight turns the call into a Claude Code permission prompt instead, and records your answer as human oversight in the log. Reading .env, running a migration, deploying: these should involve a human, not be impossible.

The distinction matters after the fact too. "The agent was blocked" and "a human approved this" are different facts, and a log that flattens them into one is not evidence of anything.

The receipt

Every decision, allowed, denied, or escalated, is Ed25519 signed and hash-chained to the one before it. Export it:

pr guard receipt
pr verify guard-receipt.json

Change one byte of that file and the verifier exits non-zero and names the broken link. Two independent implementations, a Python CLI and a browser verifier, are held in lockstep by frozen conformance vectors, so verification does not depend on running our code or trusting our servers. You can watch it happen in your own browser on a real record and on a tampered one.

Limits, stated plainly

Guardrails reduce the blast radius of a mistake. They are not a security boundary against an adversary with shell access, and nobody should sell them as one.

Questions

Does this slow the agent down?
The decision is a local pattern match with no network call, so it costs milliseconds. Recording is written to a local journal and shipped asynchronously.
Does it work in claude -p and in CI?
Yes, and that is the case it exists for. A -p session runs the hooks in the project's .claude/settings.json unless you pass --bare, and the Agent SDK loads the same settings hooks by default, so pr guard install plus a committed settings file is the setup. The same engine also runs as a Python SDK if you want to wrap tool calls in a loop of your own.
Is it free?
Everything on this page is MIT licensed and free, including for commercial use: the rules, the spend cap, the journal, the signed local receipts and both verifiers. The one paid option is $9 a month for unlimited independent RFC 3161 timestamps on the root of your receipts, which is the part you cannot issue to yourself.
Does anything leave my machine?
Not with pr quickstart. That configures a local sink and no account. Sending records off-box is a separate, explicit step, and it is the step that makes them tamper-evident to a third party.
Can the agent turn the guardrail off?
The hook runs outside the model's context and the model does not get a vote on the verdict. It can, like any process with write access, edit the settings file, and that edit is itself a tool call the guardrail sees.
← Back to Provenrail