agent-guardrail
A deterministic policy firewall for AI agent tool calls. YAML rules, no LLM calls, no risk scores. CLI + MCP server (advisory) + @enforce() decorator (unbypassable). MIT, 46 tests, zero network dependency.
Documentation
Guardrail
A policy firewall for AI agent tool calls.
Your agent wants to run a shell command, send an email, or move money.
Guardrail checks that request against rules you wrote, before it happens,
and either lets it through, asks a human, or blocks it — with a plain-
English reason every time.
60-second quickstart
git clone && cd agent-guardrail
pip install -r requirements.txt
python3 cli.py check --agent trading-agent-001 --tool wallet.transfer \
--args '{"amount": 9999, "to": "0xabc"}'Or `pip install guardrail-mcp` gives you a `guardrail`
command directly — same output, no repo checkout required (falls back to
the policy bundled in the package if you don't point `--policy` at your
own file):
guardrail check --agent trading-agent-001 --tool wallet.transfer \
--args '{"amount": 9999, "to": "0xabc"}'{
"decision": "BLOCK",
"matched_rules": [
{"rule": "numeric_cap_exceeded", "severity": "BLOCK",
"message": "amount=9999.0 exceeds cap 5 for 'wallet.transfer' (unknown agent)"}
]
}That's it — no server, no account, no API key. `policies/default.yaml` is
the file that decided this; open it and change the numbers to match your
own rules.
Why this, not another "AI risk scoring" tool
Most "AI agent security" projects (including an earlier project of mine)
lean on statistical risk scores computed from data nobody can actually
verify at build time — wallet age, "reputation," contract "risk" — which
either requires paid data feeds you don't have yet, or quietly becomes
mock data pretending to be real. Fine for prototyping, dishonest to ship.
Guardrail only makes claims it can back up. Every check is a deterministic
rule — a blocklist entry, a regex match, a numeric cap, a rate limit —
evaluated against a policy file you write and can audit yourself, backed
by a real, persistent audit log (SQLite) you can query. Nothing here
pretends to know something it doesn't.
It's also not blockchain-specific. Shell execution, email, HTTP
requests, file deletion, database writes, crypto transactions — same
engine, same policy file, same rules.
Four ways to use it
1. CLI — for testing a policy by hand
Shown above. No setup, instant feedback while you write rules.
2. MCP server (`mcp_server.py`) — the easy on-ramp, advisory
Exposes `guardrail_check`, `guardrail_record_outcome`, and
`guardrail_agent_history` as MCP tools any MCP-compatible agent (Claude
Desktop, Claude Code, custom MCP clients) can call.
{
"mcpServers": {
"guardrail": {
"command": "python3",
"args": ["/absolute/path/to/agent-guardrail/mcp_server.py"],
"env": { "GUARDRAIL_POLICY": "/absolute/path/to/agent-guardrail/policies/default.yaml" }
}
}
}Then tell your agent (in its system prompt) to always call
`guardrail_check` before spending money, deleting data, messaging someone
externally, or running code.
Be clear-eyed about its limit: like any MCP tool, nothing stops the
calling model from just not invoking it. This only helps if the agent is
instructed to always check first — for a guarantee it can't skip, see #3.
3. `guardrail.decorator.enforce` — the real guarantee
Wraps the actual Python function that performs a tool's side effect. The
check runs in your code, before that function executes — the model never
gets a chance to call the real function directly.
from guardrail.decorator import enforce, BlockedActionError
@enforce(engine, tool_name="send_email")
def send_email(agent_id: str, to: str, subject: str, body: str):
... # only runs if the decision is ALLOW, or WARN-and-confirmedUse this if you're building your own agent loop (LangChain, CrewAI, a
custom MCP host, a Slack bot with tool access). Run `python3
examples/example_agent_usage.py` to see it block a real function call.
4. `guardrail.mcp_enforced_server.EnforcedGuardrailMCPServer` — the real guarantee, over MCP
The MCP server in #2 above is honest about being advisory: the model
gets a `guardrail_check` tool, but nothing stops it from calling the
*actual* tool (exposed by some other MCP server, or by the model's own
direct access) without checking first, or checking one thing and doing
another. If the model talks to your infrastructure only over MCP - no
Python decorator possible - this is the same #3 guarantee for that case:
the operator registers real action executors (the code that holds real
credentials and performs the real side effect) as the *only* way the
model can invoke that action at all.
from guardrail.mcp_enforced_server import EnforcedGuardrailMCPServer
def do_transfer(request):
wallet = get_wallet_for(request.agent_id) # real credentials, held here - never exposed to the model
tx_hash = wallet.transfer(to=request.arguments["to"], amount=request.arguments["amount"])
return {"tx_hash": tx_hash}
server = EnforcedGuardrailMCPServer(policy_path="policies/default.yaml")
server.register_action(
"wallet.transfer", "Transfer funds from the agent's wallet.",
input_schema={"type": "object", "properties": {"to": {"type": "string"}, "amount": {"type": "number"}}, "required": ["to", "amount"]},
executor=do_transfer,
)
server.serve_stdio()The model is given exactly one MCP tool named `wallet.transfer` - there
is no separate, unguarded way to move funds through this server. A BLOCK
decision means `do_transfer` never runs. Both this and `enforce()` share
one implementation of "check, maybe route WARN to a human, run only if
not blocked, report the real outcome back" (`guardrail/enforcement.py`) -
not two independently-maintained copies of the same guarantee.
Getting a human to actually confirm a WARN
`on_warn` is the hook — Guardrail ships two ready-made implementations:
Local web UI (`guardrail/confirmation/web_ui.py`) — a tiny built-in
server (stdlib only, no Flask) with Approve/Reject buttons. The wrapped
function blocks until someone clicks one, or times out (fails closed —
timeout means reject, not "allow by default").
from guardrail.confirmation.web_ui import ConfirmationServer
confirmation = ConfirmationServer(port=8787, timeout_seconds=300)
confirmation.start(open_browser=True)
@enforce(engine, tool_name="wallet.transfer", on_warn=confirmation.request_confirmation)
def transfer(...): ...Try it live: `python3 examples/example_web_confirmation.py`, then open
http://localhost:8787.
Terminal prompt (`guardrail/confirmation/cli_ui.py`) — for scripts and
local testing where a browser is overkill:
from guardrail.confirmation.cli_ui import cli_confirm
@enforce(engine, tool_name="wallet.transfer", on_warn=cli_confirm)
def transfer(...): ...Neither is required — `on_warn` is just a function `(decision) -> bool`,
so a Slack message, a ticket, or anything else you already use works too.
Writing a policy
Policies are plain YAML — see `policies/default.yaml` for a real, working
starting point (11 confirmation-gated tools, 10 destructive-pattern
checks, numeric caps, domain rules, rate limits, all commented).
| Rule type | What it checks |
|---|---|
| `blocked_tools` | Tool names that are never allowed |
| `confirmation_required_tools` | Tool names that always produce `WARN` |
| `argument_patterns` | Regex against the JSON-serialized call arguments — destructive shell commands, SQL, leaked credentials, path traversal, SSRF, force-pushes, regardless of which tool carries them |
| `numeric_caps` | Per-tool numeric field caps, tighter for agents with no history |
| `aggregate_caps` | A cap shared across *several* tools, tracked as one running total per agent — see below |
| `domain_rules` | Allow/deny lists on a URL or email-recipient field, per tool |
| `rate_limits` | Sliding-window call limits per (agent, tool), backed by SQLite |
`numeric_caps` limits each tool independently — `wallet.transfer` capped
at 1000/day and `wallet.approve` capped at 1000/day separately means an
agent using both can still move 2000/day combined. `aggregate_caps`
closes that: every tool listed in the same group draws from one shared
running total, e.g.
aggregate_caps:
daily_money_movement:
tools:
wallet.transfer: amount
wallet.approve: amount
window_seconds: 86400
max_unknown_agent: 5
max_known_agent: 1000Only *confirmed* spend counts toward the total: a `BLOCK`ed request never
adds anything, and a request that's provisionally recorded (because its
own check passed) is refunded if the real action later turns out not to
have succeeded — `engine.record_outcome(request_id, "error")`, called
automatically by both `enforce()` and the enforced MCP server (they
share one implementation of this, `guardrail/enforcement.py`) when the
real executor raises, or when a `WARN` a human rejects results in a
`BlockedActionError`. Real enforcement of this therefore has the same
caveat as everything else that depends on `record_outcome` being called:
it works fully under `enforce()` and the enforced MCP server (see
below); under the *advisory-only* MCP server (#2 above), a
provisionally-recorded amount just stays recorded, since nothing ever
reports back whether the action actually happened. See
`guardrail/storage/aggregate_spend.py`'s module docstring for the full
picture.
No code changes needed to adjust any of this — edit the YAML, restart the
process (or the MCP server).
Running the tests
pip install -r requirements.txt
PYTHONPATH=. python3 -m unittest discover -s tests -v134 tests: rule evaluation, the full engine pipeline (real SQLite-backed
rate limiting, aggregate spend tracking, and audit persistence), the
`enforce` decorator and the enforced MCP server (both proving a `BLOCK`
genuinely prevents the real action from running, sharing one
implementation of that guarantee), the advisory MCP server's JSON-RPC
handling, the confirmation web UI over real HTTP requests against a
live server, and a dedicated suite that checks the *shipped*
`policies/default.yaml` — not just synthetic test policies — actually
catches what it claims to.
What's honestly still missing
- Single-process SQLite by default. Fine for one agent process; for
multiple replicas sharing rate limits/audit history, point every
process at the same file on shared storage, or swap in a real database
(the storage classes are small and easy to re-target).
- Secrets/PII redaction in the audit log is on by default.
`AuditLog` redacts values whose key looks sensitive (`password`,
`api_key`, `authorization`, ...) and a couple of high-confidence value
shapes (PEM private key blocks, JWT-shaped strings) regardless of key
name, recursing into nested dicts/lists - see
`guardrail/storage/redaction.py` for exactly what is and isn't caught,
and why general-purpose entropy heuristics were deliberately left out
(too many false positives on ordinary UUIDs/hashes). Pass
`AuditLog(redact=False)` to store arguments as-submitted, or
`extra_sensitive_keys={...}` to redact additional field names specific
to your tools.
- **The default policy is a reasonable starting point, not a complete
threat model.** It catches well-known destructive shell/SQL patterns
and obvious credential formats — extend `argument_patterns` for
whatever your agents actually touch.
- The confirmation web UI has no auth. It binds to `127.0.0.1` by
design (not exposed on the network), but anyone with local access to
that port can approve/reject. Fine for a single developer's machine;
put it behind your own auth if multiple people share the host.
None of these are mocked or faked — they're just not built yet, and
they're the honest next steps if you adopt this.
Publishing this / getting people to actually use it
See `PUBLISHING.md` for a concrete checklist: MCP directories to submit
to, what a listing needs, and what "done" looks like.
Related projects
Same author, same principle applied elsewhere:
a security decision layer for AI agents transacting on-chain. MIT,
112 tests.
cryptographically signed (Ed25519), independently verifiable
attestations for agent-to-agent payment policy decisions. Early
proof of concept.
vendor-neutral open spec (JWT+EdDSA) for signing agent policy
decisions, verifiable by anyone. x402-attest above uses a custom
format; this is the generalized version. Draft v0.1.
Project layout
guardrail/
__main__.py CLI implementation — also the `guardrail` console command
mcp_server.py MCP stdio server — also the `guardrail-mcp-server` console command
core/
models.py ActionRequest, RuleMatch, GuardrailDecision (stdlib only)
policy.py Policy loader (the one place PyYAML is used)
rules.py Deterministic rule evaluators
storage/
rate_limiter.py SQLite-backed sliding-window rate limiter
audit.py SQLite-backed persistent audit log
engine.py GuardrailEngine — orchestrates rules + rate limit + audit
decorator.py enforce() — the unbypassable integration point
confirmation/
web_ui.py Local web UI for human approve/reject (stdlib http.server)
cli_ui.py Terminal-prompt confirmation
policies/default.yaml Copy of the default policy bundled into the installed package
policies/default.yaml Canonical, editable default policy (git-clone workflow)
cli.py Thin shim -> guardrail/__main__.py (for `python3 cli.py`)
mcp_server.py Thin shim -> guardrail/mcp_server.py (for `python3 mcp_server.py`)
pyproject.toml Package metadata — `pip install .` gives you `guardrail` + `guardrail-mcp-server`
.github/workflows/ci.yml Runs the test suite + policy validation + package build on every push
examples/
example_agent_usage.py Decorator basics
example_web_confirmation.py Real browser-based approve/reject, live
tests/ 46 unit tests, all runnable with just PyYAML installed
CONTRIBUTING.md How to add a rule type, ground rules
CHANGELOG.md Version history
PUBLISHING.md How to actually get this in front of people
landing/index.html Static one-page site (open directly or host on GitHub Pages)Frequently asked questions
What is agent-guardrail?
agent-guardrail is A deterministic policy firewall for AI agent tool calls. YAML rules, no LLM calls, no risk scores. CLI + MCP server (advisory) + @enforce() decorator (unbypassable). MIT, 46 tests, zero network dependency.
How do I install agent-guardrail?
Open the GitHub repository and follow its README. Most MCP servers are added to your client's MCP config, then called by your agent.
Is agent-guardrail open source?
Yes — it is hosted on GitHub at https://github.com/rudimentall1/agent-guardrail and has 1 stars.
Related MCP tools
Open-source coding agent memory. Records issues, attempts, fixes and decisions, then warns your agent before it repeats an approach that already failed. Native MCP server for Claude Code, Cursor, Antigravity and Codex. 100% local, no cloud, no telemetry. MIT.
MCP server and Claude plugin for Postgres skills and documentation. Helps AI coding tools generate better PostgreSQL code.
AI-powered OSINT agent with interactive REPL, MCP server, and CLI. 19 tools. Works with Claude, GPT-4, or local models. For authorized security research only.
Decision audit trail + persistent memory for AI trading agents. Outcome-weighted recall, tamper-evident SHA-256 chain with RFC 3161 anchoring, 20 MCP tools.
Pre-build reality check for AI coding agents. Scans GitHub, HN, npm, PyPI, Product Hunt. MCP server. 290+ stars.
Open-source AI agent firewall for MCP security and agent egress. Scans mediated HTTP, MCP, A2A, and WebSocket traffic for exfiltration, SSRF, and prompt injection, and emits mediator-signed action receipts: verifiable audit evidence from outside the agent.
Run your own MCP server? See who uses it and what to fix.
Measure it with TrackMCP