trackmcp
Back to directory

Policy-as-code for MCP agents: deny risky tool calls before they run, prove what ran with verifiable evidence, and enforce egress in the kernel (eBPF/LSM, Linux). Deterministic, offline-first, bounded claims.

9 stars RustOthers Updated Sep 4, 2026
rustai-agentsmcppolicy-as-codeai-securitymcp-serverpolicy-enforcementagent-securitycicyclonedxevidence-bundlesgithub-actionsopenfeaturepromptfooprovenancesbomsupply-chain-securityebpfllm-securitymcp-security

Documentation

Assay

The open, recomputable evidence profile for privileged MCP tool actions.

Assay records what a privileged tool call decided, what was observed, and what stays unproven, so a reviewer can replay the claim offline instead of trusting the agent's account of itself. Enforcement is deterministic and fail-closed, and the enforcing proxy is the reference producer rather than the contract itself. Kernel-level (eBPF/LSM) observation on Linux is an optional stronger vantage. CI-native, no backend, bounded by design.

·

·

·

·

·


Agents got real tool access through MCP — and tool poisoning, rug pulls, and confused-deputy OAuth came with it. Most tools scan a server or filter a prompt. Assay sits at the tool-call boundary and does three things, in order.

One golden path: the release-pinned agent journey records the nine driven CLI/MCP steps and their exit/stdout contracts. Its protected-action fixture lives in examples/privileged-action-gate/.

Enforce, prove, stay honest

  • Enforce. A deterministic, fail-closed gate decides every `tools/call` before it runs, with the precise reason for each allow or deny. On Linux it adds real kernel enforcement — an eBPF/LSM IPv4/TCP connect-egress block and a Landlock TCP-connect port allowlist, both opt-in and fail-closed. A policy it cannot express exactly is refused, never half-applied.
  • Prove. Each decision and observed effect becomes an offline-verifiable, tamper-evident evidence bundle: the verdict, the pre-call establish journey, and declared-vs-observed conformance — all reviewable in CI, with no hosted backend.
  • Stay honest. Every claim carries its basis (`verified`, `self_reported`, `inferred`, `absent`), and a gate refuses to let a claim exceed what was observed. A tool returning "success" is the provider's assertion, never proof. Assay ships no single safety score and never claims more than it can prove.

Quickstart

bash
# Fast path: release installer for Linux and macOS.
curl -fsSL https://getassay.dev/install.sh | sh

# Confirm the command resolves; if setup fails, run `assay doctor`.
assay --version

# Source-build alternative (requires Rust):
cargo install assay-cli --version 6.0.0 --locked

python3 examples/mcp-quickstart/run.py

For v6.0.0, run the last command from a source checkout or an extracted published CLI archive.

The installer is binary-only and does not carry the bounded quickstart assets. The live

`getassay.dev` installer verifies the selected archive against its published SHA-256 sidecar before

extraction. Set `ASSAY_REQUIRE_PROVENANCE=1` to additionally require GitHub artifact provenance;

the default reports `provenance_not_requested` and strict success reports `provenance_verified`.

A checksum proves byte equality with the published sidecar, not producer identity. Provenance

identifies the source and build, not runtime safety or semantic correctness.

Captured runner output (the bundled local mock performs no external action):

text
assay quickstart: PASS
mcp_requests=initialize,tools/list,tools/call
decision=allow tool=read_file
decision_artifact=.assay/quickstart/decisions.ndjson
non_claim=forwarded_to_local_mock_only
Assay decides each MCP tool call before it runs, fail-closed, with the reason

Released surfaces:

  • Static project manifests are shipped for Claude Code and Cursor; Codex uses the equivalent TOML entry documented in the editor MCP recipe. Manifest presence is not host-discovery proof. `assay mcp config-path` supports Claude and Cursor only.
  • Published v6.0.0 CLI archives cover Linux x86_64/arm64, macOS x86_64/arm64, and Windows x86_64. The Python wheels cover CPython 3.12 on macOS x86_64/arm64 and Linux x86_64; other interpreters and platforms are not claimed.
  • Published `assay-mcp-server` archives cover Linux x86_64/arm64. MCPB and `server.json` package descriptors are also published; their presence is not host-discovery proof.
  • CI: GitHub Action. Core flows need no hosted backend or API key. New to the threat model? The OWASP MCP Top 10 mapping states, per risk, what Assay covers and deliberately does not.

What ships

OutputWhat it is
Policy gate`assay mcp wrap` — deterministic allow/deny before tools run, with the reason.
Evidence bundleOffline-verifiable, tamper-evident archive for audit and replay.
Trust Basis / Trust CardCanonical `trust-basis.json` (bounded claim classification) plus review-friendly `trustcard.{json,md,html}`.
External receiptsEval outcomes, runtime decisions, and model inventory as bounded receipts with JSON Schema contracts.
Tool-decision surfaceEach privileged `tools/call` recorded as `assay.tool_decision_surface.v0` — sensitive ids hashed, raw arguments never stored.
SARIF / CIGitHub Action, Security-tab integration, policy gates on PRs.
AttestationExport a bundle as an in-toto / DSSE statement (v0), anchor-pluggable.
text
Agent ──► Assay ──► MCP Server
              ├─ ✅ ALLOW / ❌ DENY  (policy, with reason)
              ├─► 📋 Evidence bundle (offline-verifiable)
              └─► 📊 Trust Basis → Trust Card → SARIF / CI

Current release: `v6.0.0`. CHANGELOG.md and release notes remain the authority for released behavior; merged changes after the tag are `Unreleased`, and crates.io publication is separate from merge state.

Is this for me?

Yes if you already have eval output, runtime decisions, inventory artifacts, or MCP tool-call tests, and you want a small reviewable CI artifact instead of a dashboard — bounded auditability, not a scalar trust badge.

Not yet if you need Assay to judge model correctness for you, want a hosted dashboard as the product, or want a compliance claim rather than a bounded evidence boundary. Assay is not a trust-score engine, a generic eval dashboard, or a hosted observability product — see what it is and is not.

See it work

An agent tries a privileged action — `github.add_deploy_key` — through the enforcing proxy, decided per call before it forwards, offline against a local mock (no real credentials):

bash
cd examples/privileged-action-gate && ./run.sh
privileged-action PR-gate demo

A deny is fail-closed caution, not a verdict on intent; an allow is the decision to forward, never proof the action happened. Declared-vs-observed conformance is recorded beside the verdict, never as a gate. Full walkthrough: privileged-action-gate.

Pick your path

You haveWhat you getStart here
Promptfoo JSONL from CI evalsEval outcome receipts + verified bundle + Trust Basis diffPromptfoo JSONL
OpenFeature `EvaluationDetails`Decision receipt + verified bundleOpenFeature
CycloneDX ML-BOM model componentInventory receipt + verified bundleCycloneDX ML-BOM
MCP tool callsAllow/deny audit trail + observed-behavior evidenceMCP Quick Start
A GitHub PR gateTrust Basis diff, gate status, SARIF/JUnit-ready outputCI Guide
A Runner archive / coverage annotationCoverage descriptors + claim-class cells + a claimed-vs-observed checkCoverage-honesty walkthrough

The workflow stays small: import or record a bounded outcome, bundle and verify it, compile `trust-basis.json`, gate the Trust Basis diff. Assay doesn't make the upstream tool the source of truth; it makes the evidence boundary inspectable. For privileged tool actions, the MCP proxy records each `tools/call` as a structured tool-decision surface — keeping the asserted-versus-verified line honest.

Policy is simple

yaml
version: "2.0"
name: "my-policy"
tools:
  allow: ["read_file", "list_dir"]
  deny: ["exec", "shell", "write_file"]
schemas:
  read_file:
    type: object
    properties:
      path: { type: string, pattern: "^/app/.*" }
    required: ["path"]

`assay init --from-trace trace.jsonl` generates the runtime-observation policy used by the trace-generation flow (`files`, `network`, and `processes`); it is not an MCP authorization policy. Migrate a legacy MCP `constraints:` policy with `assay policy migrate`. See Policy Files.

Why Assay

Canonical evidenceAssay's evidence model is the stable contract; OpenTelemetry and protocol adapters (ACP / A2A / UCP) map into it.
DeterministicSame input, same decision — not probabilistic.
Bounded claimsExplicit about verified vs visible vs absent — no score-first UX.
Offline-firstNo backend required for core enforcement and bundle verification.
Checkable provenanceWhich piece of the source-class and coverage model shipped when, as commits you can `git log` rather than claims you have to take — provenance, prior art credited first.

Learn more

Evidence epistemology, latency, and the internal Runner

Trust claims use explicit epistemology, not a single safety score: `verified` (direct evidence or offline verification), `self_reported` (emitted without independent corroboration), `inferred` (bounded, documented rules), `absent` (no trustworthy evidence). Assay ships no aggregate trust score or `safe/unsafe` badge as the main output — see ADR-033.

Tool-decision path latency on an M1 Pro fragmented-IPI harness: main protection `0.771ms` p50 / `1.913ms` p95; fast-path `0.345ms` p50 / `1.145ms` p95. These are tool-decision timings, not end-to-end model latency.

Assay-Runner is an internal measured-run subsystem behind the delegated Linux/eBPF acceptance path — `publish = false`, not a standalone product, no release commitment.

Ecosystem

Repositories that compose with Assay's evidence layer:

  • assay-action — GitHub Action: verify bundles, PR summaries, SARIF (Marketplace).
  • Assay-Harness — recipe, gate, and report layer over canonical evidence artifacts.
  • observed-effect-v0 — worked examples of the bounded observed-effect evidence record and its neutral carriers (in-toto, SCITT, MCP evidenceRef).
  • gateway-evidence-replay — deterministic offline replay verifier for gateway-path evidence bundles.
  • RGE-Bench — a conformance kit for evidence reviewability, maintained separately under its own machine-checked neutrality guard. Reproduction there is digest-scoped and does not carry forward: the v1 71-vector digest `sha256:e769822bc6c9e31085da7b1a17b163b9747fe0d04314fbb8685d4e612087c7cb` and the current v2 digest `sha256:ba0e3795d75c788fa48313ab462493f22d78759851d1b3275d8117051bb22fd0` (95 vectors) each carry one reported independent implementation by a second author on a different stack. JM-Lab reported the v2 95/95 reproduction on 2026-08-24, from the contract text and author-supplied inputs without reading `expected`. See its REPRODUCTIONS.md.

Open profile: privileged-mcp-action/v0

`privileged-mcp-action/v0`

is a composition and verification contract over evidence records that already exist: what a

privileged MCP tool call decided, what was observed of its effect, and what stays unproven. It adds

no new envelope and no aggregate verdict.

It ships with a 14-vector conformance corpus

(5 accept, 9 reject) whose digest is a candidate: it is not called reproduced until a non-author

implementation derives the expected outcomes from the specification text alone.

**That reproduction is open, and the invitation is real:

#1840.** Any language, any stack. The invitation

names the exact commit the current digest describes. The

clean-room protocol provides an

opaque, attested inputs pack, a one-command scoring action, and an implementation-report template

without supplying verifier logic or expected outcomes. The

corpus README states the authorship boundary and

the claim ceiling.

Contributing

bash
cargo test --workspace
cargo clippy --workspace --all-targets -- -D warnings

See CONTRIBUTING.md and GitHub Discussions.

License

MIT

Frequently asked questions

What is assay?

assay is Policy-as-code for MCP agents: deny risky tool calls before they run, prove what ran with verifiable evidence, and enforce egress in the kernel (eBPF/LSM, Linux). Deterministic, offline-first, bounded claims.

How do I install assay?

Open the GitHub repository and follow its README. Most MCP servers are added to your client's MCP config, then called by your agent.

Is assay open source?

Yes — it is hosted on GitHub at https://github.com/Rul1an/assay and has 9 stars.

Related MCP tools

Run your own MCP server? See who uses it and what to fix.

Measure it with TrackMCP