trackmcp
Back to directory
citarium

agentreliability-mcp

View on GitHub

MCP server over a cited knowledge graph on testing, benchmarking and auditing AI agents. Remote streamable-HTTP, 8 tools, no auth.

0 starsOthers Updated Aug 30, 2026
ai-agentsknowledge-graphllmmcpmcp-servermodel-context-protocol

Documentation

Agent Reliability — MCP server

> Testing, benchmarking and auditing autonomous AI agents — methods, harnesses, evidence

A remote MCP server over a curated knowledge graph. Every claim it

returns is bound to a registered source: the tools hand back claims *with*

their citations and a confidence value, so an agent can show its work

instead of asserting.

Nothing to install. It is a hosted streamable-HTTP endpoint:

code
https://agentreliability.dev/mcp

Also listed on Smithery.

Add it to a client

Claude Code

bash
claude mcp add --transport http agent-reliability https://agentreliability.dev/mcp

Claude Desktop / any client reading `mcpServers`

json
{
  "mcpServers": {
    "agent-reliability": {
      "type": "streamable-http",
      "url": "https://agentreliability.dev/mcp"
    }
  }
}

No API key, no account, no auth. Read-only.

Check it answers, without any client at all:

bash
curl -s https://agentreliability.dev/mcp \
  -H 'Content-Type: application/json' \
  -H 'Accept: application/json, text/event-stream' \
  -H 'mcp-protocol-version: 2025-06-18' \
  -d '{"jsonrpc":"2.0","id":1,"method":"tools/list"}'

Tools

Eight, each with an `outputSchema`, each returning `structuredContent`.

toolargumentswhat it does
`get_overview`Corpus overview: what this instance knows, counts by type, published tags, freshness. Start here when you land and do not yet know whether this corpus can answer your question.
`search``query`, `limit?`Full-text search over the knowledge graph. Accent- and apostrophe-insensitive, so query in the user's own words; every hit carries its relevance score and the fields it matched.
`answer``question`Answer a question from the corpus. Returns the matched object's claims with sources and confidence — never an unsourced answer.
`get_entity``id`Fetch one knowledge object by id, with its claims and the sources each claim cites.
`get_topic``tag`List the knowledge objects carrying a tag (topics are content-backed tags).
`get_related``id`Graph neighbours of an object: outgoing and incoming relations, each with its relation type.
`get_sources``object_id?`The whole source registry, or just the sources cited by one object. Use it to judge the corpus before trusting it.
`get_latest``limit?`Most recently verified knowledge objects — a freshness signal.

The intended path is `get_overview` → `search` or `answer` → `get_entity`

→ `get_related`. `get_overview` exists because an agent that has just

arrived needs to know whether this corpus can help *before* it spends a

call guessing.

What is in the corpus

knowledge objects38
registered sources31
published topics93
typeobjects
entity25
guide9
comparison2
faq1
glossary1

Subject matter: evals and benchmarks (GAIA, AgentBench, Inspect), LLM-as-judge and its failure modes, Goodhart and benchmark contamination, fault injection and chaos testing, approval gates and autonomy levels, grounding and faithfulness.

Questions it is built to answer

  • *How do I tell a real eval from a benchmark my agent has memorised?*
  • *What does calibration mean for an LLM judge, and how is it measured?*
  • *Which failure modes does fault injection actually catch?*

What an answer actually looks like

A real call against the live endpoint — `answer` with

*"how do I tell a real eval from benchmark contamination"* — returns this `structuredContent`, trimmed:

json
{
  "answered": true,
  "entity": {
    "id": "agent-reliability-glossary",
    "name": "Agent reliability glossary",
    "evidence_tier": "secondary",
    "confidence": 0.85,
    "last_verified": "2026-08-08",
    "canonical_url": "https://agentreliability.dev/k/agent-reliability-glossary"
  },
  "claims": [
    {
      "text": "An eval is a structured, repeatable test that measures an LLM or LLM-based system against a defined dimension; frameworks package evals as registries of reusable templates.",
      "sources": [{ "title": "openai/evals — framework for evaluating LLMs and LLM systems" }]
    }
  ]
}

Note what travels with the answer: the evidence tier, a confidence,

the date it was last verified, and the source behind the claim — not

as prose an agent has to parse, but as fields it can act on. An agent can

decline to use a weak claim, or cite the primary source directly.

When the corpus cannot answer, `answered` is `false`. It does not

improvise, and the miss is recorded so the gap can be filled.

Machine-readable surfaces

The MCP endpoint is one of several. The same corpus is served as plain

files an agent can read directly:

surfacewhat it is
`/llms.txt`the index, as `text/plain`
`/llms-full.txt`the whole corpus in one file
`/ai-index.json`every surface this instance publishes, with its content type
`/api/index.json`one JSON document per knowledge object
`/api/sources.json`the source registry, in full
`/.well-known/mcp/server.json`this server's manifest

Each knowledge object has a human page and a machine twin at the same id,

with a canonical URL that agrees across all of them.

Behaviour worth knowing before you integrate

  • `POST` only. Every other method answers `405` with an `Allow: POST, OPTIONS` header.
  • Rate limit: 120 requests per minute per client, counted in a shared

store, published on every response as `RateLimit-Limit`,

`RateLimit-Remaining` and `RateLimit-Reset` (all three exposed via CORS).

It fails open: if the store is unreachable the request is served.

  • Malformed input gets a spec-correct JSON-RPC error — `-32700` for

unparseable bodies, `-32602` for an unknown tool — never an HTML error page.

  • Request bodies are capped and validated before transport.

Privacy

No accounts, no cookies, no ads. Usage is measured in aggregate with

daily-rotating hashed identifiers and a 200-day retention; raw IPs are

never stored. Full policy: PRIVACY.md.

Provenance and licence

Knowledge content is CC-BY-4.0: use it, cite it. The source registry is

public precisely so a claim can be checked rather than trusted —

`get_sources` returns what any given claim rests on.

Claims carry an evidence tier and a `last_verified` date. Where the

evidence is weaker, the object says so rather than rounding up.

How it is built

Compiled and served by Citarium, a source-available framework for turning a

knowledge graph into a website, an API, an MCP server and agent-readable

files from a single source, under external evaluation.

The framework's code is licensed under the Business Source License 1.1 and

its repository is not public. What is public — and what actually matters for

trusting an answer — is this server, the corpus it serves, and the registered

source behind every claim: `get_sources` returns what any given claim rests

on, so it can be checked rather than trusted.

This repository is the server's public face: its manifest and its

documentation. The corpus itself lives at agentreliability.dev.

Frequently asked questions

What is agentreliability-mcp?

agentreliability-mcp is MCP server over a cited knowledge graph on testing, benchmarking and auditing AI agents. Remote streamable-HTTP, 8 tools, no auth.

How do I install agentreliability-mcp?

Open the GitHub repository and follow its README. Most MCP servers are added to your client's MCP config, then called by your agent.

Is agentreliability-mcp open source?

Yes — it is hosted on GitHub at https://github.com/citarium/agentreliability-mcp.

Related MCP tools

atlassianatlassian-mcp-server

Official remote MCP server for Atlassian. Securely connect Jira, Confluence, Jira Service Management, Bitbucket, and Compass to Claude, ChatGPT, Cursor, VS Code, and other AI tools using OAuth 2.1 or API tokens.

1,015 JavaScript
aiai-agentsatlassian+17
riponcmprojectmem

Open-source coding agent memory. Records issues, attempts, fixes and decisions, then warns your agent before it repeats an approach that already failed. Native MCP server for Claude Code, Cursor, Antigravity and Codex. 100% local, no cloud, no telemetry. MIT.

796 Python
ai-agentsai-memoryai-tools+17
IvanMurzakUnity-MCP

AI Skills, MCP Tools, and CLI for Unity Engine. Full AI develop and test loop. Use cli for quick setup. Efficient token usage, advanced tools. Any C# method may be turned into a tool by a single line. Works with Claude Code, Gemini, Copilot, Cursor and any other absolutely for free.

4,137 C#
aiai-integrationgame-development+16
skyhook-ioradar

The missing open-source Kubernetes UI with a built-in MCP server for AI agents. See what's broken, why, and what changed. Issues, Topology, event timeline, Helm, GitOps, live service traffic, and cluster audits - all in one Go binary.

3,246 Go
argocdcloud-nativegitops+17
jgravellejcodemunch-mcp

Cut AI token costs 95%+ on code exploration. The leading MCP server for precise, symbol-level GitHub code retrieval via tree-sitter AST. Works with Claude Code, Cursor & any MCP client. 313B+ tokens saved.

2,651 Python
claudeclaude-codeai-coding+17
0xMassiwebclaw

Fast, local-first web content extraction for LLMs. Scrape, crawl, extract structured data — all from Rust. CLI, REST API, and MCP server.

2,319 Rust
ai-agentsclillm+17

Run your own MCP server? See who uses it and what to fix.

Measure it with TrackMCP