What mcp-guard catches, what it deliberately does not claim, and the evidence behind both. mcp-guard
is a build-time Roslyn analyzer: it scans the model-visible strings of a C# MCP server: [Description]
text on tools / parameters / resources / prompts, tool Names, and parameter / enum-member names, and
fails the build on a poisoned one. It owns the static half of MCP defense; runtime guards are out of
scope (below).
| Attack class | Rules | Standards |
|---|---|---|
| Prompt injection / instruction override (incl. hidden Unicode, ANSI escapes, embedded markup) | MCPG001, MCPG002, MCPG005, MCPG008 | OWASP MCP03 Tool Poisoning; MCP-38 MCP-10 |
| Sensitive-data exposure / secret-file references (in descriptions and parameter/enum names) | MCPG003 | MCP-38 MCP-11 (full-schema poisoning) |
| Data exfiltration (transmit verb + external sink, markdown image/link sinks, encoded-blob payloads) | MCPG004, MCPG011 | OWASP MCP03; MCP-38 MCP-10 |
| Manipulative / authority phrasing; preference manipulation | MCPG006 | MCP-38 MCP-15 (MPMA) |
| Cross-tool / tool-shadowing references | MCPG009 | MCP-38 MCP-13 |
| Off-screen whitespace hiding | MCPG010 | MCP-38 MCP-10 |
| Capability ⇄ name mismatch (advisory) | MCPG007 | — |
| Confirmed exfiltration payload — secret + sink on one description → Error | MCPG012 | escalation |
| Description integrity / source-level rug-pull — drift from a committed baseline (opt-in) | MCPG013 | MCP-38 MCP-16 |
This is the natural ceiling of a build-time text scanner: MCP-38 Category I (semantic manipulation / poisoning) plus the rug-pull baseline. Most rules are individually high-confidence heuristics; MCPG012 escalates the co-occurrence of two of them to a build-breaking error.
Grounded in a known-attack corpus of
payloads transplanted from public defensive-research PoCs (DVMCP, Invariant Labs, Repello, CyberArk;
all exfil endpoints neutralized to example.test):
- 9 known attacks: each asserts the exact set of rules it triggers (e.g. Invariant
direct-poisoning→ MCPG001 + MCPG003 + MCPG004 + MCPG012; Repello's base64-wrapped exfil → MCPG011 + decoded MCPG003/MCPG004 + MCPG012; CyberArk full-schema poisoning in a parameter name → MCPG003). - 8 benign look-alikes — realistic clean tools that brush against the rules (OAuth login to an https
URL, artifact upload, a
curlfetch, base64/hex mentions, an authorized config read, a documentaryid_rsa_compat_modeparameter). All assert zero diagnostics. This is the false-positive guard; precision is the top priority. - 6 boundary cases — runtime-only attacks asserted to fire nothing (see below).
- Live tiers (integration tests): a poisoned
description is proven to survive serialization into the
tools/lista client receives, and a runtime rug-pull (a description swapped after load) is demonstrated end-to-end.
All of it runs on every PR: 126 analyzer + 3 integration tests, against .NET 8 and .NET 10.
A build-time analyzer cannot see runtime behavior or cross-session state. These are out of scope and belong to runtime / proxy tooling. The corpus includes them as negative-scope tests so the boundary is explicit (see the threat model):
- Runtime rug-pulls by a third-party server (the source-level analog is MCPG013).
- Indirect injection via data the tool fetches at runtime (DVMCP challenge 6, Backslash web-scraper).
- ATPA — payloads emitted in a tool's runtime output / errors (CyberArk).
- Tool shadowing across connected servers, cross-server confused-deputy, live typosquatting.
- Full-schema poisoning via non-standard JSON-schema fields: mcp-guard reads C# attributes, not the emitted schema (payloads in parameter/enum names are covered).
- Actual network exfiltration — observable only at runtime.
Pair mcp-guard with runtime defenses for the rest. It catches the poison before it ships.