What MCP Security Lab does
A pre-deployment lab for MCP servers. It maps what a server exposes, tests your policy against a worst-case agent in isolation, and records evidence you can replay. Everything on this page is built unless it is marked planned.
Read-only discovery over Streamable HTTP
The lab connects to your https endpoint as an MCP client and asks only for metadata. It never calls your tools.
- What it asks for
initialize, then the list methods for the capabilities your server advertises (tools, resources, resource templates and prompts), and one request for a method that does not exist, to see how errors are handled.- What it refuses
- Every other method, including
tools/call, is refused inside the lab before anything is sent. - Network guards
- Only https endpoints on port 443 or 8443 are accepted. Each DNS answer is checked when a connection opens, and private, loopback, link-local and cloud-metadata addresses are refused. Redirects that leave the endpoint’s origin are not followed. Requests time out, and responses over 1 MB are rejected.
- Credentials
- An optional bearer token is used for that discovery only and is never stored.
- Limits
- Up to 10 pages and 500 items per list. A truncated inventory is reported as a finding, not hidden.
The exact requests are listed in the connection docs.
Tool inventory and schema analysis
Each discovered tool is classified by what it can do, then checked for weak schemas, misleading annotations and hidden instructions. Classification is heuristic: names, parameters and annotations are signals, not proof.
- Capabilities
read, write, destructive, financial, egress, exec, filesystem, pii, credentials, admin, browser. A tool can have several; the policy decides per capability.- Schema checks
- Free-form command, SQL, path, amount and destination parameters; open schemas on state-changing tools; write tools annotated as read-only; idempotency claimed without a key.
- Injection checks
- Instruction-like text, directions to use other tools and invisible Unicode in tool descriptions, server instructions, prompts and resource descriptions. Secret-like strings anywhere in the metadata.
- Drift
- Per-tool hashes are compared with the previous snapshot of the same target, so a definition that changes after approval is flagged.
See all 56 rules with their method, evidence and limits.
A deterministic policy engine
You state what the agent may do in a versioned policy. Every proposed call is evaluated against it, with the same result every time. No language model takes part in a decision.
- Intent binding
- Denies calls that need capabilities the user’s request does not cover, and tools outside the task.
- Approvals
- Calls with the capabilities you choose wait for an out-of-band human to approve them. A
confirm: trueset by the agent does not count as approval unless your policy says so. - Egress allowlist
- Limits the email domains and hosts that outgoing calls may reach.
- Data-flow labels
- In simulations, tool outputs carry labels such as pii, secret and financial_record. Labeled data can be kept away from tools that send data out.
- Idempotency
- Requires an idempotency key for chosen capabilities, such as payments, and denies a key reused in the same run.
- Limits
- Caps the amount of a financial call, the number of actions per run and the number of calls per tool.
Deny overrides approval, and approval overrides allow. Each policy version is hashed, and the evidence records the hash. Policy reference.
Isolated simulations against synthetic fixtures
Scenarios run against in-memory MCP servers whose backends are mocks. Names, addresses and keys in them are fictional, and a destructive scenario cannot touch a real file, account or inbox.
| Fixture | What it models | Scenarios |
|---|---|---|
| Acme Billingbilling | Flagship: invoices, customers, email, files and refunds with typical integration gaps. | 3 |
| Acme Billing (hardened)billing | The same tool surface after server-side fixes. | 3 |
| Acme Billing 1.1 (rug pull)billing | A later release whose read_invoice description was silently poisoned. | 3 |
| Workspace Filesfilesystem | Unbounded read, write and delete over a mock workspace. | 1 |
| Git Repositorygit | Status, diff, commit and (force) push on a mock repository. | 1 |
| Support Mailboxemail | Inbox, drafts and immediate send. | 1 |
| Payments Deskpayment | Tickets, payments and refunds without limits or idempotency. | 3 |
| App Databasedatabase | A "read-only" query tool that executes writes, plus role management. | 2 |
| Browser Automationbrowser | Navigation and page reading with an injected exfiltration link. | 1 |
| Poisoned Utilitiesutilities | Tool poisoning, cross-tool shadowing and hidden Unicode in descriptions. | 1 |
Mock replicas of your tool surface
For a server you connect, the lab builds an in-memory copy of its tools: the same names, schemas and descriptions, with synthetic handlers. It then generates attacks for whatever that surface offers.
- Generated scenarios
- For each kind of risky tool the surface has, an injected instruction tries to use it: sending data out, deleting, moving money, changing permissions, running commands, or writing during a read-only task. A control checks that the legitimate read still works. If no read-only tool can carry the injection, no scenarios are generated and the gap is reported as not tested.
- What a result means
- It tests your policy against your tool surface. It does not show that your real server enforces the same controls; that check is reported as Not tested.
Evidence hashes and deterministic replay
Every scenario run produces an evidence record: the user’s intent, each proposed call and where it came from, the matched policy rules, the approval state, a digest of each result, the mock side effects and whether each assertion held.
- Hash
- The record is serialized as canonical JSON and hashed with SHA-256.
- Replay
- Re-running a recorded scenario with the same fixture, scenario, policy and runner versions must produce the same hash. If any version differs, the replay reports the mismatch instead of a result.
- Redaction
- Secret-like strings in arguments, results and side effects are redacted before the record is hashed.
Reports and before/after comparison
A report is one JSON object, and the PDF is a rendering of the same object. Nothing in a report is computed by a language model.
- JSON
- Schema
msl.report/v1, with a SHA-256 of the report body. - Executive summary, scope, target and protocol metadata, the policy under test, tool inventory, findings, results by evidence kind, simulation evidence, residual risk, reproduction steps, limitations and scoring.
- Comparison
- Run the same target after a fix and compare readiness, risk score, coverage and attacks contained, plus every rule whose result changed: fixed, improved, regressed, new or removed.
Release gate and Claude assistance
Two ways to act on results: block a release automatically, or ask Claude to explain a finding in plain language.
- CI release gate
- A command-line runner (
npm run msl -- scan https://… --fail-on high) that exits non-zero when a target is not ready under your policy, with an example GitHub Actions workflow. - Ask Claude (free)
- Every failing check in a run can produce a redacted prompt for your own Claude account: what the finding means, the impact, and a suggested fix. We send nothing; you paste it.
- Explain with Claude (in-app)
- Deployments with a Claude API key show explanations inside the run page, limited per workspace and per day. It is switched off on this deployment.
- Authority
- Claude explains; it never decides. Outcomes, severities, scores and hashes come only from the deterministic engine.
Planned, not built yet
These are on the roadmap. None of them exists today, and no dates are committed.
- stdio connector Planned
- Discovery for servers that run as local processes. Today only remote Streamable HTTP endpoints are supported.
- MCP 2026-07-28 support Planned
- Discovery for servers that only speak the stateless 2026-07-28 revision. Today discovery covers revisions up to 2025-11-25.
- LLM agent observation mode Planned
- Run a real model as the agent (Claude first) and report whether it follows the injection, separately from the deterministic baseline.
- Team features Planned
- Shared workspaces for teams.
- Scheduled rescans Planned
- Periodic re-discovery to catch tool definitions that change between releases.
- Authorized dynamic auth tests Planned
- Token audience and tenant-boundary tests. Today these checks are reported as not tested.
Where it fits
The lab works before release. It complements tools that work at other stages; it does not replace them.
- Static scanners
- Check metadata, source code or packages for known patterns, often in CI. The lab’s static checks overlap with theirs. Its policy simulations and replayable evidence are a different layer, so use both.
- Runtime gateways and proxies
- Enforce policy on live traffic in production. The lab blocks nothing in production. It tests a policy before release, which can inform what you enforce at runtime.
- Live penetration tests
- Exercise your real server and can find implementation bugs. The lab cannot: discovery is read-only, and simulations run on mocks.
The lab is not a runtime firewall, a source-code or dependency scanner, or a penetration test.