Skip to content
MCP Security Lab

What MCP Security Lab does

A pre-deployment lab for MCP servers. It maps what a server exposes, tests your policy against a worst-case agent in isolation, and records evidence you can replay. Everything on this page is built unless it is marked planned.

Read-only discovery over Streamable HTTP

The lab connects to your https endpoint as an MCP client and asks only for metadata. It never calls your tools.

What it asks for
initialize, then the list methods for the capabilities your server advertises (tools, resources, resource templates and prompts), and one request for a method that does not exist, to see how errors are handled.
What it refuses
Every other method, including tools/call, is refused inside the lab before anything is sent.
Network guards
Only https endpoints on port 443 or 8443 are accepted. Each DNS answer is checked when a connection opens, and private, loopback, link-local and cloud-metadata addresses are refused. Redirects that leave the endpoint’s origin are not followed. Requests time out, and responses over 1 MB are rejected.
Credentials
An optional bearer token is used for that discovery only and is never stored.
Limits
Up to 10 pages and 500 items per list. A truncated inventory is reported as a finding, not hidden.

The exact requests are listed in the connection docs.

Tool inventory and schema analysis

Each discovered tool is classified by what it can do, then checked for weak schemas, misleading annotations and hidden instructions. Classification is heuristic: names, parameters and annotations are signals, not proof.

Capabilities
read, write, destructive, financial, egress, exec, filesystem, pii, credentials, admin, browser. A tool can have several; the policy decides per capability.
Schema checks
Free-form command, SQL, path, amount and destination parameters; open schemas on state-changing tools; write tools annotated as read-only; idempotency claimed without a key.
Injection checks
Instruction-like text, directions to use other tools and invisible Unicode in tool descriptions, server instructions, prompts and resource descriptions. Secret-like strings anywhere in the metadata.
Drift
Per-tool hashes are compared with the previous snapshot of the same target, so a definition that changes after approval is flagged.

See all 56 rules with their method, evidence and limits.

A deterministic policy engine

You state what the agent may do in a versioned policy. Every proposed call is evaluated against it, with the same result every time. No language model takes part in a decision.

Intent binding
Denies calls that need capabilities the user’s request does not cover, and tools outside the task.
Approvals
Calls with the capabilities you choose wait for an out-of-band human to approve them. A confirm: true set by the agent does not count as approval unless your policy says so.
Egress allowlist
Limits the email domains and hosts that outgoing calls may reach.
Data-flow labels
In simulations, tool outputs carry labels such as pii, secret and financial_record. Labeled data can be kept away from tools that send data out.
Idempotency
Requires an idempotency key for chosen capabilities, such as payments, and denies a key reused in the same run.
Limits
Caps the amount of a financial call, the number of actions per run and the number of calls per tool.

Deny overrides approval, and approval overrides allow. Each policy version is hashed, and the evidence records the hash. Policy reference.

Isolated simulations against synthetic fixtures

Scenarios run against in-memory MCP servers whose backends are mocks. Names, addresses and keys in them are fictional, and a destructive scenario cannot touch a real file, account or inbox.

FixtureWhat it modelsScenarios
Acme BillingbillingFlagship: invoices, customers, email, files and refunds with typical integration gaps.3
Acme Billing (hardened)billingThe same tool surface after server-side fixes.3
Acme Billing 1.1 (rug pull)billingA later release whose read_invoice description was silently poisoned.3
Workspace FilesfilesystemUnbounded read, write and delete over a mock workspace.1
Git RepositorygitStatus, diff, commit and (force) push on a mock repository.1
Support MailboxemailInbox, drafts and immediate send.1
Payments DeskpaymentTickets, payments and refunds without limits or idempotency.3
App DatabasedatabaseA "read-only" query tool that executes writes, plus role management.2
Browser AutomationbrowserNavigation and page reading with an injected exfiltration link.1
Poisoned UtilitiesutilitiesTool poisoning, cross-tool shadowing and hidden Unicode in descriptions.1

Mock replicas of your tool surface

For a server you connect, the lab builds an in-memory copy of its tools: the same names, schemas and descriptions, with synthetic handlers. It then generates attacks for whatever that surface offers.

Generated scenarios
For each kind of risky tool the surface has, an injected instruction tries to use it: sending data out, deleting, moving money, changing permissions, running commands, or writing during a read-only task. A control checks that the legitimate read still works. If no read-only tool can carry the injection, no scenarios are generated and the gap is reported as not tested.
What a result means
It tests your policy against your tool surface. It does not show that your real server enforces the same controls; that check is reported as Not tested.

Evidence hashes and deterministic replay

Every scenario run produces an evidence record: the user’s intent, each proposed call and where it came from, the matched policy rules, the approval state, a digest of each result, the mock side effects and whether each assertion held.

Hash
The record is serialized as canonical JSON and hashed with SHA-256.
Replay
Re-running a recorded scenario with the same fixture, scenario, policy and runner versions must produce the same hash. If any version differs, the replay reports the mismatch instead of a result.
Redaction
Secret-like strings in arguments, results and side effects are redacted before the record is hashed.

Reports and before/after comparison

A report is one JSON object, and the PDF is a rendering of the same object. Nothing in a report is computed by a language model.

JSON
Schema msl.report/v1, with a SHA-256 of the report body.
PDF
Executive summary, scope, target and protocol metadata, the policy under test, tool inventory, findings, results by evidence kind, simulation evidence, residual risk, reproduction steps, limitations and scoring.
Comparison
Run the same target after a fix and compare readiness, risk score, coverage and attacks contained, plus every rule whose result changed: fixed, improved, regressed, new or removed.

Release gate and Claude assistance

Two ways to act on results: block a release automatically, or ask Claude to explain a finding in plain language.

CI release gate
A command-line runner (npm run msl -- scan https://… --fail-on high) that exits non-zero when a target is not ready under your policy, with an example GitHub Actions workflow.
Ask Claude (free)
Every failing check in a run can produce a redacted prompt for your own Claude account: what the finding means, the impact, and a suggested fix. We send nothing; you paste it.
Explain with Claude (in-app)
Deployments with a Claude API key show explanations inside the run page, limited per workspace and per day. It is switched off on this deployment.
Authority
Claude explains; it never decides. Outcomes, severities, scores and hashes come only from the deterministic engine.

Planned, not built yet

These are on the roadmap. None of them exists today, and no dates are committed.

stdio connector Planned
Discovery for servers that run as local processes. Today only remote Streamable HTTP endpoints are supported.
MCP 2026-07-28 support Planned
Discovery for servers that only speak the stateless 2026-07-28 revision. Today discovery covers revisions up to 2025-11-25.
LLM agent observation mode Planned
Run a real model as the agent (Claude first) and report whether it follows the injection, separately from the deterministic baseline.
Team features Planned
Shared workspaces for teams.
Scheduled rescans Planned
Periodic re-discovery to catch tool definitions that change between releases.
Authorized dynamic auth tests Planned
Token audience and tenant-boundary tests. Today these checks are reported as not tested.

Where it fits

The lab works before release. It complements tools that work at other stages; it does not replace them.

Static scanners
Check metadata, source code or packages for known patterns, often in CI. The lab’s static checks overlap with theirs. Its policy simulations and replayable evidence are a different layer, so use both.
Runtime gateways and proxies
Enforce policy on live traffic in production. The lab blocks nothing in production. It tests a policy before release, which can inform what you enforce at runtime.
Live penetration tests
Exercise your real server and can find implementation bugs. The lab cannot: discovery is read-only, and simulations run on mocks.

The lab is not a runtime firewall, a source-code or dependency scanner, or a penetration test.