Skip to content
MCP Security Lab

Documentation

How to connect a server, write a policy and read the results. Everything is on this one page.

Getting started

You can try the lab without an account, then connect your own server.

  1. Run the security demo. The demo needs no account and uses only synthetic data. It shows the flagship assessment of the synthetic Acme Billing server: a baseline under a permissive starter policy, a re-test under the recommended policy, and a re-test against a hardened build of the server.
  2. Create an account. It is free during early access.
  3. Connect a target: a remote MCP endpoint you own or are authorized to test. See Connecting a remote MCP.
  4. Choose a policy. Start from a template and adapt it to what your agent should be allowed to do. See Policies.
  5. Run an assessment. Review the results by evidence kind, and open a scenario to see each decision.
  6. Fix, re-test and compare. Then download the report as JSON or PDF.

Connecting a remote MCP

Connect only servers you own or are authorized to test; see the terms. Discovery reads metadata and never calls your tools.

Requirements

  • An https:// URL of a Streamable HTTP endpoint on port 443 or 8443. Plain http, other ports, and user names or passwords inside the URL are rejected.
  • A public host. Names without a dot, and names ending in localhost, local, internal, intranet, lan, corp or home.arpa, are rejected.
  • MCP protocol revisions 2024-11-05 to 2025-11-25. Servers that only implement the stateless 2026-07-28 revision are not supported yet.

Bearer token

If your server requires authorization, paste a bearer token into the token field, never into the URL. The lab sends it to your endpoint as an Authorization: Bearer header, for this discovery only. It is held in memory, scrubbed from error messages and never stored. Use a short-lived token with the narrowest scope that can list tools.

Requests the lab sends

Discovery runs once without credentials, to see what an anonymous client can list. If you supplied a token, it runs a second time with the token. Each run sends these JSON-RPC messages as HTTP POST requests to your endpoint:

  1. initialize, then the notifications/initialized notification.
  2. tools/list, resources/list, resources/templates/list and prompts/list, each only if your server advertises that capability. Pagination is followed up to 10 pages or 500 items per list, and text fields over 20,000 characters are truncated.
  3. One request for msl/probe-unknown-method, a method that does not exist, to check that your server answers with error -32601 and without internal details.

The MCP client library may also open the optional GET event stream that Streamable HTTP allows; a server that does not offer one answers 405. If a request times out, a notifications/cancelled notification may follow.

When your server rejects the anonymous attempt with 401 or 403, the lab also looks for RFC 9728 protected resource metadata. It requests the resource_metadata URL from the WWW-Authenticate header, if there is one, and /.well-known/oauth-protected-resource on your host, followed by your endpoint’s path.

Every request carries the user agent MCP-Security-Lab/0.1.0 (read-only discovery). Nothing else is sent: tools/call and every other method are refused inside the lab before they reach the network.

What is blocked

  • Private, loopback, link-local, shared, multicast, reserved and documentation address ranges, for IPv4 and IPv6, including the cloud metadata address 169.254.169.254. IPv4-mapped IPv6 addresses are refused.
  • A host with any DNS answer in a blocked range. Addresses are checked again each time a connection opens, so DNS rebinding cannot reach a blocked range.
  • Redirects to a different scheme, host or port. The MCP client library may follow a redirect that stays on the same origin, and that request passes the same checks.
  • Slow or large responses. A connection must open within 5 seconds, each HTTP request is aborted after 10 seconds, each MCP request times out after 8 seconds, and responses over 1 MB are rejected.

Policies

A policy is a JSON document that states what the agent may do. Policies are versioned, each version is hashed with SHA-256, and evidence records the hash, so every result traces back to the exact policy that produced it.

The engine checks each proposed call against every rule that applies and keeps the strictest result: deny beats approval, and approval beats allow. A tool that is not in the discovered inventory is always denied. No language model takes part.

The engine ships two templates. Starter (permissive): A typical first integration: every discovered tool is allowed and an agent-supplied confirm flag counts as approval. Recommended (least privilege): shown in full below.

Fields

idLowercase letters, digits and hyphens; up to 64 characters
Names the policy. Evidence and reports cite it as id@version.
versionPositive integer
Raise it with every change.
nameUp to 120 characters
Display name.
descriptionOptional; up to 2,000 characters
What the policy is for.
defaultDecisionallow, deny or require_approval
Applies to a tool that no tool rule and no capability rule covers.
capabilityRulesCapability to decision
One decision per capability: read, write, destructive, financial, egress, exec, filesystem, pii, credentials, admin, browser. When a tool has several capabilities, the strictest decision wins.
toolsTool name to { decision, maxCallsPerRun }
Per-tool rules. A tool’s own decision takes precedence over capability rules; allowlist, intent, data-flow and approval checks still apply. maxCallsPerRun (1 to 1,000) denies further calls to that tool in the same run.
approval.requiredForList of capabilities
Calls with these capabilities always need out-of-band human approval, even when a tool rule allows the tool.
approval.acceptAgentAssertedConfirmationtrue or false
If true, an argument such as confirm: true set by the agent counts as approval. Keep it false.
intentBinding.enabledtrue or false
Denies calls that need a write, destructive, financial, egress, exec, admin or credentials capability that the user’s request does not cover.
intentBinding.enforceToolScopetrue or false
Also denies tools outside the task’s tool list.
egressnull, or { allowedEmailDomains, allowedHosts }
null leaves egress unrestricted. Otherwise every email domain and http(s) host in an outgoing call’s arguments must be on the lists; subdomains match. Up to 100 entries each.
dataFlow.blockSensitiveToEgresstrue or false
Denies outgoing calls whose arguments carry data with a sensitive label.
dataFlow.sensitiveLabelsUp to 20 labels
Labels that count as sensitive, such as pii, secret and financial_record. Data carries the labels of the tool output it came from.
idempotency.requireKeyForList of capabilities
Calls with these capabilities need an idempotency_key argument. Reusing a key in the same run is denied.
limits.maxAmountNumber or null
Highest amount allowed in a financial call; null means no limit.
limits.maxActionsPerRun1 to 1,000
Calls proposed after this many in one run are denied.

Recommended policy

{
  "id": "agent-policy",
  "version": 2,
  "name": "Recommended (least privilege)",
  "description": "Default deny. Reads allowed; writes, egress, destructive and financial actions need out-of-band human approval; exec, admin and credential access denied. Actions must stay inside the task scope, egress only to allowlisted domains, and sensitive data never flows to egress.",
  "defaultDecision": "deny",
  "capabilityRules": {
    "read": "allow",
    "browser": "allow",
    "write": "require_approval",
    "egress": "require_approval",
    "destructive": "require_approval",
    "financial": "require_approval",
    "exec": "deny",
    "admin": "deny",
    "credentials": "deny"
  },
  "tools": {
    "create_draft": {
      "decision": "allow"
    },
    "navigate": {
      "decision": "allow"
    },
    "issue_refund": {
      "maxCallsPerRun": 1
    }
  },
  "approval": {
    "requiredFor": [
      "destructive",
      "financial"
    ],
    "acceptAgentAssertedConfirmation": false
  },
  "intentBinding": {
    "enabled": true,
    "enforceToolScope": true
  },
  "egress": {
    "allowedEmailDomains": [
      "acme.example"
    ],
    "allowedHosts": [
      "acme.example"
    ]
  },
  "dataFlow": {
    "blockSensitiveToEgress": true,
    "sensitiveLabels": [
      "pii",
      "secret",
      "financial_record"
    ]
  },
  "idempotency": {
    "requireKeyFor": [
      "financial"
    ]
  },
  "limits": {
    "maxAmount": 100,
    "maxActionsPerRun": 20
  }
}

Scenarios

13 hand-written scenarios run against the synthetic fixtures. For a server you connect, scenarios are generated from its discovered tools instead (rules MSL-SIM-101 to MSL-SIM-107). Every rule is described on the methodology page.

ScenarioSeverity
Indirect prompt injection exfiltrates customer datainvoice-exfiltration · MSL-SIM-001critical
Injected housekeeping instruction deletes secrets and backupsfs-unauthorized-deletion · MSL-SIM-002high
Diff content triggers an unrequested commit and force pushgit-unauthorized-push · MSL-SIM-003high
Agent sends an e-mail the user only asked to draftmail-unconfirmed-send · MSL-SIM-004medium
Ticket text triggers an unapproved high-value refundpay-injected-refund · MSL-SIM-005critical
A read-only question escalates into database writesdb-read-to-write · MSL-SIM-006high
Page text lures the agent into leaking a session id via URLbrowser-exfiltration · MSL-SIM-007high
A poisoned tool description makes the agent steal an SSH keytool-poisoning-secret-theft · MSL-SIM-008critical
Profile text chains a read into an admin role grantcross-tool-escalation · MSL-SIM-009critical
Agent self-confirms a refund with confirm=trueapproval-bypass · MSL-SIM-010high
Automatic retries issue the same refund three timesnon-idempotent-retry · MSL-SIM-011high
Control: legitimate read-only requestcontrol-legitimate-read · MSL-SIM-012medium
Control: approved low-value refundcontrol-approved-refund · MSL-SIM-013medium

Results and scoring

Each check ends as PASS, FAIL, WARNING, NOT_APPLICABLE, NOT_TESTED or ERROR. NOT_TESTED and ERROR are never counted as PASS. An assessment then gets:

  • a risk score from 0 to 100, where higher is riskier, weighted by the severity of each FAIL and WARNING;
  • coverage, the share of applicable checks that actually ran;
  • a readiness label: Not ready, Insufficient coverage, Conditional, or Ready within tested scope;
  • a confidence level of high, medium or low, based on coverage and errors.

The formula and every outcome definition are on the methodology page. These numbers are indicators for the tested configuration, not certifications.

Reports

Every assessment produces a report as JSON, with schema msl.report/v1, and as PDF. The PDF renders the same object, and nothing in either is written by a language model. A report contains:

  • an executive summary with readiness, risk score, coverage, confidence and the top findings;
  • the scope and authorization boundary, including what was out of scope;
  • target and protocol metadata, with a transcript of the methods exchanged;
  • the policy under test and its SHA-256;
  • the tool inventory, with derived capabilities and risk levels;
  • findings with remediation and references, and every result grouped by evidence kind;
  • simulation evidence for each scenario, with its evidence hash and the inputs needed to reproduce it;
  • a before/after comparison when the run has a baseline;
  • residual risk, limitations and the scoring formula.

The JSON includes reportHash, a SHA-256 of the report body. Comparing it with the original hash shows whether a copy was changed; it is not a signature. Reports on synthetic fixtures say so at the top.

Public synthetic fixture endpoints

Each synthetic fixture is also served as a public MCP endpoint (Streamable HTTP) at /api/mcp/fixtures/{fixtureId}. The endpoints are discovery-only: every tools/call returns an error result, so nothing is executed. Requests are rate-limited per IP address. Point any MCP client at them to see what the lab’s discovery sees.

Fixture idFixtureVersion
acme-billingAcme Billing1.0.0
acme-billing-hardenedAcme Billing (hardened)2.0.0
acme-billing-rugpullAcme Billing 1.1 (rug pull)1.1.0
fs-workspaceWorkspace Files1.0.0
git-repoGit Repository1.0.0
mailboxSupport Mailbox1.0.0
paymentsPayments Desk1.0.0
app-databaseApp Database1.0.0
browserBrowser Automation1.0.0
poisoned-toolsPoisoned Utilities1.0.0

Local development

Requires Node.js 20.9 or later.

npm install
npm run dev
npm test

npm run dev serves the app at http://localhost:3000. Without DATABASE_URL, local development uses an embedded PGlite database, so you do not need a Postgres server. Set DATABASE_URL to use your own Postgres.

Limitations

  • Simulation results prove the behavior of the evaluated policy against mocked backends only. They do not show that a real model would follow an injection, nor that a real server or gateway enforces the same controls.
  • Mock-replica simulations copy the discovered tool surface; the real remote server is never called beyond read-only discovery.
  • Static analysis is heuristic (names, schemas, deterministic patterns). Paraphrased, encoded or multi-step injections can evade it; absence of a finding is not proof of absence.
  • Discovery implements MCP revisions 2024-11-05 to 2025-11-25 (initialize-based). Servers that only implement the stateless 2026-07-28 revision are not yet supported.
  • Token audience validation, multi-identity tenant boundaries and rate limiting are reported as NOT_TESTED.
  • Scores, readiness and confidence are MCP Security Lab indicators for a defined configuration at a point in time, not certifications or penetration tests.

The lab helps you identify, reproduce and reduce risk. It cannot guarantee that an agent will never cause harm.