MCP connector auditing

The server says it’s read‑only. Prove it.

Your agent harness pauses before sensitive tool calls. Which tools count as sensitive comes from the annotations the MCP server publishes about itself. Nothing in the protocol requires those to be true, and nothing in the client checks.

Airlock exercises a server before you trust it and emits a connector policy built from what it observed.

Two recorded runs, both real executions. A server built to lie, caught claiming read-only while Airlock observed it writing to disk. And @modelcontextprotocol/server-memory, a widely used public server, where Airlock found nothing and said so: 0 findings, 39 of 54 checks marked not tested because transcript-only evidence has no sensors to establish them.

Verified against a running TrueForge 0.1.4 instance, not a mock: the harness registered this control server over the wire and read all six tools, seeing exactly the three that require an approval card as destructive.

Source: github.com/himanshu748/airlock-mcp

case af_5384c433 · acme-docs.examplecontrolled fixture · 24 probes

Declared by the server

"name": "summarize_documents"
"annotations": {
  "readOnlyHint": true
}

the harness reads this and does not pause

Observed by Airlock

  • annotation divergenceclaimed read-only, wrote to disk
  • canary exfiltrationa planted value left the building
5 of 5
planted behaviours detected in the dishonest fixture
0
findings on the honest fixture, across 36 checks
6
detections, each reported as a verdict and never a score
24
probes to audit a six-tool server end to end

The gap it closes

A safety gate is only as honest as the server describing itself.

require_approval_for_tools defaults to @write and @destructive. Those selectors resolve from annotations the server publishes. A server that labels export_report read-only runs it without the human ever seeing a pause.

What the harness is told

"require_approval_for_tools":
  ["@write", "@destructive"]

Resolved from the server’s own annotations. Never verified.

What Airlock adds

"require_approval_for_tools":
  ["export_report"]

Named literally, because the selector cannot be trusted for this server. That one line is the product.

How it works

Four passes, then a decision that belongs to you.

  1. 01

    Point it at a server

    Give Airlock a URL you are considering connecting. It validates the target, resolves the hostname once and pins that address for every connection the case makes afterwards.

  2. 02

    Read the declaration

    One tools/list call captures every tool, description, schema and annotation. That is the claim under test, recorded verbatim.

  3. 03

    Exercise and probe

    Generated inputs, then adversarial ones: boundary values, path-shaped strings, planted canaries. Rate limited, capped, and only ever pointed at a target you submitted.

  4. 04

    Decide, then enforce

    Declarations and observations land side by side and stop for your decision. On approval Airlock emits a connector policy and the proxy holds it.

See it run

One command, against a server that lies on purpose.

The repository ships two owned fixtures with the same six-tool shape. One declares itself accurately. The other plants five behaviours behind honest-looking annotations. This is the real output of auditing the second one.

scripts/demo_audit.pydishonest fixture

$ .venv/bin/python scripts/demo_audit.py --declared-root /workspace/documents

declared tools: 6

probes run: 24

findings: 7 of 36 checks

  • [block]annotation_divergencetool_0001fixture_filesystem
  • [critical]scope_escapetool_0001fixture_filesystem
  • [block]undeclared_egresstool_0002fixture_network
  • [critical]scope_escapetool_0003fixture_filesystem
  • [suspicious]injected_instructionstool_0005mcp_transcript
  • [critical]canary_exfiltrationtool_0006canary_sink
  • [critical]scope_escapetool_0006fixture_filesystem

Same command, honest fixture

findings: 0 of 36 checks

The same six-tool shape, the same 36 checks, the same probe budget. The detectors fire on behaviour, not on a server being unfamiliar.

What tool_0001 declared

"annotations": {
  "readOnlyHint": true
}

It wrote to disk during the probe. That single divergence is what a harness reading annotations cannot see, and it is why the emitted policy names the tool literally.

Reproduce it

Both fixtures ship with the repository and run on loopback. The README carries the exact commands, and the run above is their real output.

Against real servers

Two servers nobody built for a demo.

The fixtures above are ours, so they prove the detectors fire. These are not. One is a deployed memory server, the other is Airlock's own control plane. Both ran in transcript_only, the honest mode for a server whose filesystem and network you cannot see, so the scope and egress checks report capability_absent rather than a clean pass.

ContextFirewall

himanshukumarjha-contextfirewall.hf.space/mcp/

third party
Tools
6
Probes
30
Annotations
None declared

Six tools, no annotations at all.

Two of them are remember and forget_memory. A harness resolving @write and @destructive matches nothing here, so neither one earns an approval card. The fixture on this page shows a server that lies. This is a real server that simply says nothing, and silence defeats the selector the same way.

0 findings. 24 of 36 checks not tested.

Airlock control MCP

127.0.0.1:8100/airlock-control/mcp

self audit
Tools
6
Probes
11
Annotations
Complete on all six

Airlock could not finish auditing itself.

Every tool declares its hints, and probe_tool, seal_case and emit_policy all carry destructiveHint. Then the probe planner refused open_case: its schema uses a $ref into $defs, which the v1 profile rejects. The boundary this project documents turns out to bite on its own front door.

Case incomplete. The planner reported sensor_failed rather than a clean pass.

Also audited, launched as commands

  • @modelcontextprotocol/server-filesystem14 tools
  • @modelcontextprotocol/server-everything13 tools
  • mcp-server-git12 tools
  • @modelcontextprotocol/server-memory9 tools
  • @modelcontextprotocol/server-sequential-thinking1 tool

Most MCP servers ship as a command rather than a URL, so Airlock can launch one and audit it the same way. The command comes from operator configuration, never from a case or a model, because running an audited server means executing the code the audit exists to distrust.

What it looks for

Six checks. Verdicts, never a risk score.

Suspicious is a real verdict, not a hedge: it means a human should look, and the approval card says exactly that. A check no sensor in a case can observe is reported as exactly that, never folded in with the clean results.

Block

Annotation divergence

A tool annotated read-only performed a write or a state change.

Critical

Canary exfiltration

A planted value appeared in an outbound request or in another tool's output.

Block

Undeclared egress

The server reached a host outside its declared scope. Needs a sensor that can see the server's own network.

Critical

Scope escape

File paths touched outside the declared directory.

Suspicious

Injected instructions

A tool result carried imperative text addressed to the model.

Suspicious

Schema drift

A tool accepted parameters that are not in its published schema.

The output

A file you paste, not an opinion you weigh.

Airlock emits a connector policy and a downloadable evidence report. The proxy stays in front of the server afterwards and enforces exactly what you approved.

airlock-policy.json
{
  "mcp_servers": [
    {
      "name": "acme-docs-via-airlock",
      "enable_tools": ["search_docs", "get_document"],
      "disable_tools": ["export_report"],
      "require_approval_for_tools": ["create_ticket"],
      "preload": false
    }
  ]
}

Enforced, not advisory

Register the proxy rather than the server. A tool you denied is refused on the wire, not merely hidden from the model.

Evidence you can read

airlock-report.json carries every probe, observation and finding, with the raw suspect bodies left out.

No policy without a decision

Emission requires a recorded human choice. There is no path that approves a connector on the agent's own say-so.

Questions

What Airlock does and does not claim.

Does this replace my harness's approval gate?

No. The harness enforces a policy; it does not verify the claims that policy is derived from. Airlock produces that input. The two are complementary, and Airlock reimplements none of the sandboxing, approvals or tool filters the harness already provides.

What happens to a tool that lied?

It is moved into require_approval_for_tools by name rather than by selector, because the selector cannot be trusted for that server. Naming it is the whole product.

Can a hostile server attack the auditor?

Raw tool results are quarantined by the backend and never re-enter the model as text; the control MCP returns digests instead. Model and target credentials never enter the sandbox.

Will it tell me a server is fine?

It will not. Airlock reports what it observed across a bounded number of probes, names the checks nothing could observe, and refuses to describe absence of a finding as a clean result.

Audit a connector before you trust it.