Know what an MCP server does before your agent connects to it.

Preflight is a trust layer for the Model Context Protocol. Claude agents install each server in an isolated sandbox, call every tool, and probe for prompt injection, over-broad permissions and undeclared data egress. Each inspection produces a report pinned to one version, with the evidence attached.

Stage
Pre-launch
Based in
Istanbul, Türkiye
Funding
Bootstrapped
Builds on
Claude

Preflight

Inspection seal

Nº 0001

Server
com.example/notes-server
Version
1.4.21.4.3 · description changed
  • Tools enumerated through a live MCP client
  • Descriptions and outputs probed for injection
  • Permissions compared with declared scope
  • Network egress compared with declared hosts
  • Canary secrets never echoed back

Inspected by
Claude test agent

InspectedSeal broken
Re-inspection queued

Illustrative specimen for a fictional server, not a real inspection.The tool description changed after inspection, so the seal is void until Preflight inspects the new version.

Agents trust MCP servers on metadata alone.

The official MCP Registry tells you who published a server and where it runs. It does not record what the server does once your agent calls it. We indexed the registry to measure that gap.

MCP Registry snapshot, 5 October 2026
Measured in the registryServers
Active servers indexedOut of 4,800 records received. We excluded 34 inactive entries and quarantined 7 that had no safe link or transport.4,759100%
Expose a remote endpointCalls to these servers send your agent’s data to a host someone else runs.4,31691%
List no source repositoryThere is no public code to read before you connect.2,67556%
Ship an installable packagenpm 472 · PyPI 90 · MCPB 35 · OCI 22 · NuGet 1360513%
Behavioural test results on recordThe registry schema has no field for them. That empty column is what Preflight fills.0

Source: official MCP Registry API, observed 5 October 2026. The crawl stopped at a checkpoint after 4,800 records, so these figures come from a partial traversal. We counted metadata only and executed no server.

How Preflight inspects a server

Every inspection runs the same six steps and keeps the evidence from each one. A report states nothing that the evidence does not show.

  1. 1/6

    Ingest

    Read the registry record for each server: namespace, version, packages and remotes. Access is read-only and refreshes at most hourly.

    Registry record, content hash

  2. 2/6

    Sandbox

    Install the package or connect to the remote inside an isolated container. It holds canary credentials only, and every outbound connection is logged.

    Image digest, egress log

  3. 3/6

    Exercise

    A Claude test agent connects as a real MCP client. It lists every tool and calls each one with ordinary inputs and hostile ones.

    Full tool-call transcript

  4. 4/6

    Probe

    Run targeted checks: injection in tool descriptions and outputs, tool shadowing, scope wider than declared, undeclared hosts, and secrets echoed back.

    Probe inputs and responses

  5. 5/6

    Adjudicate

    A second Claude pass re-reads every finding against the raw evidence. A finding with no evidence behind it is dropped or marked inferred.

    Adjudication notes with citations

  6. 6/6

    Publish

    Issue a report pinned to one version. Maintainers see findings first. A new version or a changed tool description breaks the seal and queues a fresh inspection.

    Signed report, version diff

Built on Claude, end to end

Preflight is designed so that a Claude model makes every judgement, and each job goes to the model that fits it best. The test agent and adjudication pass are in build. Humans make the disclosure decisions.

Claude models and their roles in Preflight
Claude HaikuTriage clerkClassifies every registry entry, diffs tool descriptions between versions, and decides which servers need inspection next.Every entry, several times a day (planned)
Claude SonnetTest agentDrives the sandbox through the Claude Agent SDK. It connects to the server under test, decides which tools to call and how, reads the outputs and writes findings.Every inspection (planned)
Claude OpusAdjudicatorChecks each flagged finding against the evidence, assigns severity, and writes the report the maintainer receives.Flagged findings only (planned)

What the pipeline is built on

Claude Agent SDK
Will run the test agent’s loop, its sandbox tools and its permission boundaries.
Tool use with MCP
Claude will reach the server under test the same way your agent would.
Prompt caching
The inspection rubric and probe library will be cached once and reused on every run.
Message Batches
Will re-triage the whole registry every night at batch pricing.
Structured outputs
Every finding and report will follow one JSON schema.
Citations
Every sentence in a report will point to the evidence behind it.

Why Claude is the right inspector

  • The inspector is the target

    MCP servers are written to be called by agents like Claude. If a tool description can redirect a Claude agent in our sandbox, that is a demonstrated risk for every Claude user who installs the server, not a hypothetical one.

  • Long, stateful tool use

    One inspection means dozens of tool calls, and the agent has to carry state across all of them. This is the agentic work Claude is built for. Here a dropped thread produces a missed finding.

  • Resistance where it counts

    The test agent reads hostile input on purpose, so we need a model trained to resist prompt injection. When it does fail, the failure becomes a finding.

  • One family, three price points

    Haiku, Sonnet and Opus share one API and behave consistently, so we route each job by difficulty. With caching and batches, a full pass over the registry stays affordable for a bootstrapped company.

What a report looks like

Reports are pinned to one version and written for the maintainer first. The rule under each piece of evidence shows how far it goes.

verified
checked against raw sandbox evidence
inferred
a reasoned conclusion that is not directly observed

Preflight inspection report

Illustrative format · fictional server

Server
com.example/notes-server
Version
1.4.2
Transport
stdio · npm
Inspected by
Claude test agent · Claude adjudicator
  • High

    The description of search_notes tells the agent to read ~/.ssh/config and add it to the query.

    Description text v1.4.2 · transcript call 14· verified

  • Medium

    The first tool call opens a connection to an analytics host that the package does not declare.

    Egress log · 2 connections· verified

  • Low

    export_notes accepts any file path. 40 probes showed no traversal, but its scope is wider than the package declares.

    Input schema · probe set P-07· inferred

  • Pass

    None of the 212 tool outputs echoed the canary token.

    Output scan · 212 responses· verified

Where Preflight is today

Preflight is pre-launch. We have no customers and no public reports yet, and this page shows no numbers we do not have.

Done

  • Ingestion and normalisation of the official MCP Registry: 4,759 active servers indexed, with de-duplication and quarantine rules (snapshot of 5 October 2026, partial traversal).

Building

  • Sandbox runner: container isolation, egress logging and canary credentials.
  • Claude test agent on the Agent SDK, with the probe library.
  • Report schema v0 and the adjudication pass.

Planned

  • The first 100 public reports, starting with the most-installed npm and PyPI servers.
  • Coordinated disclosure with maintainers before every publication.
  • A CI check for MCP authors that runs on every release.
  • A pre-install check for Claude Code users.
  • An allowlist API for platform and security teams.

What we’re asking Anthropic for

Preflight is built Claude-first. Every model call in the inspection pipeline goes to Claude.

API credits
A full registry pass is thousands of Sonnet agent sessions plus Opus adjudication, and it repeats whenever servers change. With credits we can inspect the whole registry instead of a sample, and publish the first 100 reports.
Technical guidance
We want a review of our injection probe library and eval design, and advice on keeping a test agent useful while it reads hostile input.
A disclosure path
We need a contact for coordinating with Anthropic when a server puts Claude users at risk, so fixes land before reports go public.
Ecosystem reach
We would like introductions to the MCP and Claude Code teams, so findings reach developers at the moment they install a server.

Responsible by design

  • Maintainers hear first

    Maintainers receive findings privately, with a fixed disclosure window before anything is published.

  • Canary credentials only

    No real user data or secrets ever enter the sandbox.

  • Evidence or drop

    A finding with no evidence behind it is removed or labelled inferred.

  • Registry terms respected

    We aggregate CC0 metadata read-only and refresh at most hourly.

  • No paid scores

    Nobody can pay to change a finding, and popularity never changes severity.

The founder

Ali Murat Umutlu

Founder, Umutlu AI Labs

Ali Murat Umutlu is a software engineer and founder in Istanbul who builds AI-native products and teaches builders to ship them.

  • Consulted on mission-critical telecom portals and distributed API architecture at Ericsson (2019–2020).
  • Frontend consulting on a large-scale AI analytics platform at Afiniti (2022–2023).
  • Shipped production web apps serving more than 18 million users.
  • Spoke at the UN AI for Development forum in Vienna (2024) on using AI to close the digital divide.
  • Built the pipeline that indexes the official MCP Registry, the dataset Preflight starts from.