# Umutlu AI Labs > Preflight: A trust layer for Model Context Protocol servers, built on Claude. Preflight is a trust layer for the Model Context Protocol. Claude agents install each server in an isolated sandbox, call every tool, and probe for prompt injection, over-broad permissions and undeclared data egress. Each inspection produces a report pinned to one version, with the evidence attached. ## Company facts - Company: Umutlu AI Labs - Product: Preflight - Founder: Ali Murat Umutlu, Founder - Headquarters: Istanbul, Türkiye - Incorporated: Türkiye - Industry: Developer tools · AI security - Funding: Bootstrapped, no outside capital - Stage: Pre-launch, building - Builds on: Claude, by Anthropic - Contact: murat@umutlu.net - LinkedIn: linkedin.com/in/alimuratumutlu ## The problem, measured Snapshot of the official MCP Registry (https://registry.modelcontextprotocol.io/v0.1/servers), observed 5 October 2026, partial traversal: - 4,759 active servers indexed (4,800 records received) - 4,316 (91%) expose a remote endpoint - 2,675 (56%) list no source repository - Registry metadata has no field for behavioural test results ## How Preflight inspects a server 1. Ingest: Read the registry record for each server: namespace, version, packages and remotes. Access is read-only and refreshes at most hourly. Evidence kept: Registry record, content hash. 2. Sandbox: Install the package or connect to the remote inside an isolated container. It holds canary credentials only, and every outbound connection is logged. Evidence kept: Image digest, egress log. 3. Exercise: A Claude test agent connects as a real MCP client. It lists every tool and calls each one with ordinary inputs and hostile ones. Evidence kept: Full tool-call transcript. 4. Probe: Run targeted checks: injection in tool descriptions and outputs, tool shadowing, scope wider than declared, undeclared hosts, and secrets echoed back. Evidence kept: Probe inputs and responses. 5. Adjudicate: A second Claude pass re-reads every finding against the raw evidence. A finding with no evidence behind it is dropped or marked inferred. Evidence kept: Adjudication notes with citations. 6. Publish: Issue a report pinned to one version. Maintainers see findings first. A new version or a changed tool description breaks the seal and queues a fresh inspection. Evidence kept: Signed report, version diff. ## What we are building on Claude - Claude Haiku (Triage clerk): Classifies every registry entry, diffs tool descriptions between versions, and decides which servers need inspection next. Runs: Every entry, several times a day (planned). - Claude Sonnet (Test agent): Drives the sandbox through the Claude Agent SDK. It connects to the server under test, decides which tools to call and how, reads the outputs and writes findings. Runs: Every inspection (planned). - Claude Opus (Adjudicator): Checks each flagged finding against the evidence, assigns severity, and writes the report the maintainer receives. Runs: Flagged findings only (planned). Claude platform features used: - Claude Agent SDK: Will run the test agent’s loop, its sandbox tools and its permission boundaries. - Tool use with MCP: Claude will reach the server under test the same way your agent would. - Prompt caching: The inspection rubric and probe library will be cached once and reused on every run. - Message Batches: Will re-triage the whole registry every night at batch pricing. - Structured outputs: Every finding and report will follow one JSON schema. - Citations: Every sentence in a report will point to the evidence behind it. ## Status - Done: Ingestion and normalisation of the official MCP Registry: 4,759 active servers indexed, with de-duplication and quarantine rules (snapshot of 5 October 2026, partial traversal). - Building: Sandbox runner: container isolation, egress logging and canary credentials. - Building: Claude test agent on the Agent SDK, with the probe library. - Building: Report schema v0 and the adjudication pass. - Planned: The first 100 public reports, starting with the most-installed npm and PyPI servers. - Planned: Coordinated disclosure with maintainers before every publication. - Planned: A CI check for MCP authors that runs on every release. - Planned: A pre-install check for Claude Code users. - Planned: An allowlist API for platform and security teams. Preflight is pre-launch. There are no customers or public reports yet. ## Support requested from Anthropic - API credits: A full registry pass is thousands of Sonnet agent sessions plus Opus adjudication, and it repeats whenever servers change. With credits we can inspect the whole registry instead of a sample, and publish the first 100 reports. - Technical guidance: We want a review of our injection probe library and eval design, and advice on keeping a test agent useful while it reads hostile input. - A disclosure path: We need a contact for coordinating with Anthropic when a server puts Claude users at risk, so fixes land before reports go public. - Ecosystem reach: We would like introductions to the MCP and Claude Code teams, so findings reach developers at the moment they install a server. ## Principles - Maintainers hear first: Maintainers receive findings privately, with a fixed disclosure window before anything is published. - Canary credentials only: No real user data or secrets ever enter the sandbox. - Evidence or drop: A finding with no evidence behind it is removed or labelled inferred. - Registry terms respected: We aggregate CC0 metadata read-only and refresh at most hourly. - No paid scores: Nobody can pay to change a finding, and popularity never changes severity. ## Links - Website: https://web-production-956c8.up.railway.app - Company facts (JSON): https://web-production-956c8.up.railway.app/company.json - Founder on LinkedIn: https://www.linkedin.com/in/alimuratumutlu - Contact: murat@umutlu.net Umutlu AI Labs is independent and is not affiliated with or endorsed by Anthropic.