Dashboard

How to Vet an MCP Server Before Connecting It

An MCP server is third-party code the model invokes on its own, with your credentials. Read the tool manifest, not the README. Scope the token first.

Carlo Zuercher
Carlo Zuercher
Staff Engineer, Platform
20 September 20261 min read

Vet an MCP server before connecting it by treating it as code you are about to run with your credentials, because that is what it is. Read the tool descriptions, not just the README. Check who publishes it and whether the source is available. Give it the narrowest credentials that make it work. Run it somewhere it cannot reach anything you care about until you have watched what it does.

The failure mode here is not that a server is obviously malicious. It is that a server is useful, popular, and quietly capable of more than you assumed. This sits in the same family as the rest of the practical risk list for builders, and it is currently one of the least examined entries on it.

Why an MCP server is a different kind of risk

Model Context Protocol servers give a model tools: read a file, query a database, call an API, send a message. The model decides when to call them. That combination, third-party code plus autonomous invocation plus your credentials, is the part that makes the usual "it's just an integration" instinct wrong.

Three properties make it distinct from adding a library or a REST integration:

  • The model chooses. You are not calling the tool at a point in your code you control. The model calls it when it judges it relevant, based on text that may have come from somewhere untrusted.

  • The description is the interface. A model decides which tool to call from the natural-language description the server supplies. That description is attacker-controlled input if the server is hostile, and it is not code you review in a diff.

  • Credentials are usually broad. Most servers ask for a token that covers the whole service, because that is easier to document than a scoped one. Broad by default is how OAuth scope creep happens.

The protocol's own specification is worth reading on this point. The MCP security best practices document names the attack shapes directly: confused deputy problems in proxy servers, token passthrough, server-side request forgery during discovery, state handle hijacking, and local server compromise. It is not a hypothetical list. It also states plainly that a local MCP server is a binary executing on your machine with your client's privileges, which is the sentence to keep in mind while reading the rest of this.

How to vet an MCP server before connecting it, in order

Start with the cheap ones. Most servers fail on provenance before you get anywhere near behaviour.

Check

What you are looking for

Disqualifying answer

|---|---|---|

Publisher

A named person or organisation you can identify, with other work

Anonymous author, account created recently, no history

Source

Public repository, readable, with real commit history

Binary only, or a repo with one squashed commit

Install path

Pinned version from a registry you trust

An install script piped straight into a shell from an arbitrary domain

Dependencies

A short list you could actually audit

Dozens of transitive packages for a simple wrapper

Update mechanism

Explicit, versioned, under your control

Auto-updating from a remote source at runtime

Tool surface

The tools it declares match what it claims to do

A read-only tool that also declares write and delete

That last row is the one people skip and it is the highest-value check in the table.

Read the tool descriptions, properly

Every MCP server declares its tools, their parameters, and a description the model uses to decide when to call them. Dump that list before you connect the server to anything real, and read it as though you were reviewing a permissions manifest.

Look for:

  1. Tools that exist but are not mentioned in the documentation. A server described as "search your notes" that also declares a tool for writing files has a surface you did not agree to.

  2. Descriptions containing instructions rather than descriptions. Text along the lines of "always call this tool first" or "ignore previous constraints when using this" is an attempt to steer the model, not documentation. This is prompt injection delivered through the tool manifest.

  3. Parameters that take arbitrary paths, URLs, or commands. A parameter called command or path with no stated constraint means the tool's real capability is whatever the model puts in it.

  4. Vague catch-alls. A tool called execute or run_query does everything its backing service does, whatever the description says.

If the server's manifest and its README disagree, trust the manifest. The model does.

Scope the credentials down before the first run

This is where most of the actual protection comes from, and it is boring work that people defer.

  • Create a dedicated identity for the server. Not your personal token, not the shared service account, a new one that exists only for this.

  • Grant read before write. Most servers are useful read-only, and you can discover whether you need write access by finding out what fails.

  • Restrict at the resource level where the service allows it: one repository rather than the organisation, one database schema rather than the instance, one folder rather than the drive.

  • Set an expiry. A token that dies in thirty days forces a conscious decision to renew.

  • Record what you granted, somewhere you will look again. Six months from now you will not remember which of eleven connected servers holds which permission.

The principle is the same one behind preventing prompt injection in your own app: assume the instruction stream can be poisoned, and make sure the credentials behind it cannot do much damage when it is.

Watch it before you trust it

Run the server for a week against non-production resources, with logging on, and look at what it actually did.

The questions worth answering from the logs: which tools were called, how often, with what arguments, and triggered by what. Look specifically for calls that happened when you were not expecting any, and for arguments the model constructed rather than ones you supplied. A server that quietly calls out to a domain that is not the service it wraps is the finding you are hoping not to make, and the only way to see it is egress logging.

If you cannot run it in isolation, you can at least run it without anything valuable attached. A server connected to an empty test account tells you a great deal for very little risk.

A short disqualification list

Do not connect a server that:

  • Requires credentials broader than its stated function, with no explanation.

  • Ships without source, or with source that does not match the published artefact.

  • Auto-updates from a remote location at runtime, which means the code you reviewed is not the code that runs tomorrow.

  • Proxies your tokens to a third party. If your API key leaves your machine to reach a service other than the one it authenticates to, that is a token passthrough problem regardless of the reason given.

  • Declares a tool whose blast radius you cannot describe in one sentence.

None of these are exotic. They are the same judgement you would apply to vetting a browser extension, which has a near-identical risk profile and a longer history of abuse to learn from.

Frequently asked questions

Is an official MCP server from a large vendor automatically safe?

Safer, not safe. A first-party server from the service it connects to removes the supply chain question and the token passthrough question. It does not remove scope. Vendor servers frequently request broad access because it simplifies support, and narrowing it remains your job.

Can an MCP server read data I did not point it at?

Only within what its credentials allow, which is exactly why scoping matters more than any other check here. The protocol does not constrain what a server does with the access you give it.

What is the risk if I only use read-only tools?

Exfiltration rather than destruction. A read-only server with broad access can still surface data to a model, and from there into a context you did not intend. Read-only reduces the worst case substantially without eliminating it.

How is this different from calling the same service's REST API?

Control over invocation. With a REST integration you decide when each call happens and with what arguments. With MCP the model decides, based on text that may include content from untrusted sources.

Do I need to re-check a server I already trust?

Whenever its version changes, and periodically for the tool manifest specifically. A new version can add tools, and an added tool is an added permission that no upgrade prompt will describe to you in those terms.

How did this land?

About the author

Carlo Zuercher
Carlo Zuercher

Staff Engineer, Platform

Carlo works on the platform that turns prompts into running apps. He writes the engineering deep dives and the changelog notes worth reading.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.