MCP Server Security: Access Control and Guardrails for AI Agents

By DharmOps Team•September 22, 2026•11 min read
MCP server security architecture showing OAuth access control, tool-poisoning guardrails, and sandboxing for AI agents

A security review of an MCP deployment usually starts from the wrong assumption: that the risk is a relabeled version of ordinary API security, and OAuth closes most of it. The Model Context Protocol's own specification says otherwise directly.

Authorization is documented as "OPTIONAL for MCP implementations," and servers using the stdio transport are told to skip the OAuth flow entirely and read credentials from the environment instead. That single design choice, made so an individual developer could stand up a local tool server in minutes, is the reason MCP security has become its own category rather than a footnote under API security.

This guide covers what MCP server security actually means, the two attack patterns unique to how agents read tool descriptions, the access-control architecture the spec mandates once authorization is in use, and the guardrail and gateway layer that closes what the protocol leaves open by default.

Prefer to skip the debugging and have an expert handle this? Book a 30-min diagnostic →

What MCP Server Security Actually Covers

MCP server security is two distinct problems, not one. The first is authentication and authorization: verifying who is calling a server and what they're allowed to invoke, governed by the spec's own Authorization and Security Best Practices documents. The second is guardrails: constraining what a tool call can actually do once it's invoked, and what an agent will believe about a tool's behavior, since the model reads a tool's description as trusted instruction text with no protocol-level way to tell an injected command from a legitimate capability description.

Scope this correctly and it excludes some adjacent problems that get folded in by mistake: general LLM prompt-injection defense not specific to tool-calling, model-level content moderation, and general API-gateway concerns like rate limiting and WAF rules. Those are real, but they're not what makes MCP security a distinct discipline. What does is the protocol-specific attack surface: tool poisoning, rug pulls, token passthrough, and the confused-deputy failure mode in proxy architectures, each of which has no close analogue in traditional web-app security.

The Exposure Numbers Behind the 'Optional' Design

The consequence of shipping authorization as optional shows up directly in measurement of the live server population. A 2026 academic study scanning 7,973 remote MCP servers found 3,233 of them, 40.6%, expose tool interfaces with no authentication at all. An earlier, smaller Knostic study manually verified 119 internet-exposed servers discovered via Shodan and found all 119 granted access to internal tool listings without authentication.

Neither figure is a worst-case sample; both are direct measurement of what's actually running. Between January and February 2026, security researchers filed more than 30 CVEs against MCP servers, clients, and supporting tooling, with the most severe, CVE-2025-6514 in the mcp-remote package used by Claude Desktop, VS Code, and Cursor integrations, scoring 9.6 out of 10 on CVSS and affecting a package with 437,000-plus downloads. Read those together and the pattern is an adoption curve running well ahead of a defense curve: Gartner projects 40% of enterprise applications will carry task-specific AI agents by the end of 2026, up from under 5% in 2025, while the servers those agents actually call are, at scale, unauthenticated by default.

Tool Poisoning: When the Tool Description Is the Attack

MCP collapses two channels that security models normally keep separate: data, meaning what a tool returns, and instructions, meaning what the model should do next. Both arrive as the same text stream, read identically by the model. That collapse is what makes a Tool Poisoning Attack possible.

A tool's description field can carry hidden instructions, often wrapped in tags like `<IMPORTANT>`, that the model reads and follows but that never appear in the simplified approval UI a user actually reviews. Invariant Labs' original proof of concept embedded exactly this inside an innocuous-looking `add(a, b)` tool, instructing the model to also read the user's SSH private key and smuggle its contents out through a response parameter, and the model complied with zero user interaction. A related variant, tool shadowing, has a malicious server alter the behavior description of an entirely different, trusted tool from another server; Invariant demonstrated this silently redirecting emails to an attacker's address even when the user explicitly named a different recipient in the request.

Neither pattern is a jailbreak in the usual sense. Both exploit the fact that nothing in MCP distinguishes a legitimate capability description from an injected command.

Rug Pulls: The Attack That Waits for Approval to Expire

A second MCP-specific pattern exploits trust across time instead of across servers. MCP clients re-fetch tool definitions live, at call time, rather than caching the version a user actually approved. That means a server can pass an initial security review with a completely benign description, get approved once, and then silently serve a different, malicious description weeks later on a subsequent call, with no re-prompt forcing the user to notice.

This is the exact mechanism behind CVE-2025-54136, nicknamed MCPoison: commit a benign configuration, wait for it to be approved, then swap the payload. The common misconception it exploits is that reviewing a server once at install time is sufficient. Under MCP's live-fetch model, that assumption is false by construction, and the only real defense is procedural rather than one-time: pin tool definitions by cryptographic hash, and force a fresh, visible re-consent screen any time that hash changes between calls, treating a definition diff the same way a code review treats an unexpected diff in a pull request nobody submitted.

OAuth 2.1, Audience Binding, and the Confused Deputy

When authorization is used, MCP mandates OAuth 2.1: the server acts as a resource server, the client as an OAuth client, and every token request carries an RFC 8707 resource parameter identifying the specific MCP server URI the token is scoped to. Servers must validate that audience claim and reject any token not issued for them. The failure mode the spec calls out by name is token passthrough: an MCP server accepting a client-supplied token and forwarding it unmodified to a downstream API without checking it was actually issued for the MCP server itself.

The spec lists this as explicitly forbidden, for four concrete reasons: it bypasses the downstream service's own security controls, it breaks audit trails since the downstream service sees the client's identity rather than the MCP server's, it collapses a trust boundary that's supposed to exist between services, and it forecloses the ability to add controls later without breaking every integration built on the passthrough assumption. A related and subtler failure lives in proxy architectures specifically: the confused-deputy problem, where a proxy holding one static upstream client ID across many downstream MCP clients can let an attacker register a malicious client, exploit a leftover consent cookie from an earlier legitimate flow, and get an authorization code redirected straight to an attacker-controlled server, skipping consent entirely. The fix is a per-client consent registry checked before any request is forwarded upstream, not a single shared credential treated as good enough for every caller.

Still Piecing This Together Yourself?

A senior engineer looks at your actual setup, not a generic checklist, and tells you exactly what's wrong and how to fix it.

Book a Diagnostic Call

Least-Privilege Scoping Instead of Wildcard Tokens

The MCP spec formalizes scope minimization directly, warning against wildcard or omnibus scopes such as `*`, `all`, or `full-access`, and recommending a minimal initial scope set with incremental step-up elevation via `WWW-Authenticate` challenges rather than requesting everything a server might ever need up front. The anti-pattern the spec names explicitly is publishing every possible scope a server supports "to preempt future prompts" during setup. It feels efficient at configuration time and is exactly backwards: a stolen or misused token then carries a blast radius equal to the entire server's capability set, not the one operation a user actually approved.

This matters most for any MCP server that spans more than one class of resource, the common example being a server that can both read issues and merge changes to a production branch under a single scope. The practical test is straightforward and worth running before trusting any deployment: enumerate the minimum scope set the tools a server actually exposes require, and diff that list against what's configured. Where the two don't match, the gap is the attack surface a compromised or leaked token would actually have.

Sandboxing, SSRF, and Session-Level Failures

Most MCP servers today run locally, spawned as child processes carrying the same privileges as the client itself, which means a compromised local server has full user-level filesystem and network access by default unless something restricts it. The spec's own guidance is to use the stdio transport to limit exposure to the client alone, bind any HTTP transport strictly to 127.0.0.1, and run untrusted server code inside platform-appropriate sandboxing, container or microVM isolation rather than the bare host process. Two transport-level failure modes sit underneath OAuth entirely and are easy to miss if a review stops at "authorization is configured." The first is SSRF during OAuth metadata discovery: a malicious server can point its resource-metadata or token-endpoint fields at internal infrastructure, including cloud metadata endpoints that leak IAM credentials, and a client that follows discovery URLs without validation will fetch them.

The second is session hijacking in stateful HTTP deployments, where a guessable or leaked session ID either lets an attacker inject events into a shared queue a different server instance later delivers as if resumed, or, more simply, lets a session ID stand in for authentication with no re-check of the underlying user. The spec is explicit that sessions must never be used as authentication on their own, and that session IDs need to come from a CSPRNG, not a sequential counter.

The Guardrail and Gateway Layer That Closes the Gap

Closing what the protocol leaves optional is a tooling problem as much as a configuration one, and a handful of purpose-built tools now cover most of it. MCP-Scan, open-sourced by Invariant Labs and now the foundation of Snyk's commercial Agent Scan product, connects to a client's configured MCP servers, retrieves their tool descriptions, and checks them for prompt injection, tool poisoning, cross-origin escalation, and rug-pull drift, making it the closest thing to an industry-standard pre-deployment gate. General-purpose guardrail frameworks like NVIDIA's NeMo Guardrails and content-safety classifiers like Meta's Llama Guard 4 cover an adjacent but different risk, unsafe content in what a tool returns, rather than whether the tool description itself can be trusted.

At the infrastructure layer, gateway products like Kong's AI MCP Proxy and Docker's MCP Gateway centralize OAuth termination, per-tool access-control lists, and audit logging across many servers, so no MCP client talks directly to a server without passing through one enforcement point. A defensible production stack layers these rather than picking one: identity and token issuance from an existing provider, a gateway in front of every server, a pre-deployment scan gate with the resulting tool-description hash stored as the approved baseline, execution isolation for anything locally spawned, and a runtime layer that strips instruction-like tags from tool outputs before they re-enter the model's context.

The OWASP MCP Top 10: Where Everything Above Fits

OWASP opened a dedicated MCP Top 10 project in 2025, still in beta, and it's the first attempt at a standard taxonomy for this attack surface rather than a vendor's private list. The ten categories: MCP01 token mismanagement and secret exposure, hard-coded credentials or long-lived tokens sitting in model memory or protocol logs; MCP02 privilege escalation via scope creep, the wildcard-scope problem described above; MCP03 tool poisoning, the hidden-instruction pattern this guide opened with; MCP04 software supply chain attacks and dependency tampering; MCP05 command injection and execution, an agent building a system command from untrusted input with no sanitization; MCP06 prompt injection via contextual payloads; MCP07 insufficient authentication and authorization, the optional-by-default problem; MCP08 lack of audit and telemetry; MCP09 shadow MCP servers, unapproved deployments running outside any security review with default credentials still set; and MCP10 context injection and over-sharing, a shared context window leaking one task's or user's data into another's. A 2026 survey of 2,614 live MCP servers found 82% exposed to path traversal and 34% exposed to command injection, MCP05 and a close cousin of MCP01, which means the two categories this guide has spent the most time on, authorization and tool poisoning, are real but not the only two gaps worth checking.

A review scoped only to OAuth and tool descriptions will still miss supply-chain and audit-logging failures that the same taxonomy flags as equally common. Scale makes MCP09 harder to dismiss than it sounds: the official MCP Registry, in preview since September 2025, lists close to 2,000 servers, while community directories run far larger, Glama's index alone lists over 31,000 and PulseMCP and mcp.so each list more than 20,000. The large majority of that count is unmaintained experiments and duplicates rather than production software, and the set of genuinely maintained, first-party servers from real vendors is in the low hundreds, not the tens of thousands.

That gap between total listings and actually-maintained servers is exactly the shadow-MCP-server risk MCP09 names: nothing in the protocol stops a team from pointing a production agent at whichever of those thousands of servers looked convenient, with no governance process ever having reviewed it.

Get a Straight Answer, Not a Sales Pitch

Tell us what you're running into. We'll tell you directly what's causing it and what it takes to fix, before you sign anything.

Contact Us

Two More Real Incidents: A GitHub Toxic Flow and an RCE in mcp-remote

Invariant Labs disclosed a second real-world case beyond the SSH-key proof of concept already covered, this one against the official GitHub MCP server, which had roughly 14,000 GitHub stars at the time. The flaw wasn't a bug in the server's code. It was a default configuration: a single access token scoped to both a user's public and private repositories, handed to an agent that reads public repository content, including issues, as untrusted input.

An attacker files an issue on any public repo the victim's agent can see, wraps a prompt-injection payload inside it, and the agent, following what it reads as an instruction, pulls data out of the victim's private repositories and exposes it through an action the attacker actually controls, all without the private repository ever being directly reachable to the attacker. Invariant classified this as a "toxic agent flow": three individually reasonable permissions, read public issues, read private repos, take an externally visible action, that become dangerous only in combination. The fix isn't a patch to the GitHub MCP server; it's separating the token scopes so public, attacker-reachable input and private, sensitive data are never available to the same agent session at once.

The mcp-remote package, used to bridge Claude Desktop, VS Code, and Cursor to remote MCP servers, shipped a more conventional bug: CVE-2025-6514, a CVSS 9.6 OS command injection in the `sanitizeUrlz` function inside `utils.ts`. A malicious MCP server could return a crafted `authorization_endpoint` value in its OAuth metadata response, and that value reached a shell command without adequate sanitization, giving a hostile server arbitrary code execution on the client machine simply because the client connected to it. Every version from 0.0.5 up to 0.1.16 was affected; 0.1.16 carries the fix.

Both incidents share a lesson the OWASP taxonomy above formalizes: the failure is rarely in one component read in isolation, it's in a combination or a code path nobody threatened-modeled because the protocol is new enough that the obvious attack surface hasn't been fully mapped yet.

What to Actually Verify Before Calling a Deployment Secure

Configuration review answers whether the right settings exist on paper. It doesn't confirm any of them actually hold under an adversarial test, and the gap between the two is where most of these incidents live. A review worth trusting runs a live audience-mismatch test against every HTTP-transport server, sending a token minted for a different resource and confirming it's rejected rather than silently accepted.

It runs MCP-Scan or an equivalent against the full server inventory rather than a sample. It captures tool-description hashes as an approved baseline with drift detection wired to force re-consent, and then actually simulates a rug pull, swapping a tool's description after approval, to confirm the client fails closed instead of executing the new definition without comment. It attempts the confused-deputy sequence against any proxy server in the stack.

And it confirms, directly, that a compromised local server cannot read files outside its declared working directory and that OAuth discovery requests to `169.254.169.254` or an internal IP range are actually blocked at the client, not just documented as a known risk in a policy no one has tested against the running system.

Related: Data Governance and Security Engineering, AI Data Infrastructure

MCP's authorization model isn't under-specified. It's specific, testable, and largely ignored by default because the protocol was built to make a local developer server trivial to stand up, and that convenience carried straight into production. The gap that matters isn't "does this deployment use OAuth" — it's whether tokens are audience-bound and validated, whether scopes are minimal rather than wildcarded, whether tool descriptions are hash-pinned against silent rewrites, and whether a locally spawned server is sandboxed instead of running with a user's full privileges.

Every one of those is checkable today, against the spec's own normative language, without waiting for a bigger incident to justify the review.

Frequently Asked Questions

Is MCP server security different from regular API security?

Partly. MCP servers still need standard OAuth 2.1 access control, but two attack patterns have no close analogue in traditional API security: tool poisoning, where a tool's description carries hidden instructions the model follows but a user's UI never shows, and rug pulls, where a server swaps its tool description after approval because clients re-fetch definitions live rather than caching what was reviewed.

Does MCP require authentication by default?

No. The MCP specification states authorization is optional, and servers using the stdio transport are explicitly told to skip the OAuth flow and read credentials from the environment instead. Academic measurement of live remote MCP servers found 40.6% expose tool interfaces with no authentication at all.

What is a tool poisoning attack in MCP?

A malicious or compromised MCP server embeds hidden instructions inside a tool's description, often in tags like `<IMPORTANT>`, that the model reads as legitimate instruction text but that never appear in the simplified approval screen a user sees. Invariant Labs' original proof of concept used this to exfiltrate SSH private keys with zero user interaction.

What is a rug pull attack in MCP?

Because MCP clients fetch tool definitions live at call time rather than caching the approved version, a server can pass initial review with a benign description and later silently serve a different, malicious one. CVE-2025-54136 (MCPoison) is a confirmed real-world instance of this exact pattern.

How should MCP server access be scoped?

Per the spec's own guidance: minimal initial scopes with incremental step-up authorization, never wildcard scopes like `*` or `all-access` published up front. Every token should also carry an RFC 8707 resource parameter binding it to one specific MCP server, and servers must reject tokens issued for a different audience.

What tools exist to scan MCP servers for security issues?

MCP-Scan, from Invariant Labs and now the basis of Snyk's Agent Scan product, is the closest thing to an industry-standard pre-deployment gate, checking configured servers for tool poisoning, prompt injection, and rug-pull drift. Gateway products like Kong's AI MCP Proxy and Docker's MCP Gateway add centralized OAuth enforcement and per-tool access control on top.

Who can audit my MCP server deployment for security issues?

DharmOps runs MCP security reviews under Data Governance and Security Engineering, using the same checklist this guide ends on — live audience-mismatch tests, MCP-Scan against the full server inventory, and a simulated rug-pull — rather than a configuration-only review.

Do I need a security review before putting an MCP server in production?

If it's internet-exposed or handles anything beyond a local, single-user tool, yes — the 40.6% unauthenticated-server figure in this guide is measured from production deployments, not test environments. A Diagnostic Call is the fastest way to find out what your specific deployment is missing.

How much does an MCP security review cost?

It scales with how many servers, transports, and proxy layers are in the deployment, not a flat rate — a single locally-spawned stdio server is a different scope than a multi-server gateway architecture. Start with a Diagnostic Call to get an actual scope, not a guess.

What percentage of MCP servers actually have security vulnerabilities?

One 2026 analysis of 100 public Claude MCP servers found 43% allowed command injection, 30% would fetch any URL handed to them (classic SSRF), and 22% leaked files outside their intended directory. A separate scan found prompt injection in 31% of servers and hardcoded secrets or leaked credentials, like API keys and OAuth tokens, in 28%. These are on top of, not instead of, the 40.6% with no authentication at all.

How do I know if an MCP server is safe to install before connecting it?

Check how it authenticates first: prefer OAuth over a long-lived static token, and be wary of any server that only offers the latter. Then check whether it exposes narrow, purpose-built tools or a generic shell/SQL-execution capability, since the latter carries the blast radius of whatever the host process can reach. Run it through MCP-Scan before connecting it to anything holding real credentials, and don't skip that step just because the server looks popular.

What is "shadow AI" and how does it relate to MCP security?

Shadow AI is the MCP-era version of shadow IT: a server someone on the team connected to production because it looked convenient, with no security review, sitting outside whatever governance process is supposed to catch it. It's exactly what OWASP's MCP09 category names, and it's a bigger problem than the raw registry numbers suggest — community directories list tens of thousands of servers, but the number that are actually maintained by an identifiable vendor is in the low hundreds, and nothing in the protocol stops a team from pointing an agent at whichever of the rest looked useful.

Are MCP servers from the official registry safer than ones from community directories?

Listing in a registry isn't itself a security guarantee on either side — the official MCP Registry (in preview since September 2025) and larger community directories like Glama, PulseMCP, and mcp.so don't run the kind of security audit that would catch tool poisoning or command injection before listing. The practical filter is provenance and maintenance activity, not which directory a server came from: a server from an identifiable vendor with recent commits and a security contact is a materially different risk than an anonymous, unmaintained listing, regardless of which directory indexed it.

What's the difference between an MCP security review and an MCP penetration test?

A configuration review checks whether the right settings exist on paper: OAuth is enabled, scopes look reasonable, sandboxing is documented. A penetration test actually tries to break them: sending a token minted for a different resource to confirm it's rejected, simulating a rug pull to see if the client fails closed, and attempting the confused-deputy sequence against any proxy in the stack. Most MCP deployments that pass a review have never had the second one run against them.

Get a Straight Answer on Your Setup

Tell us what you're running into. We read every message personally and reply within 24 hours with times for a free call.

We'll reply within 24 hours. Your information is never shared.