MCP Security: The Risks Every CTO Should Understand Before Going Live

Connecting AI agents to your systems through MCP adds a few new ways for data to leak or actions to misfire, and the teams that ship safely are the ones that ask the right questions before launch rather than after an incident.

tech content13 min read

The Model Context Protocol lets AI assistants call your systems as tools (if you need the background first, start with our plain-English guide to MCP servers). The protocol itself is sound, but it changes who decides what data gets combined and shown, and it opens a few routes for instructions to reach the model from places you didn't expect.

This briefing is for CTOs and directors at fintech, ticketing, and other regulated or high-traffic platforms. It covers the four risks worth understanding, why they're harder to audit than classic access control, ten questions to put to your team or vendor before launch, and what governance looks like once the server is live.

Four MCP Security Risks Worth Knowing

The most useful reference points here come from OWASP, which now maintains three relevant lists: the Top 10 for LLM Applications, the Top 10 for Agentic Applications released in December 2025, and an MCP Top 10 that's currently in beta. In May 2026, the NSA's AI Security Center added its own guidance on MCP, concluding that "established cyber defense strategies unfortunately do not adequately address these new risks" (NSA, 2026).

Across those sources, four risks come up again and again for teams putting MCP into production.

1. Over-Scoped Permissions

Most MCP servers start life as a proof of concept, and proofs of concept tend to run on whatever credential was quickest to set up. That shortcut has a way of surviving into production. OWASP calls the result excessive agency and breaks it into three parts, all of which apply to MCP (OWASP, 2025):

  • Excessive functionality: the server exposes tools that the use case doesn't need, such as a "delete record" tool on an assistant that only answers questions.
  • Excessive permissions: the tools run under a credential that can reach more data than any single user should see.
  • Excessive autonomy: high-impact actions happen without a person approving them in real time.

The MCP specification is specific about one version of this problem. An MCP server must not accept tokens that weren't issued for it and pass them straight through to downstream APIs, because doing so bypasses the controls those APIs rely on and leaves logs that can't tell one caller from another (MCP specification). OWASP's MCP list adds scope creep, where permissions loosely granted at launch keep widening as new tools are added.

Even well-run providers get caught by isolation gaps. In June 2025, Asana took its MCP server offline for almost two weeks after finding a bug that could have exposed information from one customer's workspace to other organizations' MCP users (The Register, 2025).

2. Tool Poisoning

Every MCP tool comes with a description written in plain language, and the AI model reads that description to decide when and how to use the tool. The model treats it as guidance, not as data to be inspected, which makes the description a place to hide instructions.

A tool described as a currency converter can carry a line, invisible in most user interfaces, telling the model to include the contents of a configuration file in its next request. Invariant Labs demonstrated this class of attack in April 2025, and OWASP's MCP list now ranks it third.

How often does it work? MCPTox, an academic benchmark presented at AAAI, tested poisoned tool descriptions across 45 live MCP servers and 20 well-known AI agents.

The most susceptible model followed the hidden instructions in 72.8% of cases, and the researchers found that more capable models were often more susceptible, because the attack exploits their stronger instruction-following (Wang et al., 2025). A separate study of 1,899 open-source MCP servers found tool poisoning patterns in 5.5% of them (Hasan et al., 2025).

A quieter variant is the "rug pull," where a tool is clean when your team reviews it and changes later. If tool descriptions aren't treated like code, with changes reviewed before they reach production, an approval granted in March says little about what the tool tells the model in October.

3. MCP Prompt Injection Through Returned Data

Tool descriptions aren't the only text the model reads. Everything a tool returns, from a support ticket to a product review or a transaction memo, lands in the same context window, and the model has no reliable way to separate "information to summarize" from "instructions to follow." OWASP calls this indirect prompt injection and is candid about the state of defenses: "it is unclear if there are fool-proof methods of prevention".

The clearest public example came in May 2025, when Invariant Labs showed that a malicious issue posted to a public GitHub repository could steer an agent using the GitHub MCP server into copying data from the user's private repositories into a public pull request. The researchers stressed that this wasn't a bug in GitHub's server code but an architectural problem: the agent had access to both public and private repositories in the same session, and untrusted text reached it through a legitimate tool (Invariant Labs, 2025).

For a fintech or ticketing platform, the equivalent openings are anywhere outsiders can write text your tools will later return:

  • free-text fields on payments, invoices, and refund requests
  • customer support messages and chat transcripts
  • event descriptions, venue notes, and reviews submitted by organizers or fans
  • documents and emails pulled in for summarization

Since no filter catches every injected instruction, the practical goal is to limit what a successful one can do. A model that reads untrusted text shouldn't, in the same session, hold a credential that can send data somewhere new or take an irreversible action without someone approving it.

4. Unvetted Third-Party MCP Servers

Not every MCP server in your environment will be one your team built. Developers add community servers to their AI coding tools, business teams connect off-the-shelf connectors, and vendors ship MCP servers alongside their products. Each one runs with whatever access it's given, and OWASP's MCP list names "shadow MCP servers," deployments that operate outside security governance, as a risk in their own right.

The first known malicious MCP server showed how this plays out. A package called postmark-mcp, posing as an official connector for an email provider, published fifteen clean versions before version 1.0.16 added a single line that blind-copied every email it sent to an outside address. Postmark had no involvement with the package and the teams that reviewed it early and let it update automatically would have missed the change entirely.

Why AI Agent Security Is Harder to Audit Than Classic Access Control

In a traditional application, a security reviewer can list every screen and every query, check who is allowed to reach each one, and sign off. The set of possible behaviors is fixed when the code ships, so a review done in March is still meaningful in May unless someone changes the code.

An MCP-connected agent doesn't work from a fixed set of screens. It composes each answer at runtime from the user's request, the descriptions of every tool it can see, and whatever data those tools return, and the same request can take a different path on a different day. A tool description edited by a third party, a new field in an upstream API response, or a cleverly worded support ticket can all change what the agent does without anyone on your team touching your code.

That shifts the audit question. Instead of asking "who can access this data?", teams also need to answer "what did the agent do with this data in this session, and why?" Answering the second question needs logging and controls that most platforms haven't built yet, which is why agentic AI security tends to be decided by choices made before launch rather than by testing after it.

10 Questions to Ask Before You Ship an MCP Server

These questions work for an internal team and for a vendor building or supplying an MCP server for you. A confident, specific answer to each is a good sign. "We'll handle that later" deserves a follow-up conversation.

Graphic placement: "10 Questions Before You Ship" checklist card

  1. Whose identity does each tool call run under, and could it ever fall back to a shared service account? A good answer explains how the user's identity travels from the AI client to every downstream system, and confirms the server only accepts tokens issued for it.
  2. Which tools can change data or move money, and what stops the agent from calling them without a person approving? Look for a short list of write actions, each with a defined approval step or limit enforced in the server, not in the prompt.
  3. Who wrote each tool description, and how would we know if one changed? Descriptions should live in version control and go through review like code, with alerts when a third-party tool's description differs from the approved version.
  4. What untrusted text can reach the model through our tools? Your team should be able to name the free-text fields, documents, and messages that outsiders can write to, and explain what limits apply when the agent reads them.
  5. Can a single response combine data that should never meet? Ask how the server keeps one customer's, tenant's, or department's data out of another's session, including when two tools are used together.
  6. Which third-party MCP servers are in use, including ones developers installed themselves? There should be an inventory with a named approver for each server, not a list assembled after an incident.
  7. Are third-party servers pinned to a reviewed version? Automatic updates on an MCP server are automatic changes to what your agent can do, so updates should be reviewed before they're applied.
  8. Where do credentials live, how long do they last, and can we revoke one agent without breaking everything else? Short-lived, narrowly scoped credentials per agent or per server make revocation a routine step rather than an outage.
  9. If something goes wrong, can we reconstruct exactly what the agent saw, called, and returned? That means logs tying each tool call to a user, an AI client, the arguments sent, and the data returned, kept somewhere the agent itself can't edit.
  10. Has anyone tested it with poisoned tool descriptions and injected data? A standard penetration test checks the server as an API. It's worth also checking how the agent behaves when a tool description or a returned record contains instructions.

MCP Governance After Launch

Launch is where MCP governance starts. Four areas tend to decide whether a server stays trustworthy as tools, clients, and upstream systems change around it:

Area

What goes wrong without it

What good looks like

Logging

An incident review can show that data left, but not which request, tool, or user caused it

Every tool call is logged with user, client, tool, arguments, and a record of what was returned, in storage the agent can't modify

Versioning

A tool's behavior or description changes and nobody notices until something misfires

Tool definitions and third-party servers are pinned, and any change goes through the same review as a code release

Ownership

A server built for a pilot keeps running after its builder moves on, with nobody watching its access

Each server and each high-impact tool has a named owner who reviews its permissions on a set schedule

Upstream changes

An API adds a field and the MCP server passes it straight to the model

The server returns only fields on an approved list, and tests flag any change in what the upstream API sends

The last row catches teams most often. An MCP server usually sits on top of APIs owned by other teams or other companies, and those APIs change on their own schedules. If a payments API starts returning an internal risk note or a full account number in a field the MCP server forwards without filtering, the model will read it and may repeat it. Returning only approved fields, and testing for changes in what upstream systems send, turns a silent exposure into a failed test.

Ownership also matters beyond engineering. For payment environments, the PCI Security Standards Council's AI principles say an AI system can't accept responsibility for its actions, so a named person has to (PCI SSC, 2025). Assigning that owner at launch is easier than finding one after an audit question arrives.

Read more: Case Study Snapshot: Putting AI Agents to Work on Live Infrastructure, With Guardrails First While supporting a client's multi-service platform, a Softjourn senior engineer built an agentic development workflow connected to live project systems through MCP servers. Before the workflow started, the team set strict security policies for every agent with codebase access and controls on what client data the agents could reach. It also learned two lessons along the way: agents need cross-project dependencies documented explicitly, and every package an AI suggests gets verified before it's installed. [Read the case study](LINK: AI-Augmented Development Cut a Week of Documentation Down to Minutes)

Getting MCP Security Right Before Go-Live

MCP security comes down to a handful of decisions: which identity each tool call runs under, which tools can act without a person, which descriptions and third-party servers you trust, and what you'll be able to reconstruct after something goes wrong. None of them require exotic technology, but they're all cheaper to make before launch than to retrofit afterward.

If you're planning to connect agents to customer, payment, or ticketing data, our MCP server development services start from those decisions. Contact Softjourn to get started on an MCP server with permissions, approvals, and logging built in from the first tool.

Frequently Asked Questions

What Is MCP Tool Poisoning?

MCP tool poisoning is an attack in which instructions are hidden in a tool's description, the plain-language text an AI model reads to decide how to use the tool. Because the model treats that description as guidance, it may follow the hidden instructions, for example by sending data to a place the user never asked for. An academic benchmark found that the most susceptible model tested followed poisoned descriptions in 72.8% of cases. Reviewing tool descriptions like code and pinning third-party servers to approved versions are the main defenses.

What Is MCP Prompt Injection?

MCP prompt injection happens when text returned by a tool, such as a support ticket, review, or document, contains instructions that the AI model follows as if they came from the user. OWASP calls this indirect prompt injection and notes that there may be no fool-proof way to prevent it. The practical response is to limit what the agent can do in any session where it reads untrusted text, especially sending data externally or taking actions that can't be undone.

Is MCP Secure?

The protocol includes security requirements, such as rules on how servers handle authorization tokens, but an MCP deployment is only as secure as the permissions, tools, and third-party servers behind it. Most published incidents so far, including the GitHub and Asana cases, came from how servers were scoped and isolated rather than from a flaw in the protocol itself. Treating each MCP server like any other privileged connection to your systems, with an owner, a review process, and logging, covers most of the risk.

What Does MCP Governance Cover?

MCP governance covers what happens after an MCP server goes live: logging every tool call, versioning tool definitions and third-party servers, assigning an owner to each server and high-impact tool, and handling changes in the upstream APIs the server depends on. Without it, a server that was safe at launch can drift as tools are added, descriptions change, or upstream systems start returning new data.

What Our Clients Say

  • “Your team has provided us with outstanding service and outcomes. We couldn't be happier with your work or our progress. All of the members of your team have each shown themselves experts in their respective areas and have been a pleasure to work with.”

    Ben Melton

    Product Owner at CapStorm

    Read case study →
  • “The partnership, commitment, and skill of the Softjourn team enabled us to navigate this product transformation effectively.”
    Eric Rauch

    Eric Rauch

    Co-Founder of Pivot, Pivot

    Read case study →
  • “The Softjourn team was very quick to response to issues as well. I'm happy with the result.”

    Mike Kenefsky

    Operations Director at PM Vitals, PM Vitals

  • “Softjourn's pragmatic approach spotted potential blockers early on, ensuring we stayed on track.”
    Sam Mogil

    Sam Mogil

    CEO & Co-Founder, SquadUP

    Read case study →
  • “Softjourn's pragmatic approach spotted potential blockers early on, ensuring we stayed on track.”
    Richard Bates

    Richard Bates

    Director of Product at Spektrix, Spektrix

    Read case study →
  • “Wonderful work on our platform – everything looks great, and you did such a great job!”

    Myers-Briggs

    Team Leaders, Myers-Briggs

    Read case study →

Partnership & Recognition

Want to Know More?

Fill out your contact information so we can call you