
MCP vs API vs RAG vs A2A: What Each One Does and When You Need It
MCP, APIs, RAG, A2A, and function calling often show up in the same pitch as if they were interchangeable, but each one solves a different problem, and most real systems end up using several of them together.
In the same quarter, a CTO can hear three requests that sound related. Engineering wants to build an MCP server. A vendor's pitch says its platform already handles RAG and A2A. And the board wants customers to be able to reach the product from ChatGPT. Ask whether any of this means replacing the API the business already runs on, and the room usually goes quiet.
It doesn't, and the reason is the most useful thing to understand about MCP vs API, RAG, and A2A. An API lets software talk to software. MCP lets an AI assistant use those APIs as tools. RAG gives the assistant your documents to read, and A2A lets it hand a task to another company's agent.
These are four different jobs, and a real system often needs several of them at once. (If you want the background first, our guide explains what an MCP server actually does.)
This guide gives you a one-table summary, a section on each comparison that ends with when to use each, and a walkthrough of a single refund request that touches four of them.
It doesn't name a winner, because there isn't one. What you'll get is a way to tell which job each box on the next architecture slide is supposed to be doing.
The Short Version: Five Terms in One Table
What it is | What it's for | Use it when | |
|---|---|---|---|
API | A defined way for one piece of software to request data or actions from another | Connecting systems to each other, with or without AI | A program with a known sequence of steps needs to talk to your platform |
MCP | An open standard for exposing your systems to AI assistants as a set of tools | Letting AI work with live data and take actions in your systems | An AI assistant needs to answer varied questions or complete tasks using your platform |
RAG | Retrieval-augmented generation: finding relevant passages in a document collection and giving them to the model before it answers | Grounding AI answers in your own documents, policies, and knowledge | The answer lives in text you've already written, not in a live system |
A2A | The Agent2Agent protocol, an open standard for AI agents to delegate tasks to each other | Coordinating work between agents, often across organizations | Your agent needs to hand a task to another agent that plans and acts on its own |
Function calling | A model feature that lets the AI request a call to a function your application defines | Giving one application's AI a few actions to take | A single application, with tools you own, needs its AI to take simple actions |
Graphic placement: decision tree ("What do you need the AI to do?")
MCP vs API: An Adapter, Not a Replacement
The most common misunderstanding is that MCP competes with APIs. It doesn't. An MCP server almost always sits on top of APIs you already have, and when an AI assistant asks it to look up an order, the server turns around and calls your order API to do it. If you retired the API, the MCP server would have nothing to call.
The better question is who the consumer is. An API is written for a developer who reads the documentation once, writes code, and expects that code to behave identically every time it runs.
An MCP tool is written for an AI model that reads a short description, decides in the moment whether the tool fits the request in front of it, and chooses what to call next based on what came back. Those are two very different readers, and they need different things.
MCP Server vs API: Where the Differences Show Up
Three practical differences come up once teams start building:
- Who decides the sequence: in a traditional API setup, a developer hard-codes the order of calls. With MCP, the model picks the order at runtime. That's useful for open-ended requests, but it's the wrong fit for workflows that must run identically every time, like a settlement batch or a payroll run.
- How the work is cut: an API might have forty endpoints mapped to how your database is organized. A good set of MCP tools usually has far fewer, each mapped to a task a person would ask for, such as "check whether this order can be refunded" rather than separate calls for orders, payments, and policy rules.
- What gets returned: APIs often return full records for programs to process. AI models work best with short, readable results, so an MCP tool that returns 5,000 transaction rows when the question needed five is wasting time and money and making the answer worse.
There's a cost difference too. A direct API call is fast, cheap, and predictable. A request that goes through an AI model and MCP adds model processing time and usage costs, and the same question can take slightly different paths on different days. That's a fair trade when the flexibility is what you need, and a poor one when it isn't.
One more practical point: an MCP server inherits the quality of the APIs underneath it. If your order API can't filter by customer or doesn't enforce per-user permissions, the MCP server built on top will struggle with the same gaps. Teams that plan on building an MCP server on top of your existing APIs often find the first round of work is tidying up those APIs.
Softjourn has seen this firsthand on projects connecting ticketing platforms to partner systems.
When a ticketing client needed to connect to Tessitura, the vendor's API documentation had no usage guide, disorganized endpoints, and inconsistent URL patterns, and the project had stalled for nearly two months. Softjourn's team started with a discovery phase that mapped the APIs against the client's needs, then built an adaptor layer that handled ticket sync and purchase flows on the client's behalf (Tessitura case study).
An MCP server plays a similar adaptor role, with an AI assistant as the consumer instead of a partner platform, and the same discovery work pays off.
Read More: Successful Vivenu Integration Opens New Doors for Ticketing Platform A ticketing platform needed to connect to Vivenu to support a major university's athletic events, and found gaps in the newer version of Vivenu's APIs along the way.
Softjourn's discovery work covered storefront sales, distribution, and barcode ingestion, and deliberately held back non-critical features until the APIs matured, so the launch didn't depend on them. Read the case study
When to Use Each
- Use an API when the consumer is code with a known sequence of steps: your mobile app, a partner's booking system, a nightly reconciliation job, or any workflow that has to behave the same way every time.
- Use MCP when the consumer is an AI assistant handling requests that vary from one conversation to the next, such as a support agent's questions, a finance analyst's ad hoc lookups, or a customer asking an assistant to find and compare options.
- Use both in almost every real deployment. The API stays the contract with your systems of record, and MCP becomes the layer that lets AI assistants use it safely.
MCP vs RAG: Knowing Things vs. Doing Things
RAG and MCP both feed information to an AI model, which is why they get confused. The difference is where the information lives and whether anything happens as a result.
RAG works on documents. Before the model answers, a retrieval step searches a collection of text (policies, contracts, help articles, product documentation) and passes the most relevant passages to the model, so the answer is grounded in what your organization has actually written. MCP works on live systems. The model calls a tool, the tool queries a database or triggers an action, and the result reflects the state of the business right now.
A quick way to tell them apart is to look at the question being asked:
- "What's our refund policy for postponed events?" is a RAG question, because the answer is in a policy document.
- "Can this customer get a refund on order 48213, and can you process it?" is an MCP question, because it needs live order data and an action.
Freshness matters as well. A RAG index reflects documents as of the last time they were processed, which is fine for policies that change quarterly and risky for anything that changes by the minute, like seat availability or an account balance. The two also overlap more than vendor diagrams suggest, since a document search can itself be offered to the assistant as an MCP tool.
RAG also takes more than a model and a pile of documents.
When Softjourn's R&D team turned more than 25 years of case studies, team surveys, and project feature logs into a searchable knowledge base, it built three separate pipelines to clean and structure the material before anything reached the database.
The team's own summary was that "the AI is about 10% of the work," with the rest spent on data preparation, orchestration, error handling, and testing (AI-searchable knowledge base case study). Any RAG proposal that skips over the data preparation deserves a follow-up question.
Read More: How We Built a Non-Hallucinating AI Chatbot in One Week Softjourn's R&D team built a website chatbot on Google Cloud's Vertex AI that answers only from our own case studies and documentation, with a scripted fallback for anything outside them. It's a working example of grounding answers in your own content, and the same build connects to backend systems through APIs and MCP servers. Read the case study
When to Use Each
- Use RAG when the answer already exists in written form and the main risk is the AI making something up, such as policy questions, product documentation, internal knowledge bases, and "have we done this before?" questions.
- Use MCP when the answer depends on live data or the request ends in an action, such as order status, balances, availability, bookings, or refunds.
- Use both when a request needs the rule and the record, which describes most customer support conversations.
MCP vs A2A: Using a Tool vs. Working With a Colleague
MCP connects an agent to tools. A tool does exactly what it's asked, returns a result, and has no opinion about the request. The A2A protocol connects an agent to another agent, which is a different relationship.
The agent on the other side can plan its own steps, ask follow-up questions, take hours or days to finish, and decide how to handle the work using tools and data the first agent never sees.
Google launched A2A in April 2025 and handed it to the Linux Foundation, and by its first anniversary the protocol had reached version 1.0 with more than 150 supporting organizations and support in Microsoft, AWS, and Google Cloud's agent platforms (Linux Foundation, 2026).
In August 2026, A2A joined the Agentic AI Foundation, the same Linux Foundation body that has hosted MCP since December 2025, which puts both protocols under one roof (Axios, 2026). The Linux Foundation's own summary of the relationship is a useful one-liner: A2A covers how agents "communicate and coordinate with each other across organizational boundaries, while MCP defines how agents connect to internal tools and data sources."
A practical example makes the line clearer. A bank's customer assistant can use MCP tools to look up a disputed card transaction and freeze the card, because both are actions in the bank's own systems.
If resolving the dispute needs input from the merchant's side, and the merchant runs its own agent that reviews evidence and decides on a refund, that handoff is an A2A conversation. The bank's assistant can't call the merchant's systems directly and shouldn't be able to, but it can ask the merchant's agent to do the work and report back.
When to Use Each
- Use MCP when your AI needs to read from or act on systems you control, or systems a partner exposes as plain tools.
- Use A2A when the other party is an agent that makes its own decisions, especially one run by another organization, or when a task is long-running and the result comes back later.
- Start with MCP in most cases. Agent-to-agent handoffs across companies are real but still early, and most platforms get more value first from giving their own assistants reliable access to their own systems.
MCP vs Function Calling: A Model Feature vs. a Shared Standard
Function calling is a feature of the AI model itself. A developer describes a few functions to the model inside one application, and when the model decides one is needed, it returns a structured request that the application then runs.
It's simple and effective, and each application defines its own functions in the format its model provider expects.
MCP standardizes the same idea across applications. Instead of each app defining its own functions for each model, an MCP server describes its tools once, and any compatible AI client can connect and use them.
In practice, the two work together, since AI clients often use the model's function calling to invoke the MCP tools they've connected to.
The shared-standard part is what makes MCP useful inside engineering teams as well as customer-facing products.
On one client engagement, a Softjourn senior engineer connected AI coding agents to the team's live project systems through MCP servers instead of wiring custom functions into each tool, and documentation updates that used to take about a week dropped to minutes (AI-augmented development case study).
When to Use Each
- Use function calling for a single application with a handful of tools you own, one model provider, and no plans to share those tools elsewhere.
- Use MCP when the same capabilities should be available to several assistants, teams, or AI clients, or when you want customers to reach your platform from the assistant they already use.
How MCP, APIs, RAG, and A2A Work Together
In a real system, these pieces sit in layers rather than competing for the same slot. Here's how a single request might move through a ticketing platform's support assistant. The scenario is illustrative, but each step uses the building block it would in practice.
A fan writes to the platform's assistant: "My Saturday show was moved to a date I can't make. Can I get my money back?"
- RAG finds the rule: the assistant searches the platform's policy documents and the organizer's event terms, and finds that rescheduled shows qualify for a refund within 14 days of the change announcement, unless the organizer has set a different policy.
- MCP checks the record: the assistant calls MCP tools to look up the fan's order, confirm the event was rescheduled and when, and check whether the organizer has a custom refund rule on file. Each tool calls the platform's existing order and event APIs underneath.
- A2A handles the exception: the organizer requires approval for refunds on this event, and the organizer runs its own back-office agent. The platform's assistant sends the request to that agent over A2A, and the organizer's agent approves it based on its own rules.
- MCP takes the action: with approval in hand, the assistant calls the MCP refund tool, which calls the payments API to issue the refund and the ticketing API to release the seat back into inventory.
- Function calling ties it together: throughout, the model inside the assistant uses function calling to request each tool, and the application runs those requests.
Remove any one layer and the experience breaks in a specific way. Without RAG, the assistant guesses at the policy.
Without MCP, it can quote the policy but can't check the order or act on it. Without A2A, the exception becomes an email to the organizer and a ticket in someone's queue. And without the APIs underneath, none of the tools would have anything to call.
That bottom layer is usually the most mature part of a ticketing platform, because it's the same sync and purchase plumbing partners already use.
When Softjourn connected a ticket distribution client to Ticket Evolution, the work was organized around exactly those two flows, ticket sync and ticket purchase, each delivered as its own milestone (Ticket Evolution case study). MCP tools for an AI assistant can sit on the same flows rather than duplicating them.
Choosing the Right Mix
The useful question for any pitch or roadmap isn't "MCP vs API" or "RAG vs A2A" but "which job is each piece doing here?" If a diagram includes all of them, someone should be able to explain what each one handles and what would break without it. If they can't, that's the box to question.
For most platforms, the sensible order is APIs you trust, MCP tools on top of them for AI assistants, RAG wherever answers live in documents, and A2A once there's a real partner agent to work with.
Contact Softjourn to get started on an MCP layer that puts your existing APIs to work for the AI assistants your customers and teams already use.
Frequently Asked Questions
Does MCP Replace APIs?
No. An MCP server sits on top of your existing APIs and calls them to do its work, so the APIs remain the contract with your systems of record. MCP changes who the consumer is: instead of code written by a developer, it's an AI assistant deciding which tool to use in the moment. Most platforms keep their APIs for apps, partners, and batch jobs, and add MCP for AI assistants.
What Is the A2A Protocol?
The A2A (Agent2Agent) protocol is an open standard that lets AI agents delegate tasks to each other, including agents run by different organizations. Google launched it in April 2025, it reached version 1.0 in 2026, and it's now hosted by the Agentic AI Foundation alongside MCP. It's designed for cases where the other side is an agent that plans and acts on its own, not a tool that simply returns data.
Do We Still Need RAG If We Have MCP?
Usually, yes. RAG grounds answers in documents like policies, contracts, and help content, while MCP handles live data and actions. Many support and operations assistants need both, one for the rule and one for the record, and a document search can even be offered to the assistant as an MCP tool.
Do We Need A2A Today?
Most platforms don't yet. A2A becomes useful when your agents need to hand work to agents run by partners, suppliers, or customers, and those agents are only now appearing in production. Getting reliable MCP access to your own systems first gives you something worth connecting when partner agents arrive.


