LetterMCP

What Is an MCP Server and How It Works

MCP standardizes how AI applications connect to external tools and data sources.

Staff Writer · · 9 min read · Updated
Cover illustration for “What Is an MCP Server and How It Works”
Agentic AI Foundations · August 19, 2026 · 9 min read · 2,002 words

Every AI application that needs to reach an outside system, a database, a file store, a messaging tool, faces the same wall: without a shared protocol, that connection has to be built by hand, and the work multiplies in a way that breaks teams rather than helping them. This is the N×M problem. With N AI applications and M tools or data sources, every single pairing needs its own custom connector, its own authentication approach, its own data format, its own error handling.

The pain appears differently depending on whether someone sits in engineering, maintenance, or the end-user seat. Engineering teams redo integration work from scratch every time a new model or data source enters the picture. Maintenance load climbs because any update to a tool, or any model deprecation, can silently break a connector nobody is watching closely enough. Behavior drifts across integrations that were each built a little differently, so end users get inconsistent results from what looks like the same feature. For the person actually using the AI tool, the cost is simpler and more annoying: without connected systems, even a sophisticated model leaves someone manually copying information back and forth between a source and the AI interface, a copy-and-paste tango that undercuts the whole point of having an assistant.

The fix is arithmetic before it's anything else. A shared protocol turns N×M into N+M: each AI application implements the protocol once, each external system gets one server built for it once, and from there any compatible client can reach any compatible server with no custom integration code in between. That's the problem the Model Context Protocol was built to solve, and the rest of this piece is about how it does it.

What MCP is

Diagram: From N×M Custom Connectors to N+M with MCP. Visualizes: Visualize the arithmetic shift at the heart of MCP's value proposition.

MCP is an open standard, built on JSON-RPC 2.0, that gives AI applications one consistent way to discover, connect to, and call external tools, data sources, and services. It doesn't invent a new idea so much as formalize one that already existed in pieces. Function calling and tool use were already common ways for models to act on the world, and MCP takes those concepts and locks them into a single specification that any model vendor or toolmaker can build against. It doesn't replace APIs. It standardizes how agents reach them.

Anthropic open-sourced MCP in November 2024. Within months, OpenAI, Google DeepMind, and Microsoft had adopted it, and current SDK download numbers (around 97 million monthly downloads across the Python and TypeScript SDKs, per 2026 data) point to a protocol that has become the default way to wire AI into real systems. The spec itself keeps moving: a July 28, 2026 revision made the core protocol stateless, pushing optional capabilities like interactive apps and enterprise-managed authorization out into extensions that sit outside the core spec. That's the shape of MCP to hold onto going in: one wire format, one way to discover and call tools, built to stay small at the core and grow at the edges.

The three-role architecture: host, client, and server

MCP assigns every participant in a deployment exactly one of three roles: host, client, or server. Each role carries a distinct job and a distinct trust boundary, and mixing them up is where most early confusion about the protocol starts.

The host is the AI application a person actually interacts with, a desktop assistant, an IDE, a custom-built agent. It holds the LLM, decides what context gets fed into the model, and owns the overall session. Think of the host as the browser: it's the environment someone sits inside, not something most teams are building from scratch.

The client lives inside the host. It translates a model's requests into MCP-formatted calls and hands results back to the model. A client keeps a one-to-one relationship with a single server, but a host can run several clients at once, which lets one user request reach multiple servers in the same breath. You rarely write a client by hand, because the host framework usually supplies it.

The server is the external program doing the actual work: it wraps one or more backend systems, declares what tools and resources it offers, handles its own backend authentication, and returns structured results. A server has no idea which model is on the other end of the conversation. It receives a call, does the work, and returns an answer, and that indifference to what's asking lets the same server serve different models and different hosts without modification. Carrying the analogy forward: the server is the website. Teams building for MCP are mostly building servers, not browsers.

All of this travels as JSON-RPC 2.0, a lightweight, agreed-upon format for saying "call this function with these arguments" and getting back "here is the result," written in JSON. The SDK handles the serialization, so in practice you almost never look at a raw message. What matters more than the wire format is a structural fact that sets up everything that follows: the model never touches the backend directly. Every bit of access runs through the server.

The lifecycle from user prompt to structured result

Diagram: One User Prompt, Four Ordered Steps. Visualizes: Illustrate the four-step lifecycle every MCP agent action follows: Step 1 — Discovery (client sends request; server returns schema of available tools, resources, prompts); Step 2 — Tool…

Every action an agent takes follows the same four-step sequence, and tracing it closely shows exactly where control, filtering, and logging need to live.

Step one is discovery. The AI client sends a discovery request, and the server answers with a schema describing what tools, resources, and prompts it offers. The agent learns what's available at runtime, with nothing hardcoded in advance. Picture a user asking, "which open pull requests are blocking our next release?" The host surfaces whatever GitHub MCP capabilities are available, and the model recognizes it needs repository data to answer the question.

Step two is tool selection. The model, not the user, decides which tool to call and with what parameters. The model does, based on the tool's description, and that single fact has security implications addressed later.

Step three is invocation. The client packages a structured request and sends it to the server, carrying whatever inputs the tool requires.

Step four is execution and return. The server does all the system-specific work, calling an API, querying a database, reading a file, validates the request, and sends back a structured result that the model folds into its answer. This is the step that matters most for anyone thinking about control. Because the model never touches the backend directly and every call passes through the server, the server is the one place permissions, filtering, and logging can actually be enforced. Nothing happens anywhere else in the chain that a server-side check can't catch.

None of this is limited to one server per request. A single user question can route through several servers working together, a Slack server and a project management server both contributing to one composed answer, with the host stitching the results together seamlessly.

The three primitives: tools, resources, and prompts

An MCP server can expose exactly three kinds of capability, and what matters is who gets to initiate each one, since that single detail determines how much human oversight each type needs.

Tools are model-controlled. The model decides on its own, based on the tool's description, when to query a database, create a ticket, send a message, or trigger a deployment. Tools carry side effects, they're the write side of MCP, and that's where most of the protocol's power and most of its risk both live. Because the model is the one invoking tools without a human in the loop at that moment, a tool's description functions as its real interface. A description that's vague or too broad invites calls nobody intended.

Resources are application-controlled. The host decides what files, database records, or documents get pulled into context. Most servers get by on tools alone; resources start to matter once an agent needs to read a whole body of data rather than just perform an action.

Prompts are user-controlled. A person invokes them explicitly, often as a slash command, and they work as reusable instruction templates that standardize how the AI approaches some recurring task or domain.

| Primitive | Who initiates it | What it does | Example | |---|---|---|---| | Tools | Model | Actions with effects | Create a ticket, run a query | | Resources | Application | Data to read | Files, records, documents | | Prompts | User | Reusable templates | Slash commands, workflow starters |

How servers are transported: local stdio versus networked HTTP

A server can run one of two ways, and the choice sets the entire security and operational posture of the deployment, including what kinds of authentication are even possible.

With stdio, the host launches the server as a subprocess on the same machine, and messages pass over standard input and output. The security boundary here is the user's own machine. The server inherits whatever access the host application already has. This is the default setup for developer tooling and local assistants, and it's the right starting point for anything personal or single-user.

With Streamable HTTP, the server runs as its own independent process, reachable over a private network or the public internet. It can serve many clients at once and slots into the infrastructure teams already run: load balancers, proxies, CDNs. Running this way means taking on everything a production API requires: authentication, authorization, rate limiting, TLS, observability. As of the July 28, 2026 revision, the core protocol is stateless, and the older SSE-based stateful transport has been superseded by Streamable HTTP.

The decision rule is simple: local and personal work stays on stdio, anything shared or hosted moves to HTTP and takes on the full production API checklist that comes with it. Networked HTTP is where enterprise deployments live, and that's exactly where authentication and governance stop being optional.

How MCP handles authentication and authorization

Authentication in MCP depends entirely on which transport a server uses, and the space between what the protocol specifies and what an organization actually needs in production is where most enterprise governance work gets done.

Local stdio servers carry no protocol-level auth mechanism. Authentication rides on the host application's own permissions, so the server inherits the user's machine access, and security lives at the OS boundary rather than anywhere in the protocol.

HTTP-based servers work differently, built around OAuth 2.1. An MCP server acts as an OAuth resource server, while a separate authorization server handles login and issues tokens. A request that shows up without a valid token gets a 401 challenge back. From there, the client discovers the authorization server through Protected Resource Metadata, registers itself, and completes an authorization code flow with PKCE. The token that comes out the other end is bound to that specific server and starts out with the smallest set of scopes it can get away with. No MCP credential should ever be issued without a defined TTL: the recommended pattern is OAuth 2.1 with short-lived tokens for SaaS integrations, vault-issued dynamic credentials for internal services, and scoped, vaulted API keys reserved only as a last resort.

Enterprise-Managed Authorization, EMA, reached stable status as an extension in June 2026, and it formalizes enterprise identity as the authority that provisions MCP server access. Under EMA, a user authenticates once through the identity provider the organization already runs, and that IdP automatically provisions whichever MCP servers the user is authorized to reach, no separate OAuth consent screen for each one. Access decisions sit in the IdP admin console, so you get a single auditable trail rather than a scattered set of per-server logs. Okta was the first identity provider supported, and Anthropic, Microsoft, and Okta were among the earliest to adopt EMA.

Even with EMA in place, one gap remains outside the protocol's reach. OAuth 2.1 and EMA both govern who gets to connect. Neither governs what an agent is allowed to do once it's connected. Scope management, per-call authorization, and audit logging of what actually happened after the connection was made are left to the organization to build, not something the protocol hands over for free.

Sources

  1. modelcontextprotocol.io
  2. k2view.com
  3. mindstudio.ai

More in Agentic AI Foundations