Fundamentals

API, MCP, and Skills: The Three Layers Behind How an AI Agent Actually Calls a Tool

API, MCP, and Skills: The Three Layers Behind How an AI Agent Actually Calls a Tool

An earlier article in this series broke an AI agent down into perception, planning, memory, and tool use — but “tool use” was left deliberately abstract. This one gets concrete: when an agent “calls a tool,” what is actually happening, mechanically? The honest answer runs through three distinct layers — API, MCP, and Skills — that get casually lumped together in conversation but are genuinely different things solving different problems.

Layer 1: API — the general-purpose idea, decades older than AI

An API (Application Programming Interface) is just a defined contract: here’s how you’re allowed to ask my program to do something, and here’s the shape of the answer you’ll get back. A weather service’s API might say: send a GET request to a specific URL with a city name, and you’ll get back a JSON object with a temperature field. The program exposing the API doesn’t need to reveal how it actually computed that temperature — you just need to speak the contract correctly.

This idea predates AI by decades and underlies essentially all of modern software — every app on your phone that talks to a server is making API calls constantly. The relevant point for this article: when people say an AI model “used a tool,” at the lowest mechanical level, that almost always means the model’s software wrapper made an API call on the model’s behalf, using arguments the model generated, and then fed the API’s response back into the model’s context so it could use the result.

The problem this creates at scale is straightforward: every API is designed differently — different authentication schemes, different request formats, different ways of describing what’s even possible to do with it. If you want an AI application to be able to use dozens or hundreds of different tools, someone has to write custom integration code translating the model’s intentions into each tool’s specific API format, one at a time. This is sometimes called the M×N integration problem: M different AI applications, each needing custom code for N different tools, means roughly M×N separate integrations to build and maintain.

Layer 2: MCP — a shared protocol that turns M×N into M+N

Model Context Protocol (MCP) is a specification Anthropic released in November 2024, and it exists specifically to collapse that M×N problem. Instead of every AI application writing bespoke code for every tool, MCP defines one standard way for a “server” (something exposing a capability — a database, a file system, a search engine, a company’s internal ticketing system) to describe what it can do, and one standard way for a “client” (the AI application — Claude Desktop, Claude Code, or any other MCP-compatible app) to discover and call those capabilities.

Concretely, an MCP server exposes three kinds of things: tools (actions the model can invoke, like “create_ticket” or “run_query”), resources (data the model can read, like a file or a database record), and prompts (reusable prompt templates). Build one MCP server for, say, a company’s internal CRM, and every MCP-compatible AI application can connect to it without anyone writing app-specific glue code. Build one MCP client capability into an AI app, and it can talk to every MCP server that already exists, growing without the app’s developers writing anything more. That’s the actual mechanism behind the “M+N instead of M×N” claim — a comparison people often make to USB-C: before a shared standard, every device needed its own proprietary cable; after, one port works with everything that speaks the same protocol.

It’s worth being precise about MCP’s real limitation, not just its pitch: MCP standardizes the plumbing — discovery and calling conventions — not the judgment about when and how to use a tool well, and not the quality of what’s on the other end. A badly designed MCP server with a confusing tool description will still get misused by a model, the same way a badly documented API confuses a human developer. MCP removes an integration cost; it doesn’t remove the need for a tool to be genuinely well designed.

Layer 3: Skills — packaged know-how, not a live connection

Skills solve a different problem entirely, and the distinction matters. An MCP connection reaches an external, live system to fetch fresh data or take a real-world action — it’s inherently about talking to something outside the model. A skill is closer to a bundled unit of instructions, reference material, and sometimes small helper scripts that teaches an agent how to competently carry out a specific kind of task — say, filling out a particular company’s expense report format, or following a specific code-review checklist — loaded into the agent’s context only when that task is actually relevant, rather than crammed into every conversation by default.

The reason this separation exists is a real, practical constraint: a model’s context window — how much text it can actively “think about” at once — is finite, and stuffing it full of instructions for every conceivable task the model might never actually need in a given conversation wastes that budget and can dilute the model’s focus. Packaging specialized procedural knowledge as an on-demand skill, loaded only when relevant, keeps the default context lean while still letting the agent become genuinely competent at narrow, specific tasks when they come up. It’s less “a new way to talk to a database” and more “a reference manual and toolkit the agent picks up off the shelf exactly when the job calls for it.”

How the three layers actually fit together

Put concretely: an API is the raw contract any two programs use to talk to each other, and it’s been the foundation of software integration for decades. MCP is a standardized way for AI applications specifically to discover and call APIs — and other capabilities — across many different tools without bespoke per-tool integration work. A skill is a separate, complementary idea: a packaged bundle of task-specific knowledge and instructions that an agent loads on demand, which may or may not involve calling anything external at all.

None of these three layers is “the” way an AI agent does real work — they’re each solving a different part of the same underlying problem: letting a model that’s fundamentally good at generating text also reliably take real, useful actions in the world. Understanding which layer you’re actually talking about is what separates a useful conversation about AI agents from one where “it can use tools” is doing all the unexamined work.

Frequently asked questions

Is MCP the same thing as an API?

No — MCP is built on top of the same idea as an API, but standardizes it specifically for AI models. An API is any defined interface for one program to call another; there are countless incompatible API designs. MCP is one specific, shared protocol so an AI model doesn't need custom integration code for every different tool it connects to.

Do I need MCP if a tool already has an API?

The tool having an API doesn't automatically mean an AI model can use it — someone still has to write code translating between the model's tool-calling format and that specific API's request format. An MCP server is exactly that translation layer, built once per tool and then reusable by any MCP-compatible AI application, instead of every AI app writing its own bespoke integration.

What's the practical difference between an MCP tool and a Skill?

An MCP connection reaches out to a live external system — a database, a web service, your calendar — and gets fresh, current data or triggers a real action. A skill is closer to a pre-packaged bundle of instructions, reference material, and sometimes small scripts that teach the agent how to competently perform a specific kind of task; it doesn't necessarily talk to anything external at all.