Almost every post on MCP vs API hedges, so here is the answer up front: an MCP server is not an alternative to your API, it is a client that calls it, and if you are not wiring a language model to a tool you do not need one. We build on both sides of that line. biz collect is an OpenAPI 3.1 REST API for local business data, and we wrapped that same API as an MCP server to find out what the protocol actually changes. What follows is the MCP server explained against the API it wraps: one operation written both ways, then what REST still does better, what MCP genuinely buys, and what the whole thing costs, including the cost nobody else on this page mentions.
What is the difference between an API and an MCP server?
An API is a contract between two pieces of software: your code knows the endpoint, the parameters and the response shape ahead of time, and calls it deterministically. An MCP server is a contract between a language model and that API: it publishes a list of tools, each with a name, a description and a JSON Schema, so a model that has never read your documentation can discover what is available and decide, at runtime, which call to make and with what arguments.
The Model Context Protocol is an open standard, introduced by Anthropic in late 2024, that formalizes the second contract. Its current specification revision is 2026-07-28. It defines three server primitives, and the spec's own control table is the fastest way to see the difference between MCP and API integration as you already do it:
| Primitive | Control | What it is |
|---|---|---|
| Prompts | User-controlled | Templates a person invokes deliberately, such as a slash command |
| Resources | Application-controlled | Context the client attaches, such as file contents or git history |
| Tools | Model-controlled | Functions the model itself decides to call, such as an API POST request |
Only the third row touches anything in your stack, and it does not compete with your API. It competes with the glue code you would otherwise hand-write to expose that API to a model. For the definition on its own, the glossary entry for MCP server covers it in a paragraph.
Is an MCP the same as an API?
No. An API defines how two systems exchange data; MCP defines how a model discovers and invokes those exchanges. The relationship is layered rather than competitive, in the way a REST framework sits on top of HTTP without replacing it.
Model Context Protocol vs REST API is a confusing framing partly because, on the wire, MCP looks like an API. The protocol uses JSON-RPC 2.0 messages, and its Streamable HTTP transport requires the server to expose "a single HTTP endpoint path" accepting POST. Every tools/call is literally an HTTP POST to something like https://example.com/mcp, carrying JSON, answered with JSON or a Server-Sent Events stream. It is an API, with a fixed method vocabulary.
Are MCP servers just APIs?
Structurally, mostly yes: an MCP server is an API with a standardized method set (tools/list, tools/call, resources/list, resources/read, prompts/list) and a standardized description format. What makes it more than a wrapper is that the description format is machine-consumable at runtime, so a client needs no recompile or redeploy when your tools change.
That is a real difference and a narrow one. In practice an MCP server for a REST API is a proxy: it holds the credential, maps a tool name onto an endpoint, forwards the call. If your API already had a client library, the MCP server is a second client library whose consumer happens to be a model. Anyone calling it a fundamentally new architecture is describing a diagram, not an implementation.
Will an MCP server replace an API?
No, and it structurally cannot, because an MCP server is a consumer of APIs rather than a substitute for one. Remove the API behind an MCP server and there is nothing left to call. Remove the MCP server and the API still serves every non-model caller you have, which on most systems is all of them.
The hedging on this question conflates two claims. The weak claim, which is true, is that MCP may replace the bespoke integration code teams write to connect one model to one tool: the hand-rolled function schema, the argument parser, the auth plumbing. The strong claim, that MCP replaces REST or GraphQL as an interface style, is false and nobody shipping either believes it. Your dashboard, mobile client, cron jobs, webhook receivers and billing reconciliation script all keep calling the API directly, because none of them need a model to pick their arguments.
The protocol also keeps changing under the people who adopted it. Revision 2026-07-28 removed protocol-level sessions and the standalone GET stream from Streamable HTTP, moved protocol version and client capabilities into per-request _meta fields, and replaced the initialize handshake with per-request version negotiation. The spec now calls the handshake-based revisions (2025-11-25 and earlier) "legacy" and publishes a compatibility matrix for talking to them, while the original HTTP+SSE transport from 2024-11-05 is formally deprecated and "eligible for removal in a future revision". A REST endpoint you shipped in 2024 still works. An MCP server you shipped in 2024 needs a rewrite.
The same operation, written both ways
One operation, expressed twice: find dentists within 10 km of Zurich and enrich each with contact emails scraped from its own website. Nothing about the work changes between the two versions. Only the caller does.
The REST call
This is the real biz collect surface. wait: true collapses the async create-then-poll cycle into one blocking request that returns completed results inline.
# Your API base URL and key are both shown in the docs at /docs
BIZ_API="https://kindly-lyrebird-376.convex.site/api/v1"
curl -X POST "$BIZ_API/search" \
-H "Authorization: Bearer $BIZ_COLLECT_API_KEY" \
-H "Content-Type: application/json; charset=utf-8" \
-H "Idempotency-Key: zurich-dentists-2026-08-28-run1" \
-d '{
"location": "Zurich",
"keywords": ["dentist"],
"radius_km": 10,
"result_pages": 1,
"scrape_emails": true,
"wait": true
}'
The response is the completed job envelope, abridged here to the fields this post uses:
{
"job_id": "j_a91k7d2m",
"status": "completed",
"results_count": 34,
"businesses": [
{
"id": "b_7f31c0",
"name": "Zahnarztpraxis Beispiel",
"address": "Beispielstrasse 12, 8001 Zurich, Switzerland",
"phone": "+41 44 123 45 67",
"website": "https://example.ch",
"emails": ["praxis@example.ch"],
"social_links": [
{ "platform": "instagram", "url": "https://instagram.com/example" }
],
"rating": 4.7,
"user_rating_count": 128
}
]
}
The MCP tool call
For MCP the same operation first has to be described. The server advertises this through tools/list:
{
"name": "search_local_businesses",
"title": "Search local businesses",
"description": "Find local businesses in a geographic area and optionally enrich each one with contact emails, social profiles and outgoing links scraped from the business's own website. Call this when the user asks for businesses of a given type in a given city, region or radius.",
"inputSchema": {
"type": "object",
"properties": {
"location": {
"type": "string",
"description": "City, region, or any recognizable place name."
},
"keywords": {
"type": "array",
"items": { "type": "string" },
"description": "Business types or terms to search for."
},
"radius_km": { "type": "number", "default": 10 },
"result_pages": { "type": "integer", "default": 1 },
"scrape_emails": { "type": "boolean", "default": true }
},
"required": ["location", "keywords"]
}
}
The model reads that description, decides this is the right tool, fills in the arguments itself, and the client issues a tools/call. On the Streamable HTTP transport under revision 2026-07-28 that is one HTTP POST carrying JSON-RPC, with the method and tool name mirrored into headers so proxies can route without parsing the body:
POST /mcp HTTP/1.1
Content-Type: application/json
MCP-Protocol-Version: 2026-07-28
Mcp-Method: tools/call
Mcp-Name: search_local_businesses
{
"jsonrpc": "2.0",
"id": 1,
"method": "tools/call",
"params": {
"name": "search_local_businesses",
"arguments": {
"location": "Zurich",
"keywords": ["dentist"],
"radius_km": 10,
"scrape_emails": true
},
"_meta": {
"io.modelcontextprotocol/protocolVersion": "2026-07-28",
"io.modelcontextprotocol/clientInfo": { "name": "ExampleClient", "version": "1.0.0" },
"io.modelcontextprotocol/clientCapabilities": {}
}
}
}
The result comes back as content blocks, with the structured payload alongside a text rendering of it:
{
"jsonrpc": "2.0",
"id": 1,
"result": {
"resultType": "complete",
"content": [{ "type": "text", "text": "{\"results_count\":34,\"businesses\":[]}" }],
"structuredContent": { "results_count": 34, "businesses": [] },
"isError": false
}
}
What changed, and what did not
The handler behind that tool is three lines of intent: hold the API key, hard-code wait: true so the model never manages a poll loop, POST the arguments to /api/v1/search, return the businesses array. The build is genuinely small, and the runnable recipe lives in the MCP server for business data guide rather than here, because this page is about the decision and that one is about the code.
What did not change is where the marketing blurs:
- The upstream HTTP request is byte-for-byte identical. Same endpoint, same parameters, same JSON back.
- The credential is the same and still lives server-side. MCP added a place to put it, not a way to avoid needing it.
- The latency is the same, plus one proxy hop.
- The cost of the data operation is identical. Credits are charged by the API, not by the protocol talking to it.
What did change: a model that has never read the biz collect docs can now find this tool, understand it, and call it correctly. That is the entire value proposition. It is real, and it is narrow enough that most readers of this page do not need it.
What the REST API still does better
This is the section the vendor posts skip. Five things are meaningfully worse through a tool call, and two of them cannot be fixed by writing a better MCP server, because they are properties of the protocol.
| Concern | Direct REST | Through an MCP tool | Fixable in your server? |
|---|---|---|---|
| Result pagination | Cursor or page params you control | Not defined for tools/call | No |
| Idempotency | Idempotency-Key header, safe retries | No protocol-level equivalent | Partly |
| Completion push | Signed webhooks on job.completed | Poll, or the Tasks extension | Partly |
| Error handling | HTTP status plus a typed error code | isError: true plus prose | Partly |
| Cost control | Rate limits and daily spend caps you enforce | Model decides how often to call | No |
Pagination. MCP supports cursor-based pagination on exactly four operations: resources/list, resources/templates/list, prompts/list and tools/list. It pages the catalogue, not the results. A tools/call returns one payload, and if that payload is a thousand business records, all thousand land in the model's context window at once. Over REST you would set result_pages, read a page, decide, read the next. Through a tool, your server has to invent a paging convention and teach the model to use it through the tool description, which is a suggestion rather than a contract.
Idempotency. biz collect accepts an Idempotency-Key header on POST /api/v1/search: send the same key twice and you get the existing job back instead of a second charged run. MCP has no equivalent concept, so a model that retries because the first response looked truncated will happily run and pay for the search twice. Synthesizing a key server-side from the arguments works, and you should do it, but that is papering over a gap rather than using a feature.
Webhooks. For long-running work the REST path is POST /api/v1/webhooks, subscribe to job.completed and job.failed, get a signed callback. MCP has no push channel to your infrastructure at all; it pushes to the client. Revision 2026-07-28 moved long-lived notifications behind a subscriptions/listen request, and asynchronous work has its own opt-in extension, Tasks (io.modelcontextprotocol/tasks), where the server returns a taskId and the client polls tasks/get until the status reaches completed, failed or cancelled. That extension fits a job-based API like ours well, and it is opt-in on both sides, so support varies by client. A webhook works everywhere today.
Error handling. MCP separates protocol errors (JSON-RPC errors, for unknown tools and malformed requests) from tool execution errors (returned in the result with isError: true so the model can self-correct). The second category is deliberately model-facing: the payload is text the model reads and reacts to. Our REST API returns a 402 with an insufficient_credits code, or a 429 with a Retry-After header and X-RateLimit-* headers on every response, and your code branches on that. A model reading "insufficient credits" in a text block might retry, might apologise, might try a smaller search. It is not deterministic and you cannot make it so.
Cost control. This one bites hardest and has no fix inside the server. With direct REST calls your code decides how many searches to run; with a tool, the model decides, and it can decide to run six to be thorough. Our per-plan daily spend caps exist precisely for this: Free is capped at 100 credits per day, Pro at 2,000, Business at 10,000, with API rate limits of 30, 120 and 300 requests per minute respectively. Expose a paid data tool to an autonomous agent without a spend cap behind it and the cap is the only thing between you and a surprising invoice.
Why use an MCP server instead of an API?
Three reasons, and only three: discovery, argument selection and no glue code. A model connected to an MCP server calls tools/list, reads the names, descriptions and JSON Schemas, and can then use tools that did not exist when the model was trained or when your client was built.
Discovery sounds abstract until you have maintained the alternative. Without MCP, connecting a model to your API means hand-writing a tool schema in whatever shape your runtime wants, keeping it in sync as parameters change, and repeating that per runtime. With MCP you write the description once, the server advertises it, any compliant client consumes it, and when you add a tool the server sends notifications/tools/list_changed and clients re-fetch.
Argument selection is the second. Nothing in the code above decided that keywords should be ["dentist"] and radius_km should be 10. A user said something in natural language and the model mapped it onto the schema. If your product's job is turning vague human intent into a well-formed API call, that mapping is the product. The third reason is simply that you write less code: a tool definition and a handler against an existing spec. None of the three apply if your code already knows what to call.
How much does an MCP server cost?
Three separate bills, and most write-ups mention only the first. In rough order of what will actually hurt: the tokens are usually the biggest, the upstream API calls second, the server itself close to free.
The server itself
Near zero, and often exactly zero. On the stdio transport the server is a subprocess the client launches on the machine already running the agent, so it costs nothing beyond CPU you already own.
A remote server needs one HTTPS endpoint accepting POST, which any small serverless runtime handles. Cloudflare's Workers pricing page, read on 28 August 2026, lists a free plan with 100,000 requests per day and 10 milliseconds of CPU time per invocation, and a paid plan at $5 per month including 10 million requests and 30 million CPU-milliseconds, with overage at $0.30 per additional million requests and $0.02 per additional million CPU-milliseconds. Because an MCP proxy spends its time waiting on an upstream response rather than computing, CPU metering is rarely the binding constraint. Budget $0 to $5 per month and stop thinking about it.
The API calls behind it
Unchanged by MCP, which is the point. The tool call and the direct call hit the same endpoint and are billed identically.
For biz collect a search costs 20 credits per result page. Top-up credits are $0.01 each with a 500-credit ($5) minimum, so a single-page search is about $0.20 at top-up rates. On a subscription the effective rate is lower: Pro is $19 per month billed yearly ($23 monthly) for 6,000 credits, Business is $76 per month billed yearly ($92 monthly) for 30,000. Every plan starts with 200 signup credits and no credit card, plus 20 credits per daily login. Whichever number applies, it is the same number whether a person, a cron job or a model triggered the call. MCP is not a pricing dimension.
The tokens the tool definitions occupy
This is the cost nobody on this search result answers, and the one that surprises people. Tool definitions are input tokens, and you pay for them on every request in the conversation, not once when the tool is registered.
Anthropic's pricing documentation, read on 28 August 2026, is explicit about what gets billed: "the tools parameter in API requests (tool names, descriptions, and schemas)", plus tool_use and tool_result content blocks, plus a fixed tool-use system prompt the API injects whenever any tool is present. That fixed overhead alone is 286 tokens on Claude Opus 5 with tool_choice: auto (406 with any or tool), 354 on Claude Sonnet 5 and 496 on Claude Haiku 4.5, before a single character of your own schema.
Your schemas land on top. The search_local_businesses definition printed earlier is 937 characters of JSON, and Anthropic's own rule of thumb is roughly 4 characters per token, so call it 235 tokens. A four-tool server covering search, job status, job listing and account balance lands somewhere near 900; add the fixed overhead and you are carrying about 1,200 input tokens of pure tool description in every request of the conversation. Measure yours with the token counting endpoint rather than trusting that arithmetic.
Now the sum that matters. A 30-turn agent session re-sends that block 30 times: roughly 36,000 input tokens. At Claude Opus 5's $5 per million input tokens (Anthropic's pricing page, read 28 August 2026) that is about $0.18 of pure schema per session, before a word of the actual conversation and before the tool results land in context. On Claude Sonnet 5 at $2 per million it is about $0.07. Small per session. Not small at ten thousand sessions a day.
Three consequences worth designing around:
- Multiply by the number of servers connected, not the number you use. A client connected to eight MCP servers loads all eight catalogues into the same context window. The one you needed and the seven you did not are billed the same way.
- Prompt caching is the real mitigation, and the spec knows it. Cache reads bill at 0.1x the base input rate and a 5-minute cache write at 1.25x, so a cached tool block costs a tenth of an uncached one after the first request. The MCP specification explicitly asks servers to "return tools in a deterministic order", noting this "improves LLM prompt cache hit rates when tools are included in model context". Non-deterministic ordering silently costs money.
- Changing the tool set mid-conversation invalidates that cache. Tools render at the very front of the prompt prefix, so adding or removing one invalidates everything after it. A server that emits
notifications/tools/list_changedoften is an expensive server.
The blunt version: a narrow, stable, deterministically ordered tool surface is a cost decision, not a style preference. Ship one good tool rather than five overlapping ones.
286 tokens
Fixed tool-use system prompt overhead on Claude Opus 5 with tool_choice auto, added to every request before any of your own tool schemas.
Anthropic pricing documentation, read 28 August 2026
Does an MCP server use an API?
Almost always, yes. An MCP server that does anything useful with remote data is calling an API behind the tool, and its job is to hold the credential, translate the model's arguments into that API's parameter shape, and translate the response back into content blocks.
A minority do not: a filesystem server calls the operating system, a SQLite server talks to a database driver, a calculator server just computes. The pattern holds anyway. MCP is a description and transport layer over something else that does the work, and when that something else is remote, it is an API.
Can you use MCP without an API?
Yes, for local capabilities. A stdio MCP server reading files, running a local binary or querying a local database needs no API at all, which is why the earliest useful servers were filesystem and git servers.
You cannot use MCP without an API for remote data, because there is nothing else for the handler to call. If a vendor's MCP server returns live business records, prices or search results, there is an API on the other side whether or not it is public. When it is not public, you have taken a dependency you cannot inspect, retry against, or fall back to.
Can I convert any API to an MCP server?
Any well-specified API, yes, and OpenAPI to MCP generators exist that read an OpenAPI document and emit tool definitions from it. The request side converts mechanically: an operation already has a path, a method, typed parameters and a response schema, which is nearly everything a tools/list entry needs.
Three things do not convert automatically, and they are the difference between a generated server and a good one.
Descriptions. OpenAPI descriptions are written for humans reading docs. Tool descriptions are read by a model deciding whether to call something, so they must state when to call it, not only what it does. Generated descriptions are usually the weakest part of an auto-converted server, and description quality is the biggest single lever on whether the model reaches for the right tool.
Surface area. A 60-endpoint API converts to 60 tools, and 60 tools is a context-window problem rather than a feature. Pick the two or three operations an agent genuinely needs. For our spec that means search_local_businesses and a job-status tool, not one tool per path.
Parameters the model should never see. wait is the clean example. Exposed as an argument, the model can set it to false, get back a bare job_id, and then need a second tool plus a loop to find the results. Hard-coded to true inside the handler, one tool call is the whole interaction. Same reasoning for spend controls: max_credits belongs in your handler, not in the model's schema.
biz collect converts comfortably because it publishes an OpenAPI 3.1 spec at /openapi.json, machine-readable docs at /llms.txt and /llms-full.txt, a stable JSON response shape and that synchronous mode. The step-by-step build is the companion to this page. If you are earlier in the process and still deciding whether an agent belongs in the loop at all, the AI agent lead generation pipeline walks the end-to-end architecture first.
Can I host my own MCP server?
Yes, and self-hosting is the normal case rather than the exception. The specification defines two standard transports and both are yours to run: stdio, where the client launches your server as a subprocess and exchanges newline-delimited JSON-RPC over its standard streams, and Streamable HTTP, where you expose one HTTPS endpoint accepting POST.
For a remote server, three requirements come straight from revision 2026-07-28. Validate the Origin header on every incoming connection and return 403 Forbidden on a bad one, because without that check a malicious website can reach a local MCP server through DNS rebinding. Bind to 127.0.0.1 rather than 0.0.0.0 when running locally. And if you want authorization, the spec builds it on OAuth 2.1 with Protected Resource Metadata (RFC 9728) for discovery and Resource Indicators (RFC 8707) so tokens are bound to your server as their audience. Authorization is optional in the spec; a server holding your paid API key on the open internet without it is still a bad idea. On stdio the spec says explicitly not to use that flow and to read credentials from the environment instead.
The heavier lift is not hosting, it is keeping up. The protocol has changed shape twice in under two years, and a hosted server is a maintenance commitment in a way a REST endpoint is not.
Can you give me an example of an MCP server?
The most widely deployed examples are reference servers for local capabilities: a filesystem server exposing read, write and search over a directory, a git server exposing repository history, a database server exposing queries. Each publishes a handful of tools over stdio and is launched as a subprocess by whichever client you point at it. The remote category looks like the example in this post: a thin server holding an API credential and exposing two or three operations as tools. A well-built agent business data tool is exactly that shape, one search tool and one status tool with a spend cap behind both, not a generated mirror of every endpoint.
One clarification, because we would rather be precise than flattering: biz collect does not host a public MCP endpoint. There is no URL you can point an MCP client at today. What we ship is the material an MCP server is built from, the OpenAPI 3.1 spec, llms.txt, a stable response shape and wait: true, plus a written recipe for wrapping it. We built that wrapper ourselves, which is where the trade-offs in this post come from. If a vendor tells you their MCP server is "just a switch", ask to see the endpoint.
When to use MCP, and when a plain API is the right call
Most readers of this page should ship a plain API integration and skip MCP entirely, and that is not a knock on MCP.
| Your situation | Ship this | Why |
|---|---|---|
| Backend service, cron job, ETL | Plain REST | Your code knows the call. A schema for a model to read is dead weight. |
| Zapier, Make or n8n workflow | Plain REST | The workflow tool is the client. It already has HTTP nodes and retries. |
| One model, one tool, one runtime | Native tool definition | One function schema is less work than running a server. |
| Chat assistant that should reach for live data | MCP server | Discovery and argument selection are the whole feature. |
| Coding agent or IDE assistant | MCP server | The client already speaks MCP. Installing beats integrating. |
| Many models or clients, one data source | MCP server | Describe once, consume everywhere. This is where MCP pays for itself. |
| Autonomous long-running agent | MCP server plus hard spend caps | Discovery helps. Unbounded tool calls do not. |
Two rows deserve a sentence each. "One model, one tool" is the row people get wrong most often: wiring a single assistant to a single API with a native tool definition is genuinely less work than standing up, hosting, securing and versioning a server. Start there and add MCP when the second client appears. And the workflow row matters because no-code tools already solved this with HTTP nodes: if you are building in n8n or Make, the n8n workflow and Make.com workflow guides show the shape, and neither needs a protocol.
The honest conclusion
MCP is a good standard solving a real and narrow problem: models could not discover tools, and every integration was bespoke. It solves that. It does not solve pagination, idempotency, push notification, deterministic error handling or cost control, and on those five it is a step backwards from a well-built REST API, because those are properties of a contract between two pieces of software and MCP's second party is a model.
So the decision rule is one question, and it is not about MCP at all: does a model choose the arguments? If yes, an MCP server is probably the cleanest way to let it. If no, you want the API, and most of what you have read about MCP this year is a distraction from shipping.
biz collect is built for both answers: an async REST API with an OpenAPI 3.1 spec, webhooks, idempotency keys, per-plan rate limits and a wait: true synchronous mode, which is exactly the surface that makes a thin MCP wrapper possible when you need one and unnecessary when you do not. Start with 200 signup credits and no credit card, make the call directly, and add the protocol layer only when a model is the one deciding what to ask for.
Read the OpenAPI spec, then ship. Open the API docs.
FAQ: MCP server vs API
Frequently asked questions
- What is an MCP server?
- An MCP server is a program implementing the Model Context Protocol, an open standard for connecting AI applications to external tools and data. It exposes three kinds of primitive: tools (functions the model calls), resources (data the client attaches as context) and prompts (templates a user invokes). A compatible client discovers what the server offers and lets the model call it with structured arguments. The current specification revision is 2026-07-28.
- Why use MCP instead of API glue code you already wrote?
- You do not use MCP instead of an API; you use it in front of one, replacing the per-runtime glue rather than the endpoint. It adds three things: discovery (a model finds your tools at runtime through tools/list), argument selection (the model fills in the JSON Schema itself from natural language) and no glue code (you describe the tool once instead of once per runtime). If your own code already knows which endpoint to call and with what, none of the three apply and a direct API call is the better choice.
- What is an MCP server in AI?
- The MCP server meaning is easiest to see from the model's side: a model only produces text, and the server gives it hands. It exposes callable functions with machine-readable schemas, so the model can request a real action such as querying a database or calling an API and receive a structured result back. MCP standardizes how that catalogue of actions is published and invoked, so one server works with many different AI clients.
- Is MCP the same as REST?
- No, but MCP runs over HTTP and is shaped like an API. Its Streamable HTTP transport requires a single endpoint accepting POST, and every message is JSON-RPC 2.0 in that POST body, answered with JSON or a Server-Sent Events stream. REST is an architectural style across many resource URLs; MCP is a fixed method vocabulary (tools/list, tools/call, resources/read and a few more) over one URL.
- Does biz collect have an MCP server?
- Not a hosted one. There is no public MCP endpoint to point a client at. biz collect ships what an MCP server is built from: an OpenAPI 3.1 spec at /openapi.json, llms.txt and llms-full.txt, a stable JSON response shape, and a wait:true synchronous mode that collapses the async search-and-poll cycle into a single tool call. Wrapping those as a server is a small amount of work you or a generator does, and the build guide has the recipe.
- Does using MCP cost more than calling the API directly?
- The API calls cost exactly the same, because the handler makes the same request. The extra cost is tokens: tool names, descriptions and JSON Schemas are billed as input tokens on every request in the conversation, plus a fixed tool-use system prompt of 286 tokens on Claude Opus 5 with tool_choice auto. Prompt caching cuts the repeat cost to roughly a tenth, so keep the tool set small, stable and deterministically ordered.





