A client asked me last quarter whether they needed an MCP server. When I asked what they wanted it to do, the answer was "connect our AI to our data." That's not a scope. That's a shrug with a budget attached.
I hear a version of that conversation most weeks now, and I hear the mirror image of it from agency founders: a prospect has asked for "an MCP server" and nobody in the room is certain enough about what that means to quote it. So the agency either lowballs a REST wrapper or panics and quotes enterprise money for a weekend of work.
Before starting Scopeyard, I spent years running a product development studio and then delivering AI automation projects across healthcare, recruitment and operations. We built the Scopeyard MCP ourselves, so I've been on both sides of this — the agency quoting it, and the product team living with it afterwards.
Here's the thing nobody says plainly: an MCP server is not a technical achievement. Wiring one up is a day or two of work. What takes the time, and what the client is actually paying for, is deciding what an autonomous agent is allowed to do inside their business.
The thesis in one line:
Useful MCP Server = Real Actions × Correct Permissions × Small Tool Surface × Audit Trail
Drop any term and you've shipped a demo.
1. What an MCP server actually is, in agency language
The Model Context Protocol is a standard way for an AI application to talk to a system that isn't the AI application. Your client has a CRM, a booking system, a billing platform, an internal database. An MCP server is a thin layer in front of one of those that exposes a fixed list of things an agent may do — "list open invoices", "create a booking", "move a deliverable to review" — in a format any MCP-speaking client can discover and call.
Two comparisons help clients get it.
It is not a chatbot. A chatbot answers. An MCP server lets an agent act — and the whole reason it's interesting is that the acting happens inside a system of record, not in a chat window that someone then has to copy-paste out of.
It is not just an API. Your client probably already has an API. The difference is that an API is documentation a developer reads; an MCP server is a tool catalogue a model reads, at runtime, with descriptions written for a model rather than for a human. That shift — from "here are 200 endpoints" to "here are nine things you may do, described in plain language" — is the actual design work.
The ecosystem is past the experimental stage, which is why the requests have started arriving. Across the four Tier 1 SDKs, the maintainers reported close to half a billion downloads a month as of July 2026. The official registry counted roughly 9,700 servers by May 2026, and Stacklok's 2026 software report put 41% of surveyed software organisations in limited or broad production with MCP servers. Honeycomb said publicly that nearly 20% of their monthly interactive queries now come from agents rather than people.
Your clients' customers are going to start showing up as agents. That's the pitch, and it's a real one.
2. The July 2026 spec changed what you're quoting
If you scoped MCP work in 2025 and haven't looked since, your estimate is wrong in both directions.
The 2026-07-28 specification moved MCP to a stateless protocol core. The initialize/initialized handshake and the Mcp-Session-Id header are gone. Every request is self-describing, so any request can land on any instance behind a plain round-robin load balancer. Method and tool names travel in Mcp-Method and Mcp-Name headers, so a gateway can route, rate-limit and authorise without parsing the body.
What that means commercially: the infrastructure line on your estimate got smaller. You are deploying a normal HTTP workload now. No sticky sessions, no held-open streams, no bespoke session store. If you were quoting infra complexity, stop.
What got bigger is everything else.
Server-initiated requests were replaced by Multi Round-Trip Requests. A tool that needs a confirmation returns resultType: "input_required", and the client retries with the answers attached. That's the mechanism for "are you sure you want to delete this?" — and designing those confirmation points is judgement work, not plumbing.
Authorisation hardened. Clients must validate the iss parameter per RFC 9207. Dynamic Client Registration is formally deprecated in favour of Client ID Metadata Documents. Credentials are bound to the issuer that minted them.
And things you may have built on are on the way out: Roots, Sampling and Logging are deprecated, as is the legacy HTTP+SSE transport. There's now a formal twelve-month minimum deprecation window, so you can put migration in a maintenance contract instead of firefighting it.
Practical move: put the spec version in the proposal. "Built to 2026-07-28" is a line item, and it gives you something to charge for in twelve months.
3. Scope tools, not endpoints
The lazy build is a script that turns every API endpoint into a tool. Do not sell this. It produces a server with sixty tools, a model that picks the wrong one, and a client who concludes AI doesn't work.
Scope the way you'd scope any automation: to a workflow with a trigger and a terminal state. Same discipline I use for scoping AI automation projects generally. Ask what an agent should be able to finish, end to end, without a human retyping anything. Then work backwards to the smallest set of tools that completes it.
A useful constraint I've settled on: if you can't fit the tool list on one slide with a one-sentence description each, the surface is too big for the first release. Ship the eight tools that close a loop. Add more when usage tells you which are missing.
Two design details that separate a working server from a demo:
Tool descriptions are product copy. They are the only instructions the model gets. "Returns project data" is useless. "Returns the deliverables in a project, filtered by lane; use this before moving anything" is a specification. Budget real writing time — I'd put a third of the design effort here and I'd have the person who understands the client's domain write it, not the person who wrote the endpoint.
State goes in arguments, not in the transport. The spec's own guidance now is that if a server needs to carry state, it should mint an explicit handle from a tool and have the model pass it back. The model can see the handle and thread it. Hidden state confuses models the same way it confuses juniors.
4. Permissions are the project
Here is the sentence I'd put in every MCP proposal: the deliverable is not access, it is bounded access.
Every tool needs an answer to three questions before it ships. Who can call it? What can it see? What can it change, and can that be undone?
On our own MCP server we made this concrete: the agent acts as the person who connected it, with that person's role. A client can approve or request changes on their own deliverables; only team members can send work for review. The agent never gets a capability the human doesn't already have. That rule is boring and it is the entire safety model.
The reason to be strict is that the threat surface is real and reasonably well documented now. Tool descriptions are unsanitised text that a model treats as instructions, which is why tool poisoning works at all — an academic scan found roughly 5.5% of 1,899 public servers exhibiting it, and a broader scan of 1,808 servers reported 66% with some security finding, including 43% vulnerable to command injection. Those numbers describe public community servers rather than the one you're building for a client, but they describe the default standard of care in the ecosystem, and the default is low.
So: no destructive tool without a confirmation round trip. No credentials in tool arguments. No tool that can read across tenants. Write the audit trail from day one — every call, who made it, what changed. Clients will ask for it in month three anyway, and retrofitting it costs more.
5. What to charge
Published ranges for MCP server work are all over the place in 2026, mostly because "MCP server" describes both a weekend wrapper and a year-long platform. Vendor guides cluster simple single-tool servers at $3,000–$8,000, production integrations with auth and permissions at $8,000–$25,000, and complex multi-tool enterprise builds at $25,000–$60,000+, while consultancies quoting full enterprise programmes talk in six figures.
Here's how I'd band it for an agency selling into small and mid-market clients:
| Tier | What it is | Tool count | Typical fee | Build time |
|---|---|---|---|---|
| Read-only pilot | One system, no writes, internal users only | 3–6 | $4,000–$10,000 | 1–2 weeks |
| Production server | Writes, OAuth, role-based permissions, audit log | 6–12 | $12,000–$30,000 | 3–6 weeks |
| Multi-system build | Two or more backends, tenant isolation, gateway, monitoring | 12–25 | $35,000–$80,000 | 8–14 weeks |
| Spec migration | Move an existing server to 2026-07-28 | n/a | $5,000–$15,000 | 1–3 weeks |
Two notes on using this. First, sell the read-only pilot as a separate, paid engagement before you quote the production build — same logic as charging for a data reality check. You will discover the client's auth situation, and the client's auth situation is where the money goes. Second, the number of tools is the honest cost driver, not the number of systems. Price per tool internally even if you quote a fixed fee.
6. Sell the maintenance contract with it
An MCP server is not a project that finishes. The client's backend will change. The spec will move. Model behaviour will shift and a tool description that worked in March will start getting misused in September.
The commonly quoted industry norm is 20–30% of build cost per year for ongoing maintenance, and for MCP I'd sit at the top of that band rather than the bottom, because two of the three change drivers are outside the client's control. Write it the way you'd write any AI maintenance contract: named tools covered, a response window, a quarterly review of which tools are actually being called, and an explicit line for spec upgrades within the twelve-month deprecation window.
That last item is the easiest retainer conversation you'll ever have. The deprecation policy is published. You can show the client the calendar.
7. Should your agency have one?
Different question from whether your client should. If you run an AI or product agency, an MCP server in front of your own delivery system means your team's agents can read the board, draft review notes, and tell a client what's waiting on them without anyone acting as a human clipboard. It also means you've built one, which is the fastest way to stop guessing at estimates.
That's why Scopeyard ships with an MCP server rather than treating it as an add-on — deliverables, reviews and client approvals are exactly the state an agent needs to be useful about, and it's free on every plan. If you're weighing up how delivery tooling should work for an AI shop specifically, the AI agencies breakdown covers the rest of it.
Final thoughts
The agencies making money on MCP work in 2026 aren't the ones who learned the SDK first. The SDK is a weekend. They're the ones who can sit with a client, list the nine things an agent should be allowed to do inside their business, and defend why the tenth is not on the list.
That's not integration work. That's judgement, and it has always been the thing worth charging for.
Anyone can expose an API to a model. Getting paid properly means being the person who decides what it may not touch.