A client once asked me, in the politest possible way, why a prompt cost eight thousand dollars.
It's a fair question. He had ChatGPT open on the other monitor. He'd typed things into it. From where he sat, we were charging him a month of a good salary for a paragraph of English he could have written himself on a Tuesday afternoon.
The honest answer is that we weren't selling him a paragraph. We were selling him the reason the paragraph worked on the four hundred and eleventh customer email as well as it did on the first. But I couldn't say that cleanly at the time, because our proposal said "prompt engineering — 40 hours" and nothing else. The line item invited the argument. We deserved it.
Before starting Scopeyard I spent years running a product studio, and most of the last three delivering AI automation into healthcare, recruitment and operations teams. The single most expensive mistake I see AI agencies make is not underpricing. It's selling an invisible deliverable and then acting surprised when the client values it at zero.
Prompt Deliverable Value = Context Design × Eval Coverage × Change Control
If any of those three is missing, you are selling a paragraph. Here is how to sell the other two.
1. The prompt is the smallest artefact you produce
Ask any agency what they shipped on an LLM project and they'll say "the prompt". Ask them what took the six weeks and they'll list twenty other things: the retrieval logic, the schema the model has to return, the tool definitions, the fallback when the model refuses, the guardrail on personally identifiable information, the four rounds of failure analysis against real client data.
That gap between what you shipped and what you named is where your margin dies.
It also isn't just a naming problem any more. The 2026 State of Context Management Report found 82% of IT and data leaders say prompt engineering on its own no longer scales their AI work. The industry has already moved to calling this context engineering — deciding what information lands in the model's window at each step, and what you deliberately keep out. Agents fail on state management, not on adjectives. If your deliverable list still stops at "the prompt", you are describing 2023 work in a 2026 market.
Rename the deliverable to what it is. Then invoice for it.
2. Write the context contract before you write a word of prompt
Every LLM feature we deliver now starts with a one-page document I call the context contract. It's boring and it prevents almost every argument that comes later.
It states, in plain language:
- What goes into the model's context at each step, and from where
- What is deliberately excluded, and why
- The exact output shape — the JSON schema, the field types, what's optional
- What the system does when the model fails, refuses or times out
- Which decisions a human must still make
That document takes two to four hours to write and it's the thing you show when a client asks what they're paying for. It's also the thing that makes the build estimable. You cannot estimate "make the AI good". You can estimate seven retrieval sources, one schema, three failure paths.
Get it signed. Same as any scope. The same discipline I've written about for scoping AI automation projects applies here at the feature level.
3. Sell the eval suite as a line item, not as goodwill
This is the change that shifted our AI work from arguments to renewals.
A prompt without evals is a demo. It works on the three examples you tried and nobody, including you, knows what it does on the fourth hundred. Clients feel this even when they can't articulate it, which is why they haggle.
Standard practice in 2026 is tiered: teams run a 20 to 50 prompt smoke set on every change, a 200 to 500 prompt regression set before merging, and a 1,000-plus prompt benchmark at release. Deterministic checks run first because they're fast and cheap — regex, JSON schema validation, exact match, code execution — and one or two LLM-judge metrics catch the open-ended failures like tone, hallucination and instruction adherence.
You do not need all three tiers for a client chatbot. You do need one. A 60-case golden set built from the client's own historical data, with pass thresholds they agreed to, changes the entire conversation. Now you're not asking them to trust you. You're showing them 57 of 60, and the three that failed, and why.
Price it separately. On a mid-sized build ours typically runs 15–25% of the feature cost, and it is the easiest line item to defend in a renewal because it's the only one that produces a number.
4. Price the artefacts, not the typing
Once the deliverable is named properly, the pricing follows. Published 2026 rates for prompt specialists sit at roughly $50–100/hour for beginners, $175–300 for senior, and $250–500 for expert-level work on agent systems — but hourly is the worst way to sell this, because it makes the client mentally price your typing speed.
Sell the artefacts instead. Rough bands I'd defend, consistent with what's being quoted in the market for 2026:
| Deliverable | What the client gets | Typical band |
|---|---|---|
| Context contract + prompt spec | One-page contract, output schema, failure paths | $1,500–3,000 |
| System prompt + eval suite (single workflow) | Working prompt, 50–80 case golden set, thresholds | $4,000–8,000 |
| Retrieval pipeline over client documents | Ingestion, chunking, retrieval, grounding checks | $8,000–15,000 |
| Full AI feature in production | The above plus integration, monitoring, handover | $15,000–50,000 |
| Advisory retainer | Review, tuning, model migration guidance | $2,000–5,000/mo |
| Embedded delivery retainer | Named engineer, ongoing eval and drift work | $8,000–20,000/mo |
The bands matter less than the structure. Six named things a client can point at beats one line saying "prompt engineering — 40 hours" every single time.
5. Put model change control in the contract
Model lifecycles have compressed from roughly 18–24 months to 6–12. Deprecation is now a standing item on every platform roadmap, and here's the part that catches agencies out: when a model shifts underneath you, nothing errors. The endpoint still returns 200s. Tool-call formatting drifts, JSON adherence slips, refusal boundaries move, and your accuracy quietly walks from 91% to 84% over four months.
Nobody notices until a client complains. Then you're doing a week of unpaid forensics on a project you closed in March.
So write it in. A change control clause that says: when the underlying model version changes, we re-run the eval suite and report; remediation is billed at the agreed rate or drawn from the maintenance retainer. Clients accept this readily once you explain that the model isn't yours and its vendor doesn't ask your permission. What they won't accept is a surprise invoice with no clause behind it.
This is the same reasoning behind what belongs in an AI maintenance contract, and it's worth reading alongside this if you're rewriting your MSA.
6. Budget the second year honestly
The prevailing guidance is to reserve 20–30% of Year 1 build cost annually for maintenance. For AI work that's the floor, not the ceiling, because Year 2 costs are almost entirely maintenance: model migrations, drift remediation, monitoring, human review queues, reliability work.
Say this out loud during the sale. A client who budgets $60,000 for the build and hears "expect $12,000–18,000 a year to keep it working" respects you more, not less. A client who hears nothing and gets a bill in month fourteen churns.
It also gives you a natural retainer conversation, which is a far better place to be than pitching the move from projects to retainers cold, eighteen months later.
7. Hand over something the client can own
MIT's Project NANDA found around 95% of generative AI pilots deliver no measurable return on the P&L, and separate research puts the share of pilots that never reach production as high as 88%. Read those numbers carefully. They are not mostly a model quality problem. They're an ownership problem — the pilot works, the person who built it leaves, nobody inside the business knows how to change it, and it dies.
Your handover pack is a deliverable. It should include the context contract, the prompt with its version history, the eval suite and how to run it, the thresholds and what to do when one fails, the model version pinned, and a named human owner on the client side.
We run this as an explicit milestone with a client sign-off gate rather than an email attachment at the end, which is exactly the workflow Scopeyard exists to enforce — deliverables, review, approval, receipt. The handover being approved is what closes the project. Not the demo working.
Final thoughts
There is nothing wrong with charging serious money for prompt work. There's a great deal wrong with charging serious money and refusing to say what it buys.
Every argument I've had with a client about AI pricing traced back to the same root cause: I'd sold something they couldn't see. The fix wasn't a better rate card or a firmer tone. It was naming six artefacts instead of one, attaching a number to each, and letting the eval suite do the arguing on my behalf.
Stop selling the prompt. Sell the evidence it works, and the plan for when it stops.