A paid AI scoping engagement is worth roughly $10K to $25K in 2026 — but only if it produces six specific artifacts. Without them, what you bought was lead generation for the downstream build. The six are a capability map, a 30-to-80 case eval set, a quality threshold tied to it, a fallback and refusal design, a cost-per-query worksheet, and a written go-or-no-go memo. Everything else a scoping page sells is decoration.
This is the buyer-side worksheet — what each artifact is, what it costs as a line, and the vendor red flags that signal you are about to pay $25K for slideware. Companion reading: the eval-first build playbook, the idea-to-product manifesto, and the AI agency discovery week.
Why AI scoping is its own paid category
Frontier models collapsed the AI build to eight weeks but did not shorten the decision surface a founder commits to before signing a build SOW. If anything, the decisions got harder — the same model that compressed the build widened the space of plausible product shapes.
Scoping is the engagement that converts those decisions into a fixed contract layer. Not “discovery” in the 2018 SaaS sense — interviews, personas, opportunity sizing — but a paid two-week sprint that produces the artifacts a builder grades the work against.
The reason scoping deserves a separate fee is structural. The most senior labour in an AI product is the eval set and the capability decomposition. Agencies that fold scoping into the build at junior-rate “discovery” hours under-resource the most important phase to protect their build margin. McKinsey’s 2024 State of AI survey found that roughly 60% of enterprise AI projects fail to reach production, with unclear success criteria as the single most common driver. A separately-paid scoping engagement is the buyer-side fix.
The six deliverables, line-by-line
A defensible 2026 scoping engagement produces these six artifacts. Anything missing is a refund line.
1. The capability map
A decomposition of the product into the model capabilities it requires — classification, extraction, generation, retrieval, planning, tool use, vision — each graded against the chosen frontier model on a probe of representative inputs. Output is a one-page matrix: capabilities on the rows, the candidate model (Claude Opus 4.8, Claude Sonnet 4.6, GPT-5, Gemini 2.5 Pro) in a column, and a 1-to-10 probe score in each cell. The map is what separates “we will build an AI assistant” from “the product needs classification at 90% and extraction at 85%; both clear on Sonnet 4.6, tool use needs a routing layer.”
2. The eval set
A fixed corpus of 30 to 80 representative test cases with rubric-scored expected outputs and a documented sampling rationale. Two independent graders score the set; the engagement does not close until inter-rater agreement clears 80%. This is the eval-first PRD’s operational artifact — see the eval rubric template for the six fillable sections a non-engineer founder verifies in twenty minutes. Without it, “production-ready” is a marketing word.
The eval set is the single most valuable artifact in the pack because it is the only deliverable that survives the next frontier model release — when the founder swaps Claude Sonnet 4.6 for a later release in 2027, the eval set tells them in an afternoon whether the swap holds the quality bar.
3. The quality threshold
A numeric pass criterion tied to the eval set, written into the SOW. Two layers — a hard floor below which the build is not shippable (e.g., 80% rubric pass) and a target it is graded against (e.g., 90% with no critical-category failures). The threshold is what makes “AI feature done” a binary state instead of an opinion, and what converts an open-ended T&M agency build into a fixed-price-able engagement. Without it the builder cannot estimate, the founder cannot accept, and the change-order economics tilt to the vendor. The stop-scoping-features-in-user-stories piece explains why this is the central swap.
4. The fallback and refusal design
A written policy for what the product does when the model is wrong, uncertain, or asked something out of scope. Four cells minimum — confident-correct, confident-wrong, uncertain, refusal — each with a user-facing behaviour, a logging requirement, and a human-loop path where applicable. The hallucination budget (wrong-with-confidence answers per 1,000 queries before the trust contract breaks) lives here. Most scoping engagements omit this and treat error behaviour as a Phase 2 problem. It is a Day 1 spec item — the product’s brand promise is a function of how visibly it fails, not how often it succeeds.
5. The cost-per-query worksheet
The unit-economics model: average input tokens, average output tokens, model price per million tokens, expected calls per session, expected sessions per active user per month. Output is a single number — cost-per-active-user per month — plus a sensitivity table for the three variables that move it most (output-token length, retry rate, multi-step agent depth).
If cost-per-active-user lands above the achievable subscription price, the product is unshippable as designed, and the scoping engagement should say so in writing before the founder spends $150K on a build that cannot make gross margin.
6. The go-or-no-go memo
A two-to-four page written recommendation with a defensible call. Not “here are the options” — that is a slide deck. A recommendation, signed by the scoping lead, with rationale and a named risk register. Three possible answers: go (build now, here is the shape), no-go (do not build, here is what needs to be true first), or build-after-conditions (build once these two things resolve; here is the test).
The memo gives the founder permission to walk away. Vendors who refuse to write a no-go recommendation are not selling scoping — they are selling a soft commit to a build.
What each deliverable costs in 2026
The honest 2026 market range for a complete scoping engagement on a focused single-task AI product is roughly $10K to $25K total, with $15K as the typical centre. The split tracks the artifact list:
| Deliverable | 2026 range | What drives variance |
|---|---|---|
| Capability map + model probes | $2K–$5K | Number of capabilities; whether two frontier vendors are probed; depth of probe corpus |
| Eval set (30–80 cases, 2-grader IRR) | $3K–$7K | Domain difficulty; whether the founder supplies gold-standard cases or a domain expert is retained |
| Quality threshold definition | Included in eval set or $1K addendum | Almost always rolled into the eval set fee |
| Fallback + refusal design | $1K–$3K | Compliance footprint; whether refusal taxonomy needs domain-counsel review |
| Cost-per-query worksheet | $1K–$2K | Number of frontier vendors modelled; whether agent depth requires simulation |
| Go-or-no-go memo | $2K–$5K | Risk-register depth; whether the recommendation requires partner sign-off |
| Project orchestration + writeup | $1K–$3K | Founder-facing meeting count; documentation polish |
| Total | $10K–$25K (typical $15K) | — |
Ranges are triangulated from BCG’s 2025 AI build cost benchmark (median six-week discovery $25K–$60K when separately scoped — the two-week version compresses on that), publicly posted agency rate cards (Designli, Markovate, Apexon), and 2026 buyer-side patterns. Below $10K means the artifact list is being cut; above $25K usually means scope-creep into early prototyping.
What the $15K does not buy: a prototype, a mockup, an architecture document for the full build, or any code commit beyond the model probes. A scoping that ships code is mis-scoped — that is the M2 phase of the idea-to-product-as-a-service engagement, billed separately at $60K–$110K. What pushes the fee above range: ensembling two frontier vendors adds $3K–$5K; SOC 2 or HIPAA refusal-policy review adds $5K–$10K; an external domain expert grading the eval set adds $3K–$8K.
What scoping engagements without these artifacts are actually selling
If the proposal in front of you does not name the six artifacts, the engagement is selling one of three things.
Lead-gen for the build. The “discovery sprint” is priced at-cost because the vendor has already underwritten the build fee. The deliverables are a slide deck and an architecture sketch. The engagement closes with a soft transition into a $150K–$400K build proposal where the scoping fee is “credited back.” Walk if the scoping cannot stand alone.
Badge of effort. The vendor delivers a polished document — stakeholder interviews, persona maps, opportunity matrices. It looks like real work. It contains nothing the builder can grade against. Six weeks in, when the build is wobbling, you discover the eval set was never written and the fallback design lives in a Slack thread. The tell is the deliverable list: if “capability matrix” and “eval set” are not in the SOW, the badge is what you are buying.
Confidence laundering. The vendor delivers a recommendation but refuses to write a no-go. Every option is “promising,” every risk “mitigatable.” You bought permission to spend, not permission to walk away. The tell is the language — “considerations,” “trade-offs,” “we recommend exploring” instead of “go, no-go, or build-after-conditions.”
Vendor red flags
Six signals, in rough order of severity.
No separately-printed scoping fee. “Free discovery” or “credited back to the build” means the scoping cannot survive on its own. The vendor has already decided you are buying the build.
No eval set in the deliverable list. The single most important artifact in modern AI scoping. If it is not named, the vendor either does not know how to build one or treats it as internal QA. Either way, walk.
Capability map without probe scores. “We assessed five capabilities” is not a capability map. A real one has numeric probe scores per capability. Without numbers, the vendor is selling vocabulary, not assessment.
No written quality threshold. “Production-ready,” “high quality,” “robust performance” — without a numeric pass rate on a named eval set, these are marketing words. Demand the number in writing.
Refusal to name a frontier vendor. A scoping that ends with “we can use any of GPT, Claude, Gemini, or Llama” has not scoped. A real scoping picks one for M1, with rationale and a swap plan.
The recommendation cannot say no. Ask directly: “Have you written a no-go memo in the last twelve months?” If the answer is “we always find a way to build,” you are paying for confidence laundering.
How to procure it cleanly
A three-step buyer-side frame.
Step 1 — RFP with the deliverable list. Send the six artifacts as the RFP body. Ask vendors to price each line. Vendors who refuse to itemise are signalling that they cannot deliver the lines discretely.
Step 2 — Ask for one redacted past artifact. A redacted eval set from a prior engagement — not a slide, not a case study, the artifact itself with rubric, grader IRR, and capability column. A vendor who cannot produce one has not delivered the artifact in production before.
Step 3 — Bind the scoping to the build SOW. The build’s acceptance criterion must reference the scoping eval set and threshold by name — not a new eval set the build vendor introduces in week three. Bind the references in the procurement docs and the bait-and-switch becomes structurally difficult.
A clean scoping engagement is the cheapest insurance a founder buys against a runaway build. Fifteen thousand dollars before signing a $150K SOW is roughly 10% of the build budget — a defensible premium against the 60% project-failure rate McKinsey publishes.
Book a 30-minute idea review. Bring a one-page product description, the user persona, and the back-of-envelope budget. We say within thirty minutes whether you need full scoping, the lighter 5-day discovery week, or upstream validation first. Schedule the review.
FAQ
How much does an AI scoping service cost in 2026?
Roughly $10K to $25K total for a complete two-week engagement on a focused single-task AI product, with $15K as the typical centre. The split: $2K–$5K capability map, $3K–$7K eval set, $1K–$3K fallback design, $1K–$2K cost-per-query worksheet, $2K–$5K go-or-no-go memo, $1K–$3K orchestration. Below $10K means the artifact list is being cut; above $25K usually means scope-creep.
Why is scoping a paid engagement at all?
Because the senior labour in an AI product is the eval set and capability decomposition. Agencies that fold scoping into a longer build under-resource those phases to protect build margin. A separately-paid fee aligns incentives — the vendor produces six artifacts a buyer can take to any builder, not a soft commit to a $150K SOW.
What is the difference between scoping and discovery?
Discovery in the 2018 SaaS sense means stakeholder interviews and persona work. Scoping in the 2026 AI sense means producing the six artifacts. If your vendor uses “discovery” as a synonym for “scoping,” ask which artifacts they ship.
Can I do AI feature scoping in-house?
Sometimes. A technical founder with eval-design experience can produce the six artifacts in two to three weeks of focused work. A non-engineer founder usually cannot — the eval set and capability probing require both domain knowledge and frontier-model fluency. The honest test: can you write a 30-case eval set with a two-grader rubric this week? If yes, run scoping in-house. If no, pay for it.
Will the scoping vendor also build the product?
They might, but the engagement does not require it. If the vendor refuses to deliver scoping unless they also get the build, the scoping was never standalone — it was lead generation, priced as a discount.
How long does the scoping take?
Two calendar weeks, roughly 50 to 80 partner hours. One week is feasible if the founder arrives with a product description, a named persona, and access to a domain expert. Past three weeks means scope-creep into M2 prototyping.
What if my eval set comes back below the quality threshold?
That is the scoping working as designed. The memo recommends either no-go (the chosen model cannot clear the bar — wait for a release or redesign the task) or build-after-conditions (a routing layer, fine-tuning, or RAG retrieval changes the rubric). A scoping that always returns “go” is confidence laundering.
Should the scoping deliverables go into the build SOW?
Yes. Bind them by name. The build acceptance criterion must reference the scoping eval set and quality threshold — not a new eval set the build vendor produces in week three. This is the structural protection against bait-and-switch on quality bars.
Do I need scoping if my product is “just a chatbot”?
Especially then. “Just a chatbot” is the highest-variance phrase in AI procurement — anywhere from a $5K weekend prototype to a $400K compliance-grade conversational product depending on failure-mode design, eval set, and cost-per-query envelope.
How does AI scoping compare to a general product discovery?
Generic product discovery (Designli, IDEO-flavour) costs $30K–$60K and ships personas, journey maps, and a feature roadmap. AI feature scoping at $10K–$25K is narrower and harder — eval sets, capability probes, and threshold contracts. If you can afford only one, buy the AI scoping; it is what the builder grades against.
Arthur Wandzel