A fixed-price AI MVP contract is only defensible when both sides agree precisely what is “in scope” — one named feature, expanded into six required layers — and what is “out of scope,” which is everything else. The gap between the partner’s reading and the founder’s reading is where 30%-plus overruns live, and where every mid-engagement change order is born. This page is the contract-signing decomposition: the one-named-feature rule, the six scope layers underneath it, the 2026 out-of-scope list, the change-order trigger taxonomy, the hallucination liability clause, and a 6-week AI MVP SOW snippet.
This is the procurement-edge companion to what a defensible idea-to-product SOW looks like. The SOW article describes the nine load-bearing sections of the contract. This article goes inside the Scope and Out-of-Scope sections and asks: when the partner writes “an AI chat agent for customer support” as the named in-scope feature, what does that string actually buy?
Why fixed-price AI MVP contracts fail
A fixed-price contract fails for the same reason a fixed-price kitchen renovation fails: the parties agreed on a price before they agreed on what was being built. The AI version is harder because the work has more scope layers and the partner has more discretion over which of those layers are inside the price.
Three failure patterns recur in 2026 AI MVP engagements:
- The scope sentence is a feature name, not a feature anatomy. “An AI chat agent for customer support” is treated as one scope line. It is, structurally, six scope layers. Five of them quietly slip out of the fixed price as the engagement progresses.
- The out-of-scope list is implicit. The contract names what is included and treats the rest as obvious. It is not — native mobile, enterprise SSO, model-inference reimbursables, second-tenant migration, multi-language localization. Each is a defensible exclusion, and each becomes a fight if not written down.
- The change-order trigger is left to the partner. “Any change requires a written change order signed by both parties” is the standard clause. The partner decides what counts as a change. The founder discovers in week 4 that what they considered a clarification is being billed.
BCG’s 2024 “Where’s the Value in AI?” study reported that 74% of corporate AI investments fail to scale beyond proof-of-concept. McKinsey’s 2024 State of AI work named scope discipline and eval acceptance as the two largest gaps between the 26% that scale and the 74% that do not. The same pattern shows up at the contract level: founders who counter-sign vague scope pay 30%-plus overruns to repair the gap during the build.
The one-named-feature rule
A defensible 6-week AI MVP fixed price covers one named feature, one persona, one task, one primary frontier-vendor model. That is the rule. Every additional named feature multiplies the scope surface, the eval surface, and the change-order surface. A second feature is a second SOW.
What “one named feature” looks like in practice:
- “An AI chat agent that answers Tier-1 customer-support questions for SaaS-product end users, using one named knowledge base, in English.”
- “An AI document summarizer that produces a 200-word executive summary of contracts up to 50 pages, output to a single named target system.”
- “An AI lead-qualification scorer that returns a 0–100 confidence score on a single named CRM record type.”
What violates the rule: “An AI assistant for our customer-success team” (multiple personas, unbounded), “An AI platform for document workflows” (open-ended), “An AI agent that handles support, leads, and onboarding” (three features dressed as one).
The partner has a structural incentive to write the broader version because it makes the deal sound bigger. The founder has the opposite incentive: the narrower the named feature, the more defensible the fixed price. Insist on the narrow form. The MVP is supposed to be the smallest shippable test of one risk; we cover the underlying argument in the case for the smallest possible AI feature in your MVP.
The 6 scope layers under one AI feature
Every named AI feature has six required scope layers underneath it. A fixed-price contract that names the feature but excludes any layer is structurally underspecified. The partner will deliver layer 1, omit the rest, and the founder will discover the omission in week 5 when they try to use the feature in production.
| Layer | What it is | Why it is required | Common failure mode |
|---|---|---|---|
| 1. Feature shell | The application code that exposes the AI feature — UI, API endpoint, integration glue | Without it the feature is unreachable | Treated as the whole engagement |
| 2. Eval set | The graded test cases and rubric the feature is accepted against | Without it acceptance is subjective | Treated as internal QA, not a deliverable |
| 3. Prompt library | The versioned prompts, system messages, and prompt templates the feature uses at runtime | Without it the feature cannot survive a model upgrade | Retained by the partner as IP |
| 4. Model contract | The frontier-vendor contract that the feature calls at runtime | Without it inference costs are uncontrolled | Billed through the partner with a markup |
| 5. Observability stack | Logging, tracing, and eval-on-production for the feature’s runtime calls | Without it failures are invisible | Out-of-scope by silence |
| 6. On-call window | The 30-day post-launch support during which the partner fixes severity-1 issues | Without it production failures are billable | Vague “we’ll help” clause |
The fixed price has to cover all six. The most common omissions are layers 2, 5, and 6 — eval set, observability, and on-call — because they are the least visible to a founder reading the proposal and the most expensive to deliver well. The single most valuable edit you can make to an inherited AI MVP SOW is to name all six layers under each in-scope feature and require the partner to confirm in writing that all six are inside the fixed fee.
The 2026 out-of-scope list
The mirror of the in-scope decomposition is the explicit out-of-scope list. The 2026 market shape for a 6-week AI MVP fixed price excludes a defensible set of items. Naming them in the SOW prevents them from being mid-engagement change-order arguments.
| Out-of-scope item | Why it is out-of-scope | When it becomes in-scope |
|---|---|---|
| Frontier-model inference costs | Pass-through directly to the founder’s vendor contract from M2 onward | Never in-scope for a 6-week MVP |
| Native mobile (iOS/Android) | A second build with a second skillset, not an MVP scope item | A second SOW, post-MVP |
| Enterprise SSO (SAML, SCIM) | Required for sales to enterprise, not for shipping the MVP | A specific milestone after MVP-1 validates demand |
| SOC 2 audit start | A 6–12-month process, not a 6-week deliverable | A separate compliance engagement |
| Multi-language localization | Second locale roughly doubles the eval set | A second SOW once English is graded |
| Second-tenant migration | Multi-tenant requires architectural work that single-tenant does not | A separate engagement |
| 24/7 production operations | The 30-day on-call is single-shift, business hours, severity-1 only | Ongoing retainer after MVP |
| Data labeling at volume | If labeling exceeds, say, 200 cases, label production is its own deliverable | Carved out as a labeling sub-SOW |
| Third-party integrations beyond one named target | The MVP names one target system; others are change orders | Each additional target is a change order |
| GDPR / HIPAA / FedRAMP-grade compliance build-out | A regulated-data engagement is a different shape | A regulated engagement, scoped separately |
A defensible SOW lists every out-of-scope item explicitly. “Anything not specified is out-of-scope” is true but procedurally weak — it puts the burden on the founder to imagine every excludable item. The defensible move is to flip it: the partner lists the typical exclusions, and the founder reviews and signs.
The change-order trigger taxonomy
The change-order clause is second only to eval acceptance in importance. It defines the conversation that will happen the first time the founder asks for something not literally written in the scope sentence. A defensible clause distinguishes three categories of mid-engagement request:
| Category | Definition | Effect on price |
|---|---|---|
| Clarification | A request that resolves ambiguity in the existing scope without changing it | No change order; no fee adjustment |
| Trade | A request to swap one in-scope item for another of equivalent partner effort | Logged change order; no fee adjustment |
| Change | A request to add scope, expand the eval set, or change the named feature | Priced change order; fee adjustment per change-order rate card |
The taxonomy matters because without it the partner has full discretion, and that discretion has dollar consequences. Every mid-engagement request becomes a negotiation about whether it counts as a change. With the taxonomy, the conversation is structural.
The associated rate card sits inside the SOW. A 2026-typical card prices each priced change at $5K (small), $15K (medium), or $40K (large), each tier defined by the eval-set delta and engineering-week delta. We cover the operational mechanics in the AI agency change-order playbook.
The absorbed-favor trap deserves a specific clause. The partner sometimes offers to handle a small change without a change order, framing it as goodwill. Each absorbed favor sets a precedent that compounds into a 30%-plus overrun argument at M3 when re-priced as scope. The clause: “Every modification to in-scope deliverables, regardless of perceived size, requires an executed Change Order before work begins.” Costless to write, expensive to omit.
Eval acceptance as the in-scope referee
Eval acceptance is the structural mechanism that decides whether the partner has delivered the in-scope feature at quality. It is the referee between “in-scope and delivered” and “in-scope and rejected.” Without a contractual eval acceptance gate, the fixed-price engagement converts to T&M by default the moment quality becomes contested.
A defensible eval-acceptance clause names four things: the eval set itself (typically 50–200 graded test cases, produced at M1, owned by the founder); the rubric (the grading standard, written before M2 starts); the acceptance threshold (a numeric pass rate, e.g. “85% of cases at rubric grade A or B”); and the remediation path (a defined 2-week window inside the fixed fee if the threshold is missed at M2; termination-for-cause if missed at the end of remediation).
The eval acceptance clause is what makes “fixed-price” mean something for an AI engagement. SaaS fixed-price contracts accept on feature-list checkmarks because SaaS features are either present or absent. AI features are present at variable quality, and quality is the load-bearing axis. Without an eval gate, fixed price for AI is fixed-price-only-if-the-partner-agrees-it-passes — which is to say, not fixed-price at all.
The eval set is also the artifact that gives the founder portability across model generations. We cover the eval-first framing in detail in the eval-first PRD.
Hallucination liability: who owns the failure
Inherited templates are silent on hallucination liability. A defensible AI MVP SOW assigns it explicitly, with three components:
- The eval set defines the partner’s quality obligation. Outputs within the rubric’s tolerance are in-scope quality. Outputs outside the tolerance are partner remediation if discovered within 30 days of M2 acceptance, founder responsibility thereafter.
- A customer-visible failure clause names the partner’s obligation if a model output causes a contractual or reputational issue inside the 30-day on-call window — root-cause within 24 hours, remediation within 5 business days, at no additional fee.
- Indemnity for negligent design — if the model is deployed without an eval set, without observability, or with a known failure mode the partner did not disclose, the partner indemnifies the founder for direct damages up to the contract value.
The clause does not make the partner responsible for every possible model failure. Frontier models hallucinate; the question is whether the system around the model catches and contains the failure. The partner ships the containment system, the eval set, and the on-call. The founder accepts the residual rate documented in the eval results.
Worked SOW snippet (6-week AI MVP)
A fragment of an actual 6-week AI MVP SOW. Names and figures are illustrative.
Engagement: AI chat agent for customer support, MVP-1, 6 weeks. Fixed Fee: $145,000, milestone-billed.
Scope (in-scope deliverables) — the Partner will deliver the following, all inside the Fixed Fee:
- Feature shell. Chat-agent UI integrated into one named customer-support portal. React front-end, Node back-end, single REST endpoint to the model layer.
- Eval set. 120 graded test cases covering Tier-1 customer-support intents, produced at M1, owned by the Client. 4-grade rubric (A–D), graded by a Client-named SME.
- Prompt library. A versioned prompt library stored in the Client’s repository, assigned to the Client at M4.
- Model contract. Integration against one named frontier-vendor model (Anthropic Claude Sonnet 4.6 as of contract date). Client owns the vendor contract from M2 onward.
- Observability stack. Structured logging, traces, and eval-on-production sampling at 5% of live traffic, written to the Client’s named observability target.
- On-call window. 30 calendar days, business-hours severity-1 on-call from M4 acceptance, 4-hour first-response and 5-business-day remediation SLAs.
Out-of-Scope (explicit exclusions) — each available as a separate engagement:
- Frontier-model inference costs (pass-through from M2); native iOS / Android; enterprise SSO; SOC 2 audit; localization beyond English; multi-tenant migration; 24/7 ops; data labeling beyond the 120-case eval set; integration with any system other than the named portal; HIPAA / GDPR / FedRAMP build-out.
Change-Order Policy — every modification to in-scope deliverables, regardless of perceived size, requires an executed Change Order before work begins. Priced per Exhibit B Rate Card: small ($5K), medium ($15K), large ($40K). Clarifications are not Change Orders. Trades of equivalent partner effort are logged but not priced.
The load-bearing structure is exactly this: one named feature, six scope layers, an explicit out-of-scope list, and a change-order policy with a rate card. Every other section of the contract derives from this skeleton. For the underlying pricing math, see the AI MVP economics playbook.
Counter-signing checklist
Before you counter-sign a fixed-price AI MVP SOW, run the proposal against this checklist.
- Is the scope statement one named feature, one persona, one task, one primary model? If broader, push the partner to narrow it before signing.
- Are all six scope layers under that feature named inside the fixed fee? Feature shell, eval set, prompt library, model contract, observability stack, on-call window. If any is silent, ask whether it is included; if the answer is “we’ll cover it as part of the build,” require it in writing.
- Is the out-of-scope list explicit and exhaustive for the 2026 norms? Mobile, SSO, SOC 2, localization, multi-tenant, 24/7 ops, regulated-data compliance, additional integrations.
- Does the change-order clause name clarifications, trades, and changes distinctly? If it lumps them together, the partner has full discretion.
- Is there a published change-order rate card with small/medium/large tiers and dollar values? Without a rate card, change-order pricing is a per-request negotiation.
- Is the absorbed-favor trap explicitly refused in writing? “Every modification, regardless of perceived size, requires an executed Change Order.”
- Is eval acceptance the contractual gate on M2? With a numeric threshold, a rubric, and a remediation window.
- Is the prompt library and the eval set assigned to the founder at handoff? If either is partner-retained, refuse and renegotiate.
- Is hallucination liability assigned via the eval set and the on-call window? Not silent. Not “as-is.” Explicit.
- Are model-inference costs carved out as direct founder pass-through? Never billed through the partner with a markup.
A 10/10 is rare. An 8/10 with the missing items added by amendment is signable. A 5/10 or worse is a deferred problem.
FAQ
Is a true fixed-price AI MVP contract even possible in 2026?
Yes — for one named feature, one persona, one task, one model, over 6–12 weeks, with eval acceptance as the gate. Outside those constraints, fixed-price collapses into T&M-with-a-cap. A partner who resists the narrow form is pricing in heavy contingency or planning to recover on change orders.
What is the single most-omitted scope layer?
The eval set. Most inherited templates treat eval as internal QA, not a contractual deliverable. M2 acceptance then becomes subjective and the engagement converts to T&M the first time quality is contested. Insist on a numbered eval set as an M1 deliverable owned by the founder.
Should model-inference costs ever be inside the fixed fee?
No. Inference is variable, usage-dependent, and the partner has no incentive to optimize prompt length if the cost is absorbed. Carve inference out as direct founder pass-through to the vendor contract from M2 onward.
What happens if the model is deprecated mid-engagement?
A defensible SOW puts model-version-change support inside the on-call window. If the founder’s primary model deprecates within the 30-day window, the partner supports the swap at no additional fee. The 2026 deprecation cadence is roughly 6–12 months across major frontier vendors.
Can I treat the proposal’s “and more” or “etc.” language as binding?
No. Any in-scope item not named is unenforceable. Strike “and more,” “etc.,” and “including but not limited to” everywhere they appear and replace them with named items. Resistance from the partner is a signal that the language was load-bearing.
What is the absorbed-favor trap?
The partner’s offer to “just handle” a small mid-engagement change without a change order. It feels generous and compounds into a 30%-plus overrun argument at M3 when re-priced as scope. Refuse it in writing — language is in the change-order playbook.
How do I size a change order if the partner’s rate card seems aggressive?
The 2026 typical rate card is $5K (small) for under one engineer-week, $15K (medium) for one to two weeks, and $40K (large) for over two weeks. Rates above this for a sub-$200K engagement are aggressive — ask the partner to justify against eval-set delta and engineering-week delta.
Who owns hallucination liability for customer-visible failures?
Inside the 30-day on-call window, the partner — if the failure is for an input case represented in the eval set. Outside the window or outside the eval set, the founder. The clause names this so neither party is surprised in week 7.
Is a 6-week timeline realistic for a true fixed price?
For one narrowly named feature, yes. We cover the calendar in the 6-week AI MVP. For two or more features, split into two SOWs, sequenced.
What if the partner refuses to name the six scope layers in writing?
Walk. A partner who cannot or will not name the six layers is either inexperienced with AI MVPs or pricing in heavy change-order recovery. Either way, the fixed price in front of you is not a fixed price.
Where to go next
The full program-level view is the idea-to-product manifesto. If you are still deciding whether to engage a partner at all, the idea validation playbook is the master walkthrough. If you have a proposal and are deciding which sections to push back on, what a defensible idea-to-product SOW looks like is the section-by-section companion. For the service-category view, see idea-to-product as a service.
When you are ready for a working session on a proposal, book an idea review. We will read the SOW, run it against the 10-point checklist, and tell you what to renegotiate before you counter-sign.
Arthur Wandzel