Home About Who We Are Team Services Startups Businesses Enterprise Case Studies Industries Commercial Real Estate Blog Guides Contact Connect with Us
Back to Guides
Enterprise Software 13 min read

AI MVP partner pricing red flags: 5 patterns to refuse

AI MVP partner pricing red flags: 5 patterns to refuse

Five pricing patterns predict AI MVP overrun with enough accuracy to be refusal-grade. Each shows up as a specific line in the SOW, each has a commercial logic the partner can defend, and each quietly transfers three AI-specific unknowns — eval-threshold variance, inference cost variance, model-fallback work — from partner to founder. This piece names the five: SOW wording, why it predicts overrun, push-back script, and the clause to write instead.

The patterns sit inside the AI MVP economics playbook and under the idea-to-product manifesto. For seven engagement-wide signals one altitude higher, see idea-to-product red flags: 7 signals you’re about to overpay.

Why pricing patterns predict AI MVP overrun

Pre-LLM software had three primary cost drivers: spec ambiguity, integration surface, team velocity. A 2026 AI MVP adds three — eval-threshold variance (the gap between “feature works” and “feature clears the rubric”), inference cost variance (token spend drift between pilot and production), and model-fallback work (engineering when the primary model deprecates, gets rate-limited, or under-performs).

McKinsey’s State of AI 2025 names scope discipline and acceptance criteria as the two largest gaps between AI projects that scale and the 78% that stall. BCG’s Where’s the Value in AI? finds 74% fail to scale, with acceptance ambiguity a dominant cause. Bain’s Technology Report 2025 adds that vendor scoping correlates with outcome more than vendor experience.

Each pattern removes one AI-specific surface from the contract before work starts. Two are defensible in narrow conditions. Refuse the line, not the partner — and replace it with one that survives an AI build.

Pattern 1: All-in lump-sum with no milestones

SOW signature. “Total engagement fee: $X, payable 40% on signing, 30% at midpoint, 30% on final delivery.”

One number, three calendar-anchored payments, one acceptance event at the end. No eval threshold, no information artifact tied to any tranche, no founder veto between signing and acceptance.

Why it predicts overrun. The midpoint payment lands before the partner has run a single eval against the production rubric — a calendar event dressed as a milestone. The final 30% is the only payment tied to a deliverable, and that deliverable is “scope” — a noun resolving to whatever the partner declares finished on day 56. Quality drift accumulates in silence. See the case against fixed-price AI development contracts.

What to ask. “At what numeric eval-threshold score against a co-authored rubric does the final 30% release?” If the answer is “mutual agreement” or “signed acceptance,” the line is the pattern.

How to counter. “The lump-sum fee works; the payment schedule does not. I need tranches 2 and 3 gated on a numeric eval clearance against a rubric we co-author in the first two weeks.”

What to write instead.

“Total fee: $X, in four tranches against a co-authored evaluation set. 15% on planning sign-off; 25% on PRD and eval-set sign-off; 45% on the build clearing the N% eval threshold; 15% on runbook acceptance. First cure attempt inside the original fee, up to 2 calendar weeks.”

See the case for milestone-based AI MVP billing for the canonical structure.

Pattern 2: Per-hour with no cap

SOW signature. “Billed at $X/hour for engineering, $Y/hour for ML and evaluation, weekly invoicing against logged hours. Estimated total: $Z (non-binding).”

The contract is the hourly rate and the weekly invoice; “non-binding estimate” is the load-bearing line.

Why it predicts overrun. A non-binding estimate is the floor of what the partner intends to bill if nothing goes wrong. In an AI MVP three things typically do: the first model choice under-performs (fallback work), production token spend exceeds the pilot estimate by 2–4x (inference variance), and retrieval needs a redesign after the first eval batch (rework). Each is billable; none is in the estimate. By week 6 the engagement has billed 60% of the estimate against 30% of a build that has not yet been eval-tested. Stack Overflow’s Developer Survey 2025 shows AI-adjacent engineering tasks overrun pre-build estimates by 30–80%.

What to ask. “What is the not-to-exceed cap, and what happens if we pass it without clearing the rubric?” If there is no cap, or the cap is 1.5x the estimate without a numeric acceptance event, the line is the pattern.

How to counter. “Per-hour with weekly invoicing is workable; the non-binding estimate is not. I need a not-to-exceed cap at the estimate level, with eval-threshold acceptance as the release condition.”

What to write instead.

“Billed at $X/hour against a not-to-exceed cap of $Y. Cap may not be exceeded without a written change order signed by the buyer. Acceptance event: build clears N% threshold on co-authored eval set; final 20% of cap releases on acceptance. Hours billed past cap without acceptance are at partner risk.”

See AI MVP fixed-price vs milestone billing for the model-by-model comparison.

Pattern 3: $20K discovery converting to T&M build

SOW signature. “Discovery: $20,000 fixed, 2–3 weeks. Build: time-and-materials at standard rates, scope and timeline to be defined at the conclusion of discovery.”

A small fixed-price discovery converts to open-ended T&M for build. The headline reads as a low-risk pilot; the conversion is the load-bearing line.

Why it predicts overrun. Discovery is a 3-week window where the partner calibrates budget tolerance, eval discipline, and willingness to push back. Two failure modes follow. First, scope inflation: the PRD lists more than the budget supports; cost surfaces gradually under T&M; by week 10 a $20K discovery has become a $180K build. Second, eval design becomes billable inside build — the partner charging by the hour for the artifact that defines what they are being paid to deliver.

The pattern is defensible for exploratory pilots under $25K. It is the modal red flag when the buyer’s intent is to ship a feature.

What to ask. “Does the $20K discovery output a co-authored eval set with numeric thresholds, and a fixed-price or capped-T&M quote for build?” If discovery outputs only a PRD, the line is the pattern.

How to counter. “The $20K discovery is fine; the conversion to open T&M is not. Discovery needs to output a co-authored eval set with thresholds and a fixed or capped build quote. The partner who can name the build quote at the end of discovery has done the discovery work.”

What to write instead.

“Discovery: $20,000 fixed, 3 weeks. Deliverables: (1) PRD; (2) co-authored evaluation set with numeric thresholds; (3) build-phase quote — fixed total or T&M with a not-to-exceed cap — naming the model dependency, retrieval architecture, and acceptance event. Build commences only after the buyer accepts the build quote; the buyer’s right to refuse at the end of discovery is unconditional.”

See how to read an idea-to-product SOW and what to negotiate for the broader SOW structure.

Pattern 4: Weekly retainer with no scope

SOW signature. “Engagement structured as a weekly retainer of $X for access to a dedicated AI team. Scope reviewed weekly; deliverables adjusted to priorities.”

The retainer prices access, not deliverables. Scope is set by the partner’s account manager based on what they find most billable.

Why it predicts overrun. Retainer pricing is fine for post-MVP work. It is wrong for a 6–12 week MVP build because it has no acceptance event — the engagement ends when the buyer stops paying, not when the feature ships. A partner billing weekly against a non-deliverable retainer earns more by extending the engagement than by shipping fast. Eval discipline, which compresses engagement length, becomes adverse to partner revenue.

What to ask. “What deliverable ends the retainer, and what is the numeric acceptance event for it?” If the answer is “ongoing engagement” or “client discretion,” the line is the pattern.

How to counter. “Retainer pricing works for post-MVP support. For the build itself I need a named deliverable and an eval-threshold acceptance event. Happy to discuss a retainer after the build clears the rubric.”

What to write instead.

“Build phase: $X total, structured as fixed total with milestone schedule or T&M with a not-to-exceed cap, ending in a numeric eval-threshold acceptance event. Post-build (optional, separately signed): weekly retainer of $Y for model monitoring, eval re-runs, and prompt iteration, with 30-day termination notice and no minimum term.”

Pattern 5: Scope to be defined post-discovery

SOW signature. “Final scope of work to be defined at the conclusion of the discovery phase. Build-phase pricing to be agreed in good faith based on discovery findings.”

The buyer is signing a contract whose central term — what they are paying for — is “to be agreed.” Good faith does the load-bearing work; good faith is not a contract clause.

Why it predicts overrun. The buyer has committed to negotiate scope and price at a future date, by which time the partner has billed 2–3 weeks, the buyer has a sunk cost, and the partner’s discovery output is the only document on the table. Bain finds vendor conversations that defer scope to post-discovery convert to overrun engagements 3–4x more often than those that define a scope range before discovery. The pattern also fails the competent-successor test: another partner cannot tell from the SOW what was promised.

What to ask. “Can the SOW name a scope range — minimum and maximum feature set — and a price for each, before discovery begins?” If the partner cannot name a range, they are pricing access to their judgment, not delivery of a feature.

How to counter. “Scope post-discovery does not work. I will sign a scope range — minimum feature set with a fixed quote, stretch feature set with a fixed quote, and a published change-order rate between. Discovery selects the point inside the range; discovery does not invent the range.”

What to write instead.

“Scope range defined pre-engagement: minimum feature set $X for $Y, full feature set $A for $B, change orders at $Z/hour. Discovery selects the build point inside the range and refines the eval set; discovery does not redefine the range. Buyer may terminate at end of discovery with fees paid in full and no further commitment.”

See the fixed-price AI MVP contract: 7 clauses worth negotiating for the protecting clauses.

A 5-line refusal worksheet

Walk the SOW against the table. Refuse the line, not the partner.

Pattern SOW signature What to refuse What to replace with
1. Lump-sum, no milestones “40/30/30, final on delivery” Tranches tied to a calendar, not a rubric Milestone schedule with eval-threshold tranche 3
2. Per-hour, no cap “T&M, non-binding estimate” No not-to-exceed cap Capped T&M with eval-threshold acceptance
3. $20K + T&M “Discovery fixed, build T&M” Build quote deferred to post-discovery Discovery outputs build quote and eval set
4. Weekly retainer, no scope “Dedicated team, weekly retainer” No build deliverable, no acceptance event Fixed or milestone build; retainer only post-launch
5. Scope post-discovery “Scope agreed in good faith later” Scope and price both deferred Scope range defined pre-engagement

Every replacement preserves the partner’s commercial logic — total fee, cashflow, team utilization — while adding an acceptance event the buyer controls. Partners selling the modal 2026 contract accept the replacements without renegotiating the number; partners selling something else push back, and the push-back is the next data point.

What to do this week

If you are signing in the next two weeks, do three things.

One: walk the SOW against the 5-pattern table. Mark each pattern. Three or more, the SOW is structurally a different product than what was pitched.

Two: send the partner the replacement clauses, one per pattern, in a single email. Frame it as “tightening the SOW.” A partner with real AI build maturity returns a redline. A partner without it defends the original language.

Three: book the next conversation around the redline. If it does not arrive in five business days, the engagement is calendar-floating before it has begun.

The cost of refusing five lines is one redline conversation. The cost of not refusing is the gap between a $150K signed scope and a $250K final invoice. Refuse the line, not the partner.

FAQ

Are all five patterns always disqualifying?

No. Lump-sum and retainer are defensible in narrow conditions — well-trodden internal-only pilots under $50K, or post-MVP support. The patterns are refusal-grade in a customer-facing $80K–$250K AI MVP build with a new partner.

What if the partner refuses to redline any of these?

It is a signal. A partner unwilling to add an eval-threshold acceptance event, a not-to-exceed cap, or a scope range is selling pricing certainty for their side, not yours. Ask “what category of work are you optimized for” — the answer tells you whether your engagement is a match.

Do these patterns apply to T&M-only freelancers?

Patterns 2 and 5 apply. Patterns 1 and 4 usually do not — solo freelancers rarely structure lump-sum or retainer engagements. Focus on the cap and the scope range.

What is the right not-to-exceed cap for a T&M AI MVP?

100% of the estimate, not 150% or 200%. A 150% cap is permission to overrun by half. The eval-threshold acceptance event makes a 100% cap workable — the partner’s incentive becomes clearing the rubric inside the cap.

How do I write an eval-threshold acceptance event without an eval set in hand?

The SOW names the process: the eval set is co-authored in week 1 or 2, the threshold is set at sign-off, and that threshold becomes the acceptance event. The buyer needs no eval set at signing — only contract language that one will exist and gate payment.

What happens if the build misses the eval threshold?

A well-formed SOW includes a cure path: one cure attempt inside the original fee, up to 2 calendar weeks, with the eval re-run before the tranche releases. Subsequent attempts run at the published change-order rate.

How do I tell a non-binding estimate from a real not-to-exceed cap?

The estimate is in the proposal; the cap is in the SOW. If the SOW says “billable at $X/hour against time logged,” the engagement is uncapped. A real cap reads as a numbered clause: “Total fees shall not exceed $Y without a written change order signed by the buyer.”

If I refuse all five patterns, am I left with no partners willing to sign?

The modal 2026 AI MVP partner accepts the replacement clauses without renegotiating the headline number. Partners optimized for fixed-price-only or T&M-only structures may decline — those partners are filtered, not lost. The survivors are the ones whose pricing tolerates buyer acceptance events, the structural requirement of an AI build.

Last Updated: Jul 24, 2026

AW

Arthur Wandzel

SFAI Labs helps companies build AI-powered products that work. We focus on practical solutions, not hype.

See how companies like yours are using AI

  • AI strategy aligned to business outcomes
  • From proof-of-concept to production in weeks
  • Trusted by enterprise teams across industries
Get in Touch →
No commitment · Free consultation

Related articles