Home About Who We Are Team Services Startups Businesses Enterprise Case Studies Blog Guides Contact Connect with Us
Back to Guides
Enterprise Software 28 min read

The AI MVP Economics Playbook: what 6–12 weeks of build actually costs

The AI MVP Economics Playbook: what 6–12 weeks of build actually costs

An AI MVP is priced wrong because buyers ask the wrong question. They ask “what does it cost?” The defensible question is “what is the smallest eval set that proves the idea — and what does it cost to clear that set inside 6–12 weeks?” The eval set drives every cost line. The planning week, the PRD, the architecture decision, the model picks, the eval-engineering hours, the integration surface, the on-call buffer, the handoff documentation — every line traces back to a fixed sample of representative inputs and a written rubric for what correct looks like on each. Without that eval set as the budget anchor, the founder is buying engineer-hours at a vendor’s preferred multiplier and hoping the result feels like a product. With it, every dollar has a defensible upstream and a defensible downstream. This playbook is how a non-engineer founder builds that budget.

This article is the cluster anchor for AI MVP build economics and timelines within the broader idea-to-product manifesto. It sits alongside the idea validation playbook and the eval-first build playbook — those two anchors define what the founder is buying; this one defines what it should cost. The next two — the DIY with AI manifesto and the founder-AI partner operating manual — define how a founder runs the engagement once the budget is signed.

Table of Contents

Why AI MVP Economics Inverts the 2018 SaaS Playbook

A 2018 SaaS MVP could be priced on engineer-hours because the work was deterministic. A login form was a login form. A row in a database was a row in a database. The budget formula was a sum: number of screens × hours-per-screen × hourly rate + a 20% buffer for QA. The 2026 AI MVP cannot be priced that way because the engineering work has shifted shape. Three structural reasons explain why an AI MVP of comparable feature scope ranges 30–80% more expensive than a 2018-style web MVP — and why pricing it on engineer-hours alone produces a number that has nothing to do with whether the product will ship at quality.

First, the dominant cost line is no longer code, it is evals. In a 2018 SaaS build, evaluation was 5–10% of the budget — a small QA pass against acceptance criteria. In a 2026 AI build, evaluation is 20–35% of the budget, and not because QA got more expensive. It is because the artifact that says “this feature works” is no longer a green CI check on a deterministic test — it is a graded sample against a representative test set, with a written rubric, anchored sub-criteria, and a regression gate that fires when a model alias updates underneath the product. Building that artifact takes evaluation-engineering hours that did not exist as a line item five years ago. The eval-first build playbook is the technical companion to this economics playbook; together, they say the same thing in different registers — the eval set is the spec, and the eval set is the budget.

Second, the inference and observability surface is recurring, not one-shot. A 2018 SaaS MVP shipped against a fixed hosting bill. A 2026 AI MVP ships against an inference bill that scales with usage, an observability bill that scales with trace volume, and a model-migration overhead that fires every time a frontier model alias updates. Frontier model behavior changes on cycles of weeks — Anthropic shipped multiple Claude Opus and Sonnet variants over the past year; OpenAI shipped GPT-5 successor updates; Google shipped Gemini 2.5 family models. Every transition risks silently changing the behavior of any system pinned to a model alias, which means the eval suite has to be re-run and the prompt scaffolding may need retuning (Artificial Analysis LLM Leaderboard). The MVP budget has to fund the first model-migration cycle inside the 6–12 week window, or the founder discovers the cost in production.

Third, the build path itself branches more sharply than a web MVP build path branched in 2018. A prompt-only feature, a retrieval-augmented feature, an agentic feature, and a fine-tuned-model feature are four genuinely different build paths with three-to-five-fold cost ranges between them. A web MVP build path in 2018 was, broadly, the same path: React frontend, Node or Rails backend, Postgres. The architecture decision sat outside the cost conversation. The 2026 AI MVP architecture decision is the cost conversation — pick wrong on stage 3 and the rest of the budget is set against a build path the team will silently outgrow by week 5.

These three shifts together explain a finding that recurs in industry reports: roughly 80–85% of AI pilots stall before reaching production at scale, a figure McKinsey reports across successive State of AI editions and that Gartner has corroborated through its 2024 and 2025 CIO surveys. The stalled pilots are not failing because the engineers were bad — they are failing because the budget was built on a 2018-shape cost frame that did not fund the lines (evals, observability, model-migration buffer, eval-engineering hours) that the 2026 build actually requires. The playbook below is how a non-engineer founder funds those lines instead of discovering them.

The Nine-Stage Economics Playbook

The playbook below decomposes a 6–12 week AI MVP into nine cost-generating stages. Each stage has a name, the artifact it produces, the line item it generates with a defensible 2026 market range, the milestone that triggers payment, and the cost of skipping or under-funding it. Run them in order. A founder reading this should be able to put each stage’s range next to the corresponding line in any vendor quote and ask “where is this number in your proposal, and what would skipping it cost me downstream?”

Stage 1. ICP and Scope Intake

The first stage is to name, in writing, which customer the MVP is for, what correct looks like for that customer, and which capability is the minimum surface that proves the idea. Most founders arrive at a vendor with a feature pitch (“an inbox triage agent”) and a customer pitch (“for revenue operators at Series B SaaS companies”). The intake stage compresses both into one artifact: a one-page customer-and-capability brief that names the ideal customer profile, the single workflow the MVP automates, the three-to-five capabilities required to automate it, the success metric, and the kill criterion.

This stage takes 12–25 hours of focused work between the founder and one senior partner. The total cost depends on whether the partner is an internal cofounder, a fractional senior engineer, or a billed agency hour. In 2026 market terms, a vendor-led intake week runs $5K–$12K depending on the partner’s day rate and the number of customer interviews funded inside the week.

  • Artifact: a one-page ICP-and-capability brief.
  • Line item: planning week / discovery.
  • Defensible 2026 range: $5K–$12K (vendor-led); $0 if founder-led using a partner skill they already pay for.
  • Milestone trigger: signed brief, both founder and partner co-author.
  • Cost of skipping: the budget downstream gets set against a feature pitch rather than a capability, and the eval set in stage 5 has nothing to anchor against. The result is a $60K–$120K build with no quality bar — the most common pattern behind the 80% pilot stall rate.

Stage 2. Eval-Bound PRD Engineering

Once the ICP and capabilities are named, the next stage is to write the PRD that turns them into something a senior engineer can estimate. A 2026 AI PRD is structurally different from a 2018 SaaS PRD — it is a workflow plus an eval contract, not a feature list. The idea validation playbook describes the nine-stage validation sequence that produces this document; this economics playbook treats the PRD as the artifact whose price is being defended.

PRD engineering takes 25–55 hours of paired founder-engineer work, distributed over 4–7 calendar days. The product of those hours is a document that, in addition to the standard PRD sections (problem, users, scope, non-goals, acceptance criteria), names the eval contract — the test set the feature will be graded against, the rubric for each capability, and the threshold the build has to clear. In 2026 market terms, a vendor-engineered eval-bound PRD ranges $12K–$22K depending on the depth of customer interviews funded and the number of capabilities the PRD covers.

  • Artifact: a 6–10 page eval-bound PRD with a separate eval-contract appendix.
  • Line item: PRD engineering.
  • Defensible 2026 range: $12K–$22K.
  • Milestone trigger: founder and partner co-sign the PRD; the engineer signs the eval contract.
  • Cost of skipping: the team builds against a feature list, ships a demo that runs on hand-picked inputs, and discovers in production that the quality bar nobody wrote down was higher than the one the system clears. The disputes that follow are scope disputes — political and unwinnable.

Stage 3. Architecture Decision

The architecture decision is the largest single cost-determining call in the playbook. The four common build paths for an AI MVP in 2026 are:

  1. Prompt-engineered — a single frontier model with sophisticated prompting, optionally with structured output. Cheapest path. Works for classification, extraction, and constrained generation. Typical infra footprint: model API only.
  2. Retrieval-augmented (RAG) — frontier model + vector store + retrieval pipeline. Required when the feature has to ground on a corpus the model was not trained on. Adds a vector database, an indexing pipeline, and a chunking strategy.
  3. Agentic — frontier model + tool definitions + an orchestration loop that decides when to call which tool. Required for multi-step workflows with branching decision logic. Adds a tool-calling harness, trajectory evals, and a safety / guardrail layer.
  4. Fine-tuned / distilled — a smaller model trained on labeled data from a frontier-model baseline. Rarely the right call for an MVP (the labeled data does not yet exist), but sometimes the right call inside a 6–12 week window if the inference cost curve at projected scale makes the frontier-model path uneconomical.

Picking the right path is a 6–16 hour exercise. The wrong call costs 2–4 weeks of rework. In 2026 vendor terms, a documented architecture-decision artifact (with a written rationale, a cost projection at projected scale, and a one-year migration option) runs $4K–$9K.

  • Artifact: a 2–3 page architecture decision record (ADR).
  • Line item: architecture & technology selection.
  • Defensible 2026 range: $4K–$9K.
  • Milestone trigger: signed ADR with cost projection.
  • Cost of skipping: the team picks an architecture by default (whatever the senior engineer is most comfortable with) and discovers in week 5 that the path cannot scale, cannot ground, or cannot pass the eval set. Re-architecture mid-build is 15–30% of the total budget in lost time.

Stage 4. Infra Picks — Model, Hosting, Observability

Once the architecture is picked, the infrastructure stack is picked against it. The 2026 stack for an AI MVP typically includes:

  • A frontier model (Claude Opus 4.6, Claude Sonnet 4.6, GPT-5 family, or Gemini 3.1 Pro), priced per million input and output tokens. Anthropic, OpenAI, and Google all publish current pricing on their docs pages; per-million-token costs for the top frontier models range from a few dollars (smaller / faster variants) to tens of dollars (top reasoning variants) per million input tokens.
  • A vector store (if RAG): Pinecone, Weaviate, Qdrant, or pgvector. Hosted vector stores run roughly $70–$700/month at MVP scale depending on document count and query volume.
  • An observability layer: Langfuse, Helicone, or Braintrust. Open-source self-hosted is free in dollar terms but consumes engineer hours; hosted plans run roughly $50–$500/month at MVP scale.
  • A backend host: a generic cloud provider (Vercel, Render, Fly.io, AWS) priced against compute and bandwidth. MVP-scale runs roughly $100–$400/month.

Adding these together, an MVP-scale operating bill runs $3K–$8K/month in 2026 — composed of inference, vector store, observability, and hosting. This is a recurring line, not a one-shot. The MVP budget has to fund 2–4 months of this operating cost (the build months plus the on-call month), which is $6K–$32K depending on the architecture and the projected usage.

  • Artifact: an infrastructure-cost worksheet projecting monthly burn at MVP scale and at projected production scale.
  • Line item: infrastructure & operating cost (recurring, 2–4 months funded inside the MVP budget).
  • Defensible 2026 range: $3K–$8K per month; $6K–$32K total funded inside the MVP window.
  • Milestone trigger: signed worksheet with named vendor accounts.
  • Cost of skipping: the founder is surprised in month 2 by a $4K/month inference bill they did not budget. Cost-per-query economics — the discipline of pricing the unit of output against the unit of revenue — gets short-circuited and the unit economics do not survive the first 100 paying customers. The companion piece on cost-per-query as a defensible unit-economics framework names this discipline at the post-MVP scale.

Stage 5. Eval Engineering Hours

Eval engineering is the largest single line item in a 2026 AI MVP budget — and the one most often missing from competing vendor proposals. The eval engineer builds the test set, writes the rubric in a machine-runnable form, picks the harness (Promptfoo, Inspect AI, Langfuse, or DeepEval), wires it into CI, and runs the baseline. For an MVP with 3–5 capabilities, this is 60–120 hours of senior engineering work over 2–3 weeks.

In 2026 market terms, a vendor-engineered eval suite at MVP scale runs $25K–$45K. This is not the same as QA — QA is part of stage 7. Eval engineering is the discipline of turning the rubric in the PRD into a runnable program that grades the build on every commit and on every model alias change. The output is the artifact the founder will use to grade the team for the rest of the engagement.

  • Artifact: a runnable eval suite covering all capabilities in the PRD, plus a refusal sub-suite of 10–30 inputs per capability.
  • Line item: eval engineering.
  • Defensible 2026 range: $25K–$45K.
  • Milestone trigger: baseline passes on a sample model; founder spot-checks 20 graded outputs and agrees with the grader.
  • Cost of skipping: the team has no way to detect a regression caused by a model update, a prompt change, or an upstream behavior shift. Every week-9 dispute about quality is unresolvable because there is no rubric to point at. This is the line whose absence most reliably converts a 6–12 week build into a 16–24 week build.

Stage 6. Integration Surface

The integration surface is the set of upstream and downstream systems the MVP has to talk to — the inbox the agent triages, the CRM the system writes to, the auth provider the user signs in through, the data warehouse the analytics flow into. Every integration adds a defined surface area, a contract to honor, and a failure mode to handle. For an MVP with 2–3 integrations (one inbound, one outbound, one auth), this is 30–80 hours of engineering work.

In 2026 market terms, a vendor-engineered integration layer at MVP scale runs $8K–$20K. Each additional integration past the third adds roughly $3K–$8K. Founders consistently underestimate this line because integrations look simple from the outside — “it’s just an API call”. They are not. The contract surface, the error handling, the auth refresh logic, and the rate-limit backoff are 60–80% of the line.

  • Artifact: a working integration layer with named adapters, contract documentation, and failure-mode handling.
  • Line item: integrations.
  • Defensible 2026 range: $8K–$20K for 2–3 integrations; +$3K–$8K per additional.
  • Milestone trigger: each integration passes its contract tests; auth refresh works end-to-end.
  • Cost of skipping: nothing. This line cannot be skipped. But it can be under-scoped — and an under-scoped integration line is the most common source of week-7 schedule slip.

Stage 7. Build and Iterate

The actual build — prompt scaffolding, retrieval pipeline, agent loop, UI, API surface — is what most founders think the budget is about. In a defensible 2026 budget, it is roughly 25–35% of the total, not the whole pie. For a 6–12 week MVP, build-and-iterate runs 2–4 weeks of focused engineering after the PRD, ADR, and eval suite are in place.

In 2026 market terms, a vendor-engineered build-and-iterate phase runs $20K–$45K depending on the architecture. The discipline inside this stage is to iterate against the eval suite on every prompt change, every retrieval-strategy change, every agent-loop refactor — and to refuse to ship a change that drops eval performance below the threshold. The team’s velocity inside this stage is set by how good the eval suite is. A team with a strong eval suite ships 3–5 substantive iterations per week with confidence. A team without one ships once a week and prays.

  • Artifact: a working MVP that clears the eval threshold on every capability.
  • Line item: build and iterate.
  • Defensible 2026 range: $20K–$45K.
  • Milestone trigger: eval suite passes at the agreed threshold; UI demo runs end-to-end.
  • Cost of skipping: nothing — but skipping the eval-gate discipline inside this stage is what produces a demo that passes review but fails in production. The stop budgeting AI projects in story points, budget them in eval runs thesis is the rhetorical version of this discipline.

Stage 8. On-Call Period

The on-call period is the 2–4 weeks after the MVP launches when the team is on standby to handle production issues, customer feedback, regressions, and the first round of prompt-tuning against real usage. Most founders forget to budget this line, which is why most AI MVPs feel “done” at week 8 and then quietly burn another $15K–$30K of unbudgeted senior-engineer time between week 8 and week 12.

In 2026 market terms, a defensible on-call line is $8K–$20K for a 3–4 week period at a reduced engagement (10–20 engineer hours per week). The eval suite is the artifact that keeps the on-call line bounded — without it, on-call becomes “the engineer eyeballs production traces”; with it, on-call becomes “the engineer fixes the inputs the suite flagged”.

  • Artifact: a weekly production-trace replay report plus a backlog of prompt / retrieval / rubric refinements informed by real usage.
  • Line item: on-call buffer.
  • Defensible 2026 range: $8K–$20K.
  • Milestone trigger: each weekly trace-replay report delivered; agreed refinement backlog accepted.
  • Cost of skipping: the founder discovers in week 10 that “done” is not done — the system is regressing under real usage, the team has rotated to the next project, and the founder is stuck with a system they cannot operate.

Stage 9. Handoff and Runbook

The handoff is the artifact the founder is left with when the engagement closes. A defensible 2026 handoff includes a runbook (how to operate the system day-to-day), a model-migration plan (what to do when a frontier model alias updates), a cost dashboard (what the unit economics look like at current scale), an eval-suite refresh plan (how to add new capabilities and how to re-baseline), and a 3–6 month roadmap for the next iteration.

In 2026 market terms, a vendor-engineered handoff runs $5K–$10K. It is the cheapest line in the budget and the one most often under-delivered. The founder reading this should refuse to sign the final invoice until each artifact named above is in their hand.

  • Artifact: a 10–20 page handoff document plus a recorded walkthrough.
  • Line item: handoff & runbook.
  • Defensible 2026 range: $5K–$10K.
  • Milestone trigger: founder reviews each artifact and signs off.
  • Cost of skipping: the founder owns a black box. Every future change requires re-hiring the original team, which is the most expensive form of vendor lock-in in 2026.

A Worked Example: An $80K AI MVP, Line by Line

The following is a defensible 2026 budget for an $80K AI MVP — a typical idea-to-product engagement for a non-engineer founder shipping a single-workflow AI feature inside a 6–12 week window. The numbers are presented as defensible 2026 market ranges rather than a single point estimate, because the actual figure inside the range depends on the architecture, the integration count, and the partner’s day rate. Every line names a deliverable.

#Line itemDefensible 2026 rangeStage
1Planning week / ICP and scope intake$6K–$10KStage 1
2Eval-bound PRD engineering$14K–$18KStage 2
3Architecture decision record$5K–$7KStage 3
4Infrastructure operating cost (3 months funded)$10K–$18KStage 4
5Eval engineering$28K–$38KStage 5
6Integrations (2 inbound, 1 outbound)$10K–$15KStage 6
7Build and iterate (3–4 weeks)$22K–$32KStage 7
8Security and privacy review$3K–$6KStage 7
9Observability and dashboard setup$3K–$5KStage 4
10On-call buffer (3 weeks)$10K–$15KStage 8
11Handoff and runbook$6K–$8KStage 9
TotalRange$117K–$172K
Founder-led compressionStages 1–3 founder-run; Stages 7 build de-scoped to 2 weeksPulls the range to $70K–$95K

The compressed $70K–$95K range is the one most founders should hold against vendor quotes. The full $117K–$172K range is the one that prices in a hands-off engagement where the partner runs every stage end-to-end. The $80K target is achievable when the founder is willing to (a) co-author the PRD rather than fully delegate it, (b) accept a 2-week rather than 4-week build phase, (c) accept 2 integrations rather than 3, and (d) operate the system themselves during the on-call period with a 1-day-per-week engineer check-in.

What the $80K does not buy: a polished consumer-grade UI (the MVP UI is functional, not designed); a fine-tuned model (frontier-model prompting + RAG only); more than 3 capabilities; more than 3 integrations; production-scale load testing (the system is sized for early-customer scale, not for 10,000 concurrent users); a full enterprise security review (a startup-appropriate review only). The companion piece inside the AI project budget — what $250K actually buys in 2026 walks the upper-end engagement; the AI project economics manifesto walks the philosophical inversion that makes both budgets defensible.

For a deeper breakdown of any single line, see the founder’s MVP cost worksheet — 11 line items or the anatomy of a $75K AI MVP — where the money actually goes.

How the 6–12 Week Timeline Maps to Spend

The 9-stage playbook compresses to a 6–12 week calendar in roughly the following shape. A founder reading a vendor proposal should be able to ask “which week funds which stage?” and get a non-evasive answer.

WeekStages activeSpend that weekCumulative
1Stage 1 (intake), Stage 2 begins$8K–$12K$8K–$12K
2Stage 2 (PRD), Stage 3 (ADR)$12K–$16K$20K–$28K
3Stage 4 (infra), Stage 5 begins$10K–$14K$30K–$42K
4Stage 5 (eval engineering, peak)$12K–$16K$42K–$58K
5Stage 5 finish, Stage 6 (integrations)$10K–$14K$52K–$72K
6Stage 7 (build), Stage 6 continues$10K–$14K$62K–$86K
7Stage 7 (build, peak iteration)$10K–$14K$72K–$100K
8Stage 7 finish, launch$8K–$12K$80K–$112K
9–11Stage 8 (on-call)$3K–$6K/week$89K–$130K
12Stage 9 (handoff)$6K–$10K$95K–$140K

Two observations matter to the founder reading this. First, the cost curve front-loads — by week 4, roughly 50% of the budget is committed, because stages 1, 2, 3, and 5 cluster in the first month. A vendor proposal that bills evenly across 12 weeks is hiding which weeks are doing the load-bearing work. Second, weeks 9–12 are the cheapest weeks but the ones founders most often cut — and cutting them is what produces the “done but not done” pattern that makes the MVP feel like a sunk cost rather than a launchable product.

The Nine-Question Cost-Defensibility Self-Test

Before signing a vendor quote, the founder should be able to answer yes to all nine:

  1. Does the quote name the eval contract as a distinct line item, with a price the founder can hold against this playbook’s $25K–$45K range?
  2. Does the quote name an architecture decision record as a deliverable in week 1 or 2, with a written rationale and a cost projection?
  3. Does the quote include an infrastructure operating cost line that funds 2–4 months of inference, observability, vector store, and hosting?
  4. Does the quote separate integrations as a distinct line item priced per integration, with a defensible 2026 range of $3K–$8K each beyond the first?
  5. Does the quote include an on-call period of at least 2 weeks at a reduced engagement after launch?
  6. Does the quote include a handoff document with a runbook, a cost dashboard, a model-migration plan, and an eval-refresh plan as deliverables?
  7. Can the founder name, in writing, the success metric the eval suite will grade against and the threshold the build has to clear?
  8. Does the cost curve front-load in line with the 9-stage shape — roughly 50% committed by week 4, 80% by week 8, 100% by week 12?
  9. Has the founder priced what gets cut if the budget is compressed below the bottom of the playbook’s range — and is the team transparent about which stages the cut comes from?

Nine yeses means the founder is signing a defensible quote. Fewer than nine means the quote is structured against a 2018-shape cost frame and the founder is buying engineer-hours, not an AI MVP that will clear its quality bar at week 12.

Frequently Asked Questions

How much does an AI MVP cost in 2026?

A defensible 2026 AI MVP for a single-workflow feature with 2–3 capabilities and 2–3 integrations ranges $60K–$150K. The $80K worked example above is the median. The lower end ($60K–$80K) assumes a founder who co-authors the PRD, accepts a compressed 2-week build phase, and operates the on-call period themselves with light engineer support. The upper end ($120K–$150K) prices in a hands-off engagement, a polished UI, and a fuller integration surface. Budgets under $50K typically skip eval engineering — which is the single most predictive sign the MVP will not clear its quality bar.

Why does an AI MVP cost more than a web MVP of the same feature scope?

Three reasons. First, eval engineering is a new line item (20–35% of the budget) that does not exist in a web MVP. Second, the recurring infrastructure cost (inference, observability, vector store) has to be funded inside the build window — a web MVP has no equivalent. Third, the architecture decision is genuinely consequential — picking the wrong build path (prompt-engineered vs. RAG vs. agentic) costs 2–4 weeks of rework, where a web MVP architecture decision is largely commoditized.

What’s the cheapest defensible AI MVP in 2026?

Roughly $40K–$55K for a prompt-engineered, single-capability, single-integration feature where the founder co-authors the PRD, runs a compressed 1-week build phase, and self-operates the on-call period. Anything under $40K typically skips eval engineering or skips the on-call period — both produce a product that feels done but is not. The DIY with AI manifesto walks the deeper compression — the path where Cursor and Claude Code substitute for some of the engineer-hours entirely.

What is the AI MVP cost curve over month 1–12?

The 6–12 week build is roughly 60–70% of the year-one cost. Months 4–6 (post-launch hardening, second-round eval refresh, first model-migration cycle) add another 15–20%. Months 7–12 (operating cost at growing scale, occasional prompt and retrieval refinements) add another 10–15%. The total year-one cost for an $80K MVP typically lands at $115K–$145K — and the founder should budget the whole curve, not just the 6–12 week sticker price. The decoding AI project TCO — 7 cost lines most CFOs miss piece is the TCO companion.

Is the eval engineering line really $25K–$45K? Can it be lower?

It can be lower if the founder co-builds the eval suite (writing the rubric in prose, sampling representative inputs, hand-grading 20–30 baseline outputs). That offload typically pulls the line down to $15K–$25K. It cannot be zero — the harness setup, the LLM-as-judge configuration, the CI gate, and the regression suite require senior engineering hours. The founder should refuse a quote where the eval engineering line is below $12K — that quote is funding a token QA pass, not a real eval contract.

How should the vendor bill milestones?

Five milestones is the defensible 2026 pattern: signed PRD + ADR (15–20% of total), eval suite baseline (20–25%), build mid-point gate (20–25%), launch and eval-pass (20–25%), handoff and on-call closure (10–15%). Avoid time-and-materials billing — it converts the eval suite into a feature the vendor controls rather than the founder. Avoid fixed-bid against a feature list — it converts every scope conversation into a scope dispute. Milestone billing against signed artifacts is the structure that keeps incentives aligned through the build.

What changes if I want a fine-tuned model instead of frontier-model prompting?

Add $20K–$40K to the build. Fine-tuning requires labeled data (which usually does not exist at MVP stage), a labeling pipeline, a training run, an evaluation against the frontier-model baseline, and a hosting decision (most fine-tuned models cannot reuse the frontier-model provider’s hosting). Fine-tuning is rarely the right MVP call. It becomes the right call after the MVP has shipped and the inference cost curve at projected scale makes the frontier-model path uneconomical. Reserve fine-tuning for month 6 onward.

Why are AI MVP timelines 6–12 weeks, not 3–4 weeks like a web MVP?

Because the eval contract takes 2–3 weeks to engineer before the build can iterate against it, and the build then needs 3–5 weeks to clear the eval threshold across all capabilities. Compressing below 6 weeks usually means skipping the eval contract — which is the most reliable way to discover at week 5 that “done” is not done. Compressing above 12 weeks usually means the team is iterating against an eval suite that is poorly anchored — push the team to fix the suite, not to extend the calendar.

Should I expect vendor proposals to follow this 9-stage shape?

Most 2026 AI agency proposals follow a 4-line shape (discovery / design / build / launch) that hides where the eval and infrastructure costs sit. A vendor who can map their proposal onto the 9 stages above and name each line’s price is a vendor who has built AI MVPs before. A vendor who cannot is one whose pricing is based on a 2018 cost frame — and whose first eval-engineering disagreement, in week 5, will be a budget dispute the founder cannot win. The field guide to evaluating an AI agency in under 90 minutes is the companion piece on vendor selection.

How does this budget change for a regulated industry (healthcare, finance, legal)?

Add 15–30% across the board. The security and privacy review line moves from $3K–$6K to $15K–$35K. The eval engineering line adds a refusal-correctness sub-suite that grows the budget by $8K–$15K. The handoff document adds a compliance attestation. The on-call period typically extends to 6 weeks rather than 3. A regulated-industry AI MVP that targets the lower end of the unregulated $80K range is signing a quote that has been mispriced — push back.

Closing

The AI MVP economics playbook is not a price list. It is the discipline that lets a non-engineer founder put a vendor quote next to a defensible 2026 market range, ask the right questions, and walk into the engagement with the same negotiating position at week 12 that they had at week 1. The 9 stages — ICP intake, eval-bound PRD, architecture decision, infrastructure picks, eval engineering, integration surface, build-and-iterate, on-call, handoff — are the cost-generating units of the build. The $80K worked example is the median shape. The 9-question self-test is the gate the founder uses to grade a quote before signing it.

A founder who runs this playbook does not become a buyer of engineer-hours at a vendor’s preferred multiplier. They become a buyer of a defensible artifact at each stage of the engagement, priced at each stage’s defensible 2026 range, with each milestone gating the next. That is the control surface the rest of the idea-to-product manifesto is built on. The sibling cluster anchors — the idea validation playbook, the eval-first build playbook, the DIY with AI manifesto, and the founder-AI partner operating manual — each assume the budget is defensible. Without it, every downstream decision (which vendor, which architecture, which cadence) is being made against a number the founder cannot defend.

Price the eval. Fund the stages. Hold the range. Then ship.

Last Updated: Jul 17, 2026

AW

Arthur Wandzel

SFAI Labs helps companies build AI-powered products that work. We focus on practical solutions, not hype.

See how companies like yours are using AI

  • AI strategy aligned to business outcomes
  • From proof-of-concept to production in weeks
  • Trusted by enterprise teams across industries
Get in Touch →
No commitment · Free consultation

Related articles