Home About Who We Are Team Services Startups Businesses Enterprise Case Studies Blog Guides Contact Connect with Us
Back to Guides
Enterprise Software 14 min read

The Idea Risk Stack: 5 risks every AI MVP carries (and which to retire first)

The Idea Risk Stack: 5 risks every AI MVP carries (and which to retire first)

Most AI MVPs ship the wrong product because the founder retired the wrong risk first. A survey green-lights user-value, eight weeks of build go in, and the capability layer turns out untested. Or feasibility passes, the demo wows, and in week ten inference cost per task makes the business unworkable. Each failure looks different at autopsy. The root cause is the same — the team treated risks as a parallel checklist when AI MVP risks form a stack with a correct retirement order.

There are exactly five, in this order: capability, user-value, cost, moat, operational. Out of order, the next layer’s signal is corrupted. In order, most 2026 AI MVP failure modes go away. This is the risk stack inside the idea validation playbook, a chapter of the broader idea-to-product manifesto.

The 5-risk stack at a glance

Every AI MVP carries the same five risks. The table is the whole framework. The rest is justification.

#RiskQuestion it answersRetirement procedureTimeArtifact
1CapabilityCan the frontier model do this well enough?Feasibility probe — 10 inputs, 3 chat consoles, a 3-tier rubric4 hoursOne-page feasibility memo
2User-valueWill users do the work to use this and keep coming back?Thin-slice prototype + observed sessions with 5 real users2 weeksSession-recording notes + retention signal
3CostIs inference cost per task less than what users will pay?Cost-per-task probe — measure tokens on the 10 feasibility inputs1 dayUnit-economics one-pager
4MoatWill this still be a business after the next frontier model release?Moat memo — what we have that the model doesn’t1 dayTwo-paragraph moat thesis
5OperationalCan a small team run this in production without it silently degrading?Observability dry-run — eval suite seed + drift plan + injection test1 weekEval suite (50 cases) + runbook outline

Total elapsed time, sequential: about four weeks. Total founder time, parallel-staffed: about eight working days. None of these layers needs a build sprint. Each procedure is intentionally cheap — the value of the stack is the order, not any single layer.

Why the order matters

Why not retire risks in parallel? Because each layer’s signal depends on the layer above being clean.

  • User-value scores are meaningless if capability is unproven. A demo that wows users is a Wizard-of-Oz if the model fails on hard inputs. Demo excitement is not user-value evidence.
  • Cost numbers are meaningless if the task isn’t defined. Cost-per-task requires the input-output contract from the feasibility test.
  • Moat is unscoreable without a working product shape. A memo against an untested idea is fan fiction. Against a 7/10-useful test, it reads like strategy.
  • Operational risk depends on the prior four. Eval suite design needs the task, the user, the cost ceiling, and the moat thesis as inputs.

The order is forced by signal dependency, not preference. Skip it and you trade four weeks of stack-walking for twelve weeks of build that ships into the wrong assumption.

The most common 2026 mistake: retiring user-value first via a landing page or a Mom Test interview. Anything labeled “AI-powered” gets fake positive interest. Founders read it as demand and burn a build sprint on capability they never tested. Capability must come first, and the feasibility check is what retires it.

Layer 1 — Capability risk

Definition. Does the frontier model — at its current floor, with a reasonable first prompt — do this task at the quality bar the product needs?

Why it goes first. Capability is the only risk that can fully kill the idea today. The other four are gradients you can move. Capability is a yes-or-no on the current model floor; if the answer is no, every downstream analysis is wishful thinking.

Retirement procedure. A four-hour senior-engineer protocol: name the capability primitive (extraction, classification, generation, planning, multi-turn, tool-use, multi-modal); build a 10-input test set covering easy/medium/hard; write a 3-tier rubric (useful / recoverable / dangerous); run against Claude Opus 4.8, GPT-5, and Gemini 2.5 Pro; score; write a one-page memo. Full procedure in the AI feasibility check spoke.

Evidence of retirement. A signed one-page memo with idea, capability, test set, score table, decision; a useful score of at least 7/10 with at most 1/10 dangerous; a documented narrowing if 7/10 was not reached.

What is left dangerous if skipped. A four-week build sprint that ends with a 5/10 ceiling and dangerous failures in production. Median cost: one full build sprint, roughly fifty to one hundred thousand dollars depending on team rate.

Layer 2 — User-value risk

Definition. Will users do the work to use this product, return to it, and pay for it — given that capability is settled?

Why it goes second. A thin-slice prototype (one working flow, ten percent of the eventual product) can be built in two weeks using Cursor or Claude Code by a non-engineer founder. Capability without value is a working demo nobody uses.

Retirement procedure. Build the narrowest thin-slice that exercises the validated capability. Recruit five real users in the target segment — not your network, not “AI-curious” friends. Watch them use it (screen-recorded sessions are fine). Score three criteria: did they complete the task without prompting; did they come back unprompted within seven days; did they say “I would pay for this” without being asked. Two of three on a majority of the five is a retirement.

Evidence of retirement. Five session recordings with timestamped notes; a one-week retention number; at least two unprompted payment-intent statements tied to a specific dollar figure the user offered.

What is left dangerous if skipped. A polished product nobody uses. Recoverable only by pivot — the most expensive recovery in startup operations.

The trap: retiring user-value with a survey, a waitlist, or a landing page. Anything mentioning AI in 2026 gets a misleading positive signal because users say yes to AI surveys at a rate disconnected from real usage. The sibling pre-prompt validation guide covers the demand half of this layer in detail.

Layer 3 — Cost risk

Definition. Is the inference cost per task lower than what users will pay per task, with enough margin to absorb error and overhead?

Why it goes third. Cost depends on a defined task (layer 1) and a willingness-to-pay number (layer 2). Without both, you measure something abstract.

Retirement procedure. A one-day probe. Run the ten feasibility inputs through the chosen model with the prototype’s prompt design. Read token counts off the API response. Multiply by current per-token pricing. Add a 1.5x–2x overhead factor for retries, eval traffic, and prompt-engineering room. Inference cost should be no more than ten to twenty percent of the user’s price.

Evidence of retirement. A unit-economics one-pager (cost per task, willingness-to-pay, gross margin, per-user monthly ceiling); sensitivity on three model choices; a note on what doubles if average input length doubles.

What is left dangerous if skipped. A product that ships and silently loses money on every transaction. Recovery options are all unattractive — raise prices, switch below the capability bar, or cap usage. The editorial economics manifesto argues 2026 AI products must be budgeted in per-query cost, not feature cost. The metric that matters is not what the model costs today, but what the most expensive seven percent of your traffic costs today. Tail-input cost kills margins.

Layer 4 — Moat risk

Definition. When the next frontier release ships — every six to nine months — does this product still have a reason to exist, or does the new model subsume the value?

Why it goes fourth. Moat analysis requires a working product shape. A memo against an untested idea is speculation; against a validated three-layer idea, it reads like strategy.

Retirement procedure. Two paragraphs. Paragraph one: what the user gets from this product that the model alone does not give them. Allowed answers: proprietary data, a workflow the model cannot orchestrate alone, a regulatory or trust position, distribution, or a UX/latency floor unreachable through chat. Disallowed answer: “better prompting”. Paragraph two: simulate the next release — assume thirty percent better capability and one new built-in tool; does the product still exist? If the answer is “the user would switch to the new model directly”, the moat is unretired.

Evidence of retirement. A two-paragraph moat memo with the next-release simulation; an identified asset that compounds with usage; a timeline on durability without active defense.

What is left dangerous if skipped. Wrapper risk. The product gets traction, the next release ships the same capability built in, and users churn in a quarter. The pattern wiped out half the GPT-3.5-era wrapper cohort at the GPT-4 transition and a second cohort at GPT-5. It has repeated through three frontier transitions and will repeat through the next.

Layer 5 — Operational risk

Definition. Can the team run this in production for six months without silent degradation, drift, or adversarial-input failures?

Why it goes last. Eval suite, drift plan, and prompt-injection probe all require the prior four outputs as inputs.

Retirement procedure. A one-week observability dry-run with three deliverables: an eval suite seed (expand the 10 feasibility inputs to 50, same three-tier rubric, monthly re-score cadence); a drift plan documenting what changes if the model upgrades or the input distribution shifts, with the tip-off metric named; and a five-minute prompt-injection probe with the result logged.

Evidence of retirement. A 50-case eval suite committed to a repo with a numeric baseline; a one-page drift-and-incident runbook outline; a logged injection-probe result with mitigations.

What is left dangerous if skipped. A system that ships, then silently degrades as the input distribution drifts, eventually producing a public failure — a hallucinated customer output, leaked instructions, or a regulatory incident. Recovery costs include reputation, customer trust, and often a re-architecture.

The 1-page risk worksheet

A single-page worksheet ties the five layers together. Five rows, four columns.

RiskStatusRetirement planEvidence
CapabilityRetired / In progress / OpenWhat you did or will doArtifact link
User-valueRetired / In progress / OpenWhat you did or will doArtifact link
CostRetired / In progress / OpenWhat you did or will doArtifact link
MoatRetired / In progress / OpenWhat you did or will doArtifact link
OperationalRetired / In progress / OpenWhat you did or will doArtifact link

The worksheet is the bridge between idea and engineering scope. A founder who hands a completed version to an AI build partner gets a faster, cheaper, more accurate engagement. The long-form version, with prompts and ready-to-fill artifacts, is the AI MVP Scoping Worksheet.

A worked sheet for a contracts-extraction product two weeks in might read: capability retired (8/10 useful, one non-English failure); user-value in progress (three of five sessions unprompted, two payment-intent statements); cost retired (twelve cents per contract against a one-to-two-dollar price); moat open (workflow integrations next quarter); operational open (eval at 20 of 50). Two layers from a defensible launch, time-bounded. Compare to a founder who ran a landing page first: 1,200 signups, no idea if the system works.

Three common mis-orderings (and what they cost)

  • User-value first via survey. Cost: the entire build sprint, because capability is discovered late. The 2026 default-instinct failure. Run the four-hour capability probe before you write landing-page copy.
  • Cost first because the founder is finance-trained. Cost: a defensible spreadsheet for a product nobody wants to use. Pristine unit economics, flat retention. Gate cost analysis behind user-value retirement.
  • Moat first because the VC asked. Cost: a moat memo against an unvalidated idea, which the founder then rationalizes around for six months. Hold moat work until capability and user-value are retired.

The discipline of the stack is in the order, not in any one layer.

FAQ

What is the AI MVP risk stack?

A five-layer framework for sequencing the risks every AI product idea carries: capability, user-value, cost, moat, operational. The risks are not a parallel checklist — they form a stack, and retiring them in the wrong order corrupts the signal from later layers.

What is the correct order to retire AI MVP risks?

Capability first, then user-value, cost, moat, operational. The order is forced by signal dependency — each layer’s evidence requires the prior layer’s output to be meaningful.

Why is capability retired before user-value?

User-value tests run on an unproven capability layer produce false positives. A demo that wows users is a Wizard-of-Oz if the model fails on hard inputs. The 2026 AI-labeled survey-positive signal is also untrustworthy, compounding the problem.

How long does it take to walk the full stack?

Sequentially, four weeks elapsed and eight to ten working days of founder time. Parallel staffing compresses the timeline. Each retirement procedure runs in 4 hours to 1 week — no build sprint required.

What is the artifact that ties the stack together?

A single-page worksheet with five rows and four columns: status, retirement plan, evidence link. It is the deliverable a founder hands to a build partner, investor, or themselves before a build sprint. The long-form version is the AI MVP Scoping Worksheet.

What is the most expensive risk to skip?

Capability. Skipping it costs a full build sprint — fifty to one hundred thousand dollars plus four to twelve weeks of opportunity cost. Other layers retire at lower cost and their failure modes develop more slowly.

How is the stack different from desirability-feasibility-viability?

The IDEO triangle treats three risks as parallel and equal. The stack orders them with a forced sequence and adds two AI-specific risks the triangle misses: moat under frontier-upgrade pressure, and operational risk in LLM-app conditions. More predictive for AI products in 2026.

Does it apply to internal AI tools?

Yes, with two adjustments. User-value becomes adoption risk. Cost becomes total-cost-of-ownership. The other three layers transfer unchanged.

When should I revisit the stack?

Re-score capability after every frontier model release. Re-score cost quarterly. Re-score moat after any significant competitive event. User-value and operational re-score continuously through normal operations.

What if my idea passes the stack but I still don’t ship?

The stack retires technical and economic risk, not execution risk. A clean stack is the precondition for execution, not a substitute. If the project still isn’t shipping, the gap is execution — the idea-to-product manifesto is the right next read.

Next step

The stack is the framework; the worksheet is the artifact. The handoff is to whichever build path comes next — solo with Cursor and Claude Code, with a build partner, or by hiring. The AI MVP Scoping Worksheet is the printable long-form version. Founders who arrive at the partnership conversation with the worksheet completed close engagements faster, cheaper, and with less rework than founders who arrive with only a hunch.

Last Updated: Jul 2, 2026

AW

Arthur Wandzel

SFAI Labs helps companies build AI-powered products that work. We focus on practical solutions, not hype.

See how companies like yours are using AI

  • AI strategy aligned to business outcomes
  • From proof-of-concept to production in weeks
  • Trusted by enterprise teams across industries
Get in Touch →
No commitment · Free consultation

Related articles