A high-confidence AI idea is not a strong opinion held loudly. It is an idea that has eight specific properties — and the presence or absence of each predicts, with embarrassing accuracy, whether the build that follows will reach production. We have watched ideas with all eight ship in six weeks and ideas missing three of them die in week six of a twelve-week sprint. The variance is not luck. It is the anatomy.
This is a working diagnostic, part of the idea-validation playbook and the broader idea-to-product manifesto. Run it after the 5-question test and before you commit a build sprint. Eight properties, one scorecard, a binary answer. Ninety minutes from open to artifact.
The 8 properties at a glance
| # | Property | Diagnostic question |
|---|---|---|
| 1 | Single user, single task | Can you finish “X does Y for Z” without an “and”? |
| 2 | Capability inside current frontier | Has a feasibility probe scored the task 7/10 or better? |
| 3 | Existing workflow being replaced | Can you name the role and the step you are automating today? |
| 4 | User pays today for the manual version | Is there a payroll, contractor, or SaaS line that disappears when you ship? |
| 5 | Eval rubric write-able in plain English | Can you describe a correct output a non-engineer can grade? |
| 6 | Failure mode is recoverable | If the model outputs the worst plausible answer, can the user undo or override it? |
| 7 | Wedge into a larger market | Is there a 50-person beachhead today that expands to 5,000 within 24 months? |
| 8 | Founder has earned domain trust | Have you spent three or more years inside the workflow you are automating? |
Eight binary questions. All eight yes — the build reaches production, gets paying users, earns the right to a second sprint. Three or more no — the build dies before week six. The middle is where the work is.
Why these eight
Most AI-idea checklists in 2026 are repackaged Lean Startup — TAM, market timing, founder passion, demo reaction, waitlist size. Those answer demand. AI MVPs do not die of demand. They die of capability surprise, eval ambiguity, recoverability mismatch, and frontier-vendor encroachment — none of which the SaaS checklist names.
The eight below were extracted from AI-MVP failure post-mortems since 2023. Each is necessary; the set is minimal. Five are AI-specific (2, 5, 6, plus the AI-sharpening of 4 and 7). Three are generic hygiene (1, 3, 8). The combination is what makes a high-confidence idea anatomically distinct from a strong-feeling idea.
Property 1 — Single user, single task
Present. A one-sentence statement: “My AI does Y for Z users when W.” No “and.” Y is a single task; Z is a single buyer persona with a budget and a job title.
Missing. “AI for everyone who manages X” — a platform pitch wearing a product costume. Multiple users, multiple tasks, a use-case map resembling a Notion database.
Why it predicts. Single user / single task is what keeps the eval suite small enough to maintain and the UX small enough to ship in eight weeks. Each “and” doubles eval cases, UI surface, and integration scope. By the third “and,” the MVP is a platform requiring a sixty-week build.
Diagnostic question. Can you write “X does Y for Z” on a sticky note in fifteen words, without an “and”? If no, you have a category, not a product.
Property 2 — Capability inside current frontier
Present. A four-hour feasibility probe scored the task at 7/10 or better — useful, recoverable, not dangerous — on Claude Opus 4.8, GPT-5, or Gemini 2.5 Pro at default settings. Ten representative inputs, three-tier rubric, documented score.
Missing. “I tried a few prompts in ChatGPT and it seemed to work.” A demo on three easy inputs. No documented rubric. The founder believes capability is there, but no one has stress-tested the hard cases.
Why it predicts. Capability is the only risk that fully kills the idea today; everything downstream is a gradient. Skipping the probe means buying a build sprint to discover a 5/10 ceiling — a $50K–$100K diagnostic. The probe procedure is in the AI feasibility check.
Diagnostic question. Where is your one-page feasibility memo and what score did it return? “It seemed to work” is a no.
Property 3 — Existing workflow being replaced
Present. You name the role (paralegal, RevOps analyst, hospital biller) and the step they perform today that your product replaces. You describe their current tools and time-per-task.
Missing. The workflow does not exist yet. The pitch is “AI will let users do X” where nobody is doing X today. Demand creation required. Most “AI assistant” ideas live here.
Why it predicts. Replacing an existing workflow targets a buyer who already has the problem, an evaluator who already knows what good looks like, and a budget line that already exists. Inventing a workflow asks for all three. Invention can succeed (Notion, Linear), but the timeline is years, not weeks, and AI MVP budgets do not survive years.
Diagnostic question. Name the human role and the step. If you cannot name both today, you have an invention, not an automation — different game, different odds.
Property 4 — User pays today for the manual version
Present. A traceable spend line: payroll for the role you replace, contractor invoices for the work, or a SaaS subscription doing 60% of what you intend to do. Aggregate spend is at least 5x your intended price.
Missing. “Users said they would pay” with no current spend evidence. Survey-validated demand. Pre-orders without revenue. Willingness-to-pay is forward-looking, which makes it indistinguishable from politeness.
Why it predicts. Revenue elsewhere is a hard signal. People reorganize spending faster than they create new spending categories. Redirecting $8K/year of a $40K manual line to your AI is a procurement story; creating a new budget line adds a quarter to every sales cycle.
Diagnostic question. What payroll, contractor, or SaaS line item does your product redirect? Name the dollar figure. “They will pay because the product is great” is a no.
Property 5 — Eval rubric write-able in plain English
Present. You can write a paragraph describing a correct output in language a domain expert with no engineering background can grade. The rubric has three tiers — useful, recoverable, dangerous — with one or two examples per tier. Two graders independently agree on more than eight of ten cases.
Missing. “Correctness depends on context” or “the user knows it when they see it.” If the founder cannot define right in plain English, no eval suite can be written, no regression test can be run, and every model update is a roll of the dice. This is the AI-specific failure mode no Lean Startup checklist anticipates.
Why it predicts. Eval ambiguity scales linearly with build cost. Each ambiguous case becomes a debate between engineering and product, a re-prompting cycle, a regression. Without a rubric, model upgrades become risks instead of free upgrades. The eval-first PRD is the artifact that proves this property.
Diagnostic question. Write the paragraph. Show it to a domain expert and a non-expert. Each grades a sample output. If they disagree more than two times in ten, the rubric does not yet exist.
Property 6 — Failure mode is recoverable
Present. When the model outputs the worst plausible answer, the user notices, undoes, ignores, or overrides — without irreversible side effects. The product surface puts a human in the loop on irreversible actions (sending email, signing documents, executing trades, posting publicly) and shows the AI as draft, recommendation, or suggestion.
Missing. Auto-executed actions, no preview, no undo, no audit trail. A hallucinated answer ships to a customer or a regulator before anyone catches it. Most “agentic” ideas fail this in their first product spec and are rescued only by an explicit human-in-the-loop UX choice.
Why it predicts. Recoverable failure modes survive partial model accuracy. A clause-extraction tool finding 85% of clauses is useful because the user reviews the output. An autonomous contract signer finding 85% of clauses is a lawsuit waiting to happen.
Diagnostic question. Describe the worst plausible single output your product can emit. Is the user’s response “annoying” or “irrecoverable”? Only the first is shippable in week eight.
Property 7 — Wedge into a larger market
Present. A 50-person beachhead today — vertical, use case, or workflow narrow enough that the first 10 customers can be found in a week. The same product, with documented expansion paths, addresses a 5,000-person market in 24 months. Investors call this entry-narrow, expand-fast. We call it a wedge.
Missing. Day 1 targets the 5,000-person market. The product is generic enough to address all of them, which makes it specific enough to delight none. Or the inverse: the 50-person market has no expansion path, and you are building a lifestyle business with a fundraising story.
Why it predicts. A wedge gives you a small enough user pool to build a real eval suite from real users in week three, and a large enough surrounding market to fund the 24-month build that follows. The wedge is where AI MVPs improve fastest.
Diagnostic question. Who are the first 10 customers? Where are they? Which adjacent 50 do you sell to next? Three answers, in writing.
Property 8 — Founder has earned domain trust
Present. Three or more years inside the workflow you are automating. You recruit users from your network. You know what good looks like before the engineer does. The eval rubric writes itself because you have lived the failure modes. Investors call this founder-market fit; we call it the highest free factor of build velocity.
Missing. You read about the workflow on a podcast. The founder-market-fit story is reverse-engineered. User recruitment requires cold outreach. The eval rubric requires interviewing experts the founder does not have on speed dial.
Why it predicts. Domain trust accelerates every other property — user access, eval clarity, distribution, partner / data agreements arrive as side effects. The biggest predictor of an MVP shipping on time in 2026 is not engineering quality — it is whether the founder can text the first 10 users and get a Friday meeting.
Diagnostic question. Without LinkedIn outreach, can you put 10 real users in a Zoom room next week? If no, domain trust is the gap, and it is the slowest gap to close.
The scorecard
Score yes or no, no maybes, no halves. One sentence of evidence per yes.
| # | Property | Yes / No | Evidence |
|---|---|---|---|
| 1 | Single user, single task | ||
| 2 | Capability probe ≥ 7/10 | ||
| 3 | Existing workflow replaced | ||
| 4 | User pays today for manual version | ||
| 5 | Eval rubric in plain English | ||
| 6 | Failure mode recoverable | ||
| 7 | Wedge into larger market | ||
| 8 | Founder domain trust |
Honest filling time, with documents open: about 90 minutes. The slowest properties to evidence are usually 2 (the four-hour feasibility probe) and 5 (the two-hour rubric, with a domain expert). If the scorecard takes less than 30 minutes, you are guessing.
Interpretation
| Yes count | Reading | What to do |
|---|---|---|
| 8 / 8 | High-confidence. Build risk is execution, not idea. | Commit the sprint. Hand the scorecard to your build partner as the scoping artifact. |
| 6–7 / 8 | Conditional. One or two named gaps. | Spend a focused week closing each gap before sprint commit. Most close in days. |
| 4–5 / 8 | Borderline. Structural ambiguity. | Reframe against the failed properties — usually 1, 3, or 7 — and re-run. Many ideas at 4 become 7 next pass. |
| ≤ 3 / 8 | Low-confidence. Do not commit a build sprint. | Pivot, narrow, or shelve. These are categories or feelings, not products. |
The most common pattern in the wild is 5 / 8, with the missing properties being 2 (capability probe never run), 5 (rubric never written), and 6 (recoverability never specified). All three are AI-specific; all three close in a week of work; all three are the difference between a $200K build sprint that ships and one that produces a demo nobody pays for.
A completed scorecard is also the cleanest artifact you can hand a build partner before a scoping call. A partner who reads it and asks two sharp follow-ups — usually about the rubric or the failure mode — is doing the kind of work we describe in decoding the AI agency case study. A partner who quotes a six-month build without asking about the scorecard is selling time.
FAQ
How is this different from the 5-question test?
The 5-question test is a 30-minute reality check that filters out SaaS-ideas-with-an-LLM-bolted-on. The 8-property anatomy is a 90-minute confidence check that filters out ideas which are real but unbuildable in a six-to-twelve-week window. Run the 5-question test first; run this second.
Why eight properties and not five or twelve?
Five is too few — recoverability and eval-rubric ambiguity each kill builds, and neither folds into the others. Twelve is too many — properties beyond eight (pricing model, hiring plan, fundraising story) are downstream artifacts that depend on the eight. The eight are necessary and minimal; remove any and predictive power drops.
Does this work for B2C ideas?
Mostly yes, with one substitution. Property 4 (user pays today for manual version) becomes “user spends 2+ hours per week on the manual version.” Time is the consumer currency the way payroll is the enterprise currency. Properties 1, 3, 5, 6, 7, and 8 transfer directly. Property 2 is identical.
What if I score 8 / 8 but cannot build it?
You have an asset. The 8 / 8 scorecard is the artifact you walk into capital-raising or technical-hiring conversations with. It demonstrates pre-build diligence most founders pitching this quarter have not done. The idea risk stack covers how to retire each remaining risk before contract.
How do I score property 5 without a domain expert handy?
Hire one for two hours. A subject-matter expert in your target workflow will charge $200–$500 to read your output paragraph and grade five samples against it. Two hours of expert time settles a property that otherwise costs a build sprint to discover.
Can I cheat the scorecard by answering yes to everything?
You can, and the scorecard stops being useful. The defense is the evidence column. Each yes requires one sentence of evidence — a memo file path, a payroll line, a named user. Evidence that does not survive a re-read in a quiet room is a no.
Does the score change as the frontier moves?
Properties 2, 5, and 6 are sensitive to the model frontier. A 3 / 8 today can become 5 / 8 next quarter if a release closes the capability gap. Re-run the scorecard every two quarters. Properties 1, 3, 4, 7, and 8 are stable — they move only when the founder reframes or invests in domain trust.
Which property closes the most build risk per hour spent?
Property 5 (eval rubric in plain English) on most ideas we score. It costs two hours, settles three downstream properties (2, 6, and large portions of the build), and is the property most often skipped because it feels like documentation work. It is not — it is the contract the founder writes with the model, the user, and the build partner.
What does a “no” on property 8 mean for a first-time founder?
A real but not disqualifying gap. A first-time founder without domain trust still ships when they buy the gap closed: a co-founder, a paid advisor on retainer, a paid pilot with the first 10 users. The scorecard is not an identity test — it is a list of resourcing decisions the founder has not yet made.
Next step
If you scored 7 or 8, the AI MVP Scoping Worksheet converts a passing scorecard into the one-page brief your build partner will work from. If you scored 4 to 6, the idea risk stack is the next page — failed properties usually map onto a risk you have not yet retired. If you scored 3 or lower, the 5-question test is the page to revisit: the idea may be a SaaS idea with an LLM bolted on, not the AI idea you set out to validate.
Arthur Wandzel