Most “AI ideas” are SaaS ideas with an LLM bolted on top. That is a structural observation about the 2026 idea market, not a slight. A non-engineer founder watches a demo, has a flash of recognition, writes a sentence starting with “AI that…”, and within a week it is a Notion doc and a domain registration. None of that answers the prior question: is the idea real, or a workflow concept wearing an AI costume? The five-question test below is the 30-minute diagnostic, runnable before any code, prompt, or pitch deck.
This is the front gate of the idea-validation playbook and the second checkpoint of the broader idea-to-product manifesto. The week-long procedure that comes next is in the sibling spoke on validating an AI idea before you write a single prompt. Run this first.
Decision Scope
This is an editorial diagnostic, not legal, financial, or technical advice. Treat the questions and thresholds as planning heuristics. Validate against your own domain, users, and the current frontier-model leaderboard.
Why the test exists
The Lean Startup canon was written when the binding constraint on software was demand and tech risk was negligible — specify a CRUD app and it ships. AI inverts that. Demand for AI products in 2026 is enormous and largely uncontested. The hard part is the new question: can the model do this reliably, cheaply, and durably enough to be a business? That splits into five sub-questions, each binary enough to answer in five minutes.
Most “AI ideas” fail not because demand evaporates, but because one of the five answers comes back wrong and the founder did not check. Six months later, the wrong answer surfaces as churn, broken unit economics, or a competitor shipping the same feature. The test costs half an hour. Skipping it costs two quarters.
How to run the test
Set a 30-minute block. Write your idea down as “My AI does X for Y users when Z.” Open claude.ai, chatgpt.com, and gemini.google.com. Answer each question yes or no — no maybes — with one sentence of reasoning. Count the yeses. Do this alone, before social proof contaminates the result.
Question 1 — the model-reality test
Would users still pay for this if the AI worked 80% as well as you imagine?
Most founders pitch against the AI in their head — an idealized version that never hallucinates, runs in 200ms, costs a tenth of a cent. The 80% version is closer to what you ship: reasoning-heavy tasks lose accuracy; long inputs degrade; niche domains paraphrase; agentic chains compound latency; cost-per-query refuses to land.
Subtract twenty percent and ask whether the product still has a buyer. A yes means it has a usable floor below the imagined ceiling — a clause-extraction tool finding 85% of clauses still beats a paralegal at thirty dollars an hour. A no means the entire economic argument depends on the model behaving better than current models do.
2026 example. Founder A pitches meeting notes promising “perfect transcription, summaries, and action items.” Founder B pitches “95% transcription, summaries reviewed before send, action items as suggestions.” At 80% performance, A is unusable because A promised perfection. B still saves the user thirty minutes. A is selling vapor; B is selling a product.
Question 2 — the AI-lift test
Does the AI carry the value, or does the wrapper?
This decides whether you have an AI product or a SaaS product with an LLM call inside. The honest framing: if you removed the AI and replaced it with a forty-dollar-an-hour contractor, would the rest of the product still be valuable?
If yes, the AI is a cost optimizer. The product is a workflow tool — value lives in integrations, UI, permissions, data model, audit log. The AI shrinks per-task cost from forty dollars to forty cents, which is real, but does not change what the product is. These are workflow companies dressed in AI clothing.
If no, the AI is the product. Strip it out and you have a website. The moat is what you have built around the model: data flywheels, domain fine-tunes, retrieval indexes, evaluation harnesses, distribution.
Founders confuse the two and pay for it. An AI-native founder builds a wrapper, raises at an AI multiple, and discovers six months in that the upgrade arrived for free at the workflow competitor. A workflow founder pitches an AI vision and ships a CRUD app with a summary feature.
2026 example. Founder C pitches “an AI that listens to your calls and tells you what to say better.” Founder D pitches “a CRM with call recording, talk-track templates, and an AI that scores your talk-tracks.” C passes — strip the AI and the product is nothing. D fails — strip the AI and the CRM still holds most of the value.
Question 3 — the commoditization test
Would a six-month model improvement make your product obsolete or stronger?
Frontier models improve roughly every two quarters — typically 10 to 30 percent on accuracy, half the latency, or both. Think one cycle ahead: when Claude Opus 4.8, GPT-5, and Gemini 2.5 Pro each ship a successor, does that make your product better or eat it?
If stronger, your product rides the frontier curve. The next release is a free upgrade that raises accuracy, drops cost, or unlocks a feature your roadmap could not previously support. Your moat is everything around the model — workflow, data, distribution, trust.
If obsolete, your product is the gap between current capability and what users want. The next release closes part of that gap and your differentiation evaporates — first-generation summarization startups died this way, each release eating a layer of their UI. Building on “we make the raw model marginally better at X” is a footrace against teams of a thousand researchers with multi-billion-dollar compute. You will lose.
2026 example. Founder E pitches “AI deep research, better than the chat consoles.” Founder F pitches “AI deep research integrated with your firm’s internal documents, CRM, and compliance audit trail, citations linked to PDFs in your DMS.” E fails — every release narrows that gap to zero. F passes — the model getting better is a free upgrade to a product whose value lives in the integrations.
Question 4 — the defensibility test
Could a frontier vendor ship this as a feature next quarter?
OpenAI, Anthropic, and Google each ship a quarterly release cadence with new features bundled in. If your product is a feature one of those three could plausibly ship in their next release, you are building on rented land. Open the roadmaps and last four releases of those three and ask: is your idea closer to “feature they would ship” or “product they would not”?
Features they would ship: chat with PDF, summarize this page, draft an email in my style, extract a table, generate an image, code an app from a sketch, browse the web with citations. All shipped inside the chat consoles in the last eighteen months, and the boundary keeps expanding.
Products they would not ship: multi-tenant integration with named third-party SaaS, compliance certifications they do not hold (HITRUST, FedRAMP High), proprietary user-specific data they cannot access, or a buyer who is a job function (legal counsel, RevOps, compliance officer) rather than a consumer or developer.
This is an organism question, not a moat question. Frontier vendors eat anything inside their natural reach. Your idea has to be outside it.
2026 example. Founder G pitches “an AI that remembers your past conversations and personalizes future answers.” Founder H pitches “an AI that integrates with a private-equity firm’s IR CRM, deal-pipeline store, and SEC-filing archive to draft LP letters in the firm’s prior-year voice.” G fails — OpenAI has shipped that in pieces since 2024. H passes — there is no plausible release in which OpenAI ships a private-equity LP-letter drafter.
Question 5 — the moat test
Do you have proprietary data, distribution, or trust the model can’t replicate?
Once the model becomes a commodity — which, for most general capabilities, it has — defensibility moves outside the model. Three categories survive: data the model has not been trained on, distribution the model cannot buy, and trust the model cannot earn.
Proprietary data. Customer-specific inputs (calendar, CRM history, internal documents, transaction stream) the model needs and the foundation lab does not have. The most durable category — it compounds with every user.
Proprietary distribution. Channels the labs cannot replicate at your cost — a vertical sales motion, an enterprise customer base, a regulated marketplace, a developer community. The model is free; the customer is not.
Proprietary trust. Certifications, audit history, named insurance liability, vertical reputation. A hospital does not adopt the chat console for clinical decision support not because the model is worse, but because it cannot buy malpractice insurance against a chat console.
Not on this list: prompts, fine-tunes on public data, UX, brand. All four are reproducible in days by anyone with a credit card and a credible engineer. Features, not moats.
2026 example. Founder I pitches “a radiology assistant trained on public datasets.” Founder J pitches “a radiology assistant integrated with three hospitals’ PACS, trained on their de-identified scans under a data-use agreement, reviewed by their malpractice carrier.” I fails — public datasets are table stakes. J passes — the data, the partnerships, and the carrier review are proprietary in ways the model cannot replicate.
Scoring the result
Count the yeses.
| Yes count | Reading |
|---|---|
| 5 of 5 | Your idea is real. Proceed to the 7-step validation procedure to harden the eval suite, then to a PRD. |
| 4 of 5 | Real with one named risk. The “no” question is the work. Spend a focused day on it before committing build budget. |
| 3 of 5 | Borderline. Two answers are structural — they do not soften with research. Reframe X, Y, or the user, then re-run. Many ideas at 3 become real at 5 with tighter scope. |
| 2 of 5 or fewer | Not a real AI idea. It may still be a real SaaS idea, but not the AI startup you set out to validate. Pivot, narrow, or shelve. |
The most common pattern is 3 of 5, with the failed questions being AI-lift and commoditization. Founders pass demand-side questions easily and fail capability-side ones because the Lean Startup canon never taught them to ask the latter. The good news: the failed-question diagnostic is precise — you know what to fix.
What to do with the result
A 5-of-5 earns the right to spend a week on the deeper 7-step procedure, which ends in a go/no-go memo and an eval rubric you can hand a build partner. A 4-of-5 needs the failed question solved first; a 3-of-5 needs reframing; a 2-of-5 needs pivoting.
The scored answer is also the cleanest artifact to hand a build partner before a scoping call. A partner who reads it and asks three sharp follow-ups is doing the kind of work we describe in what makes an AI agency case study real. A partner who skips the diagnostic and quotes a six-month build is selling time. The companion piece on why an “AI idea” is not yet a product covers what the artifact looks like once the test is passed.
FAQ
How long should the 5-question test take?
Thirty minutes for a focused founder who already has the one-sentence idea written down. Two hours if you have not written the sentence yet. If it takes longer, you do not have an idea — you have a feeling. Compress it into “My AI does X for Y users when Z” before retrying.
Can I cheat the test by answering “yes” to everything?
You can, and the cost is that the test stops being useful. Founders rationalize a yes because the alternative is admitting the idea is not real. The defense: write a sentence of reasoning under each answer and re-read all five at the end. Reasoning that does not survive a re-read is a no.
What if my idea scores 5 of 5 but I cannot build it?
That is the case for raising capital, hiring a technical co-founder, or engaging a build partner. The diagnostic is the artifact you walk into those conversations with — it demonstrates work most “AI founders” pitching this quarter have not done.
Does the test apply to agentic ideas?
Yes, with one modification. Question 1 has a tougher floor for agents — they compound errors across steps, so 80% step-level becomes 50% end-to-end on a five-step plan. Apply it to the end-to-end task. Question 3 is often a stronger yes for agents because each release improves reasoning and tool use disproportionately.
What if the test contradicts what investors have told me?
The test is for you, not your investors. They evaluate against a portfolio thesis; founders evaluate against “will I still want to be working on this in 2028?” If the test says no and an investor says yes, the investor may be right about fundability and wrong about durability. Run it anyway.
What if I fail the AI-lift test (question 2)?
Two honest options. First, accept the idea is a workflow product with AI features, and price, raise, and build accordingly. Second, reframe so the AI is doing the work the user pays for — usually by narrowing X and Y until the AI is the entire job. Pretending the wrapper is an AI product is the dishonest move.
Why do AI ideas fail the commoditization test?
The value proposition reduces to “we make a current model better at X with our prompts and UI.” Every release closes part of X natively. Add something the release cannot eat — proprietary data, integrations, distribution, or a vertical workflow no foundation lab will build.
Should I re-run the test as the frontier moves?
Yes, every two quarters. A 3-of-5 today becomes a 4-of-5 if a release closes the gap blocking question 1, or drops to 2-of-5 if the same release commoditizes the differentiation behind question 4. The diagnostic is a snapshot, not a verdict.
Next step
If your idea scored 4 or 5, proceed to the week-long validation procedure. If lower, the question is what to reframe before re-running. The idea-validation playbook takes a tested idea to a defensible PRD, and the idea-to-product manifesto sets the program in context.
Arthur Wandzel