In 2026, an AI investor walks into a pitch with a different scorecard than they brought to a SaaS pitch in 2018. Team and market still matter. But four or five evaluation lenses that did not exist in the 2018 deck are now what the deal gets killed on. A non-technical founder who builds against the old scorecard pitches into a frame the room is no longer using.
This article names the lenses. It walks a non-engineer through how a 2026 venture investor evaluates an AI idea, why each lens exists, and what to put on slide 7 before it is asked. It is a spoke in the broader idea-validation playbook and the idea-to-product manifesto.
Decision Scope
This is an editorial framework synthesized from publicly published VC positions, not investment advice. Treat the lenses as planning heuristics. Investor practice varies by stage and fund — validate against the specific firms you pitch.
Why AI evaluation diverges from SaaS evaluation
The 2018 SaaS scorecard had a familiar shape: team, market, traction (ARR, MRR growth, NRR), unit economics, competitive position. The 2026 AI scorecard inherits all of that — and layers a second scorecard on top that asks a different question: what fraction of the value being pitched is created by a model the founder does not own, on a cost curve they do not control, against a frontier that will move under their feet within 18 months?
That second scorecard is the one founders underprepare for. A pitch that nails the SaaS-era questions and fumbles the AI-specific lenses gets a polite pass. The room says “we’ll watch the space,” but the lens is what closed the door.
Investors in 2026 have converged, in different vocabularies, on four or five AI-specific lenses. a16z talks about “Wave 2” and durable advantage; Sequoia frames “systems of intelligence” and the AI 50; Bessemer publishes roadmap memos and cost-curve research; Madrona writes about model portability. The questions overlap. The lenses below are the synthesis.
Lens 1: Capability dependency — model or wrapper?
The first lens is the simplest to articulate and the hardest to answer well: how much of the value the founder is pitching is carried by the model, and how much by everything around it?
The 2023–2024 “GPT wrapper” objection became a meme because it was right often enough to be useful. By 2026 it has matured into a structured question, framed loosely in a16z’s “Wave 2” essays and Sequoia’s framing of AI as moving from “system of record” toward “system of intelligence”. The investor is not asking whether you call a frontier API. They are asking three things underneath:
- Where in the value chain does the model sit? Is the model the product (a coding agent), an embedded reasoning step (a clause extractor in a contract workflow), or a marginal improvement on a non-AI workflow (smarter search inside a legacy CRM)?
- What fraction of buyer surplus survives if the model commodifies? If Claude Opus 4.8, GPT-5, and Gemini 2.5 Pro become equally good and equally cheap at your core task, what is left to charge for?
- Is the founder thinking like a product company or a feature? Investors funded “feature companies” in 2023 and watched them get absorbed in 2024–2025. They are allergic to it now.
The healthy answer is not “we are not a wrapper.” Every AI product is one in the literal sense. The healthy answer names the non-model assets compounding around the call: a proprietary eval set, a fine-tune on customer data, a workflow integration with switching cost, a labeled dataset built from real usage, or distribution into a buyer the frontier vendor cannot reach. One is the floor. Two is comfortable. Three reads as a durable AI product. Zero is a feature.
Lens 2: Model-vendor risk — what happens when the provider moves?
The second lens treats the frontier provider the way 2014 SaaS investors treated AWS concentration: a known risk that must be priced, not a taboo to avoid. Bessemer’s roadmap memos and Madrona’s writing on model portability frame this as an explicit due-diligence question.
Three sub-questions an investor will probe, often without naming them:
- Pricing sensitivity. What happens to gross margin if Anthropic, OpenAI, or Google raises your tier 30%? “We’d pass it through” without showing the model is guessing. A one-line gross-margin sensitivity table earns trust.
- Deprecation exposure. Frontier aliases get deprecated. Quality bands shift between versions. What is your fallback if your current model is sunset on six months’ notice? Have you tested the next-tier model against your eval set?
- Portability. Can you swap providers without re-engineering the prompt scaffold and re-validating the eval? The honest answer in 2026 is usually “partial” — prompts overfit to provider behavior. The investor wants you to have measured the gap, not pretended it does not exist.
The polished answer fits in two lines: “Core capability passes our eval at parity on Claude Opus 4.8 and GPT-5; Gemini 2.5 Pro is within 8% and tested quarterly. A 30% provider price hike moves blended gross margin from 71% to 64%; a fine-tuned Llama 3.3 escape hatch backstops at 58%.” Two lines. Whether the numbers are exactly right matters less than that the founder has done the work.
Lens 3: Eval defensibility — frozen test set, not a demo
Pre-revenue AI startups have a problem the 2018 SaaS founder did not: no ARR to point at. The substitute the 2026 investor accepts is not a demo, not a waitlist, not a thread of testimonials. It is a frozen evaluation set with a pass-rate curve over time.
Sequoia’s “AI 50” framing treats measurable intelligence — defensible, repeatable improvement against a held-out test set — as the closest analog to traction for pre-revenue AI products. Anthropic and OpenAI public engineering posts on evaluation echo the same point: the eval is the artifact, not the demo. See the eval-first build playbook for the founder-side procedure.
What the investor wants on the deck:
- A held-out test set of 50 to 500 realistic inputs, frozen and version-controlled.
- A clear pass criterion (a written rubric — not “looks good”).
- A pass-rate curve over time, ideally with at least three weekly data points showing the rate climbing.
- A baseline against a stock frontier model so the investor sees the lift your scaffolding adds.
A founder who walks into a 2026 pitch with a demo and no eval has done the SaaS-era homework but not the AI-era homework. A founder who walks in with an eval and no demo can usually close the room. A demo is theater; an eval is evidence. This is the single largest mismatch between founder expectations and investor expectations today.
Lens 4: Inference-cost slope — does cost outrun price?
The fourth lens is the easiest to overlook because it asks the founder to model a curve, not a snapshot. Investors who have watched frontier inference pricing since 2023 know the headline: the per-token cost of a fixed capability has dropped roughly 4× to 10× per major generation across OpenAI and Anthropic public pricing pages.
The lens question: does cost-per-task drop fast enough to outpace the pricing pressure your category will face?
This is structural. Even if the founder holds list price, competitors will not. The investor probes three things:
- Cost-per-successful-task today. Not the API list price — the realized cost per usable output, accounting for retries, multi-step calls, and post-processing.
- The slope. What did the same successful task cost six months ago, and what would the founder bet it costs six months from now? Investors who read Bessemer’s State of the Cloud research and the public pricing histories will sanity-check the answer.
- Pricing-power independence. Does pricing power depend on cost staying high (margin lives in the model bill) or on value delivered (margin lives in workflow lock-in)? The first is fragile. The second is durable.
A founder who has done this work shows it on one chart: a downward-sloping cost curve, gross margin overlaid, and a one-sentence answer to “what happens to unit economics in 18 months at projected costs?” The chart does not need to be precise. It needs to show the founder thinks in slopes.
Lens 5: Moat vs model improvement — survive the next 10×?
The fifth lens closes the most AI deals when answered well and kills the most when answered badly. The question, in plain language: when frontier models get 10× better at your task in 18 months, does your moat survive or evaporate?
a16z has discussed a version as “wrapper graduation” — whether a company accumulates enough non-model assets fast enough to survive the day the model does the wrapper’s job natively. Sequoia’s AI 50 framing treats durability the same way: are system-of-intelligence assets compounding faster than the frontier is improving? Greylock partner essays argue in similar language that durable AI companies turn the frontier’s improvements into product gains for themselves, not for competitors.
Three honest answers an investor accepts:
- The moat is data. The product captures a labeled dataset competitors cannot easily replicate, and it improves with usage. Frontier improvements compound the moat rather than erase it.
- The moat is workflow integration. The product is embedded in a buyer’s daily process with switching costs. Frontier improvements arrive as upgrades to your product, not as a competitor.
- The moat is distribution. The product reaches a buyer the frontier vendor structurally cannot — regulated industries, channel-only markets, white-label embedded use. Frontier improvements lower your cost of goods.
The unhealthy answer — and the most common — is some variant of “we are just better at prompting.” A 2026 investor assumes that gap closes inside two model generations. The founder who cannot name a moat surviving the next 10× is asking the room to underwrite a feature, not a company. The editorial-600 piece on reading the artifacts AI agencies produce — decoding the AI agency case study — has related advice on separating signal from theater.
What a non-technical founder should put on slide 7
Slide 7 — the slide that earns the meeting or ends it — should answer the five lenses in compressed form:
Capability dependency. [The AI capability the product depends on; one or two non-model assets compounding around it.]
Model-vendor risk. [Eval parity across at least two frontier providers; margin sensitivity to a 30% provider price hike.]
Eval defensibility. [Frozen test set size and current pass rate; pass-rate slope over the last 4 to 8 weeks.]
Inference-cost slope. [Cost-per-successful-task today; trajectory in the last 6 to 12 months.]
Moat vs model improvement. [Whether the moat is data, workflow, or distribution, and why frontier improvements compound it rather than erase it.]
Five lenses, ten lines, one slide. The investor reads it in 30 seconds and either leans in or politely closes. Founders who can write the slide are the same founders who can hold the conversation. Founders who cannot are usually still thinking in 2018 SaaS terms — see how to validate an AI product idea before you write a single prompt for the upstream capability check, and what is product-market fit for AI products for the matching demand-side discipline.
FAQ
Are these the only lenses investors apply?
No. The 2018 SaaS scorecard is still active underneath — team, market, GTM, unit economics, traction. These five are AI-specific additions. Necessary, not sufficient.
Do all AI investors use this exact frame?
No two firms have identical scorecards. a16z, Sequoia, Bessemer, Madrona, Greylock, and Lightspeed lean differently. The lenses above synthesize the union of questions you will hear.
How does this differ from a 2018 SaaS investor scorecard?
The 2018 scorecard treated technology as deterministic. The 2026 scorecard treats the model as a moving variable the founder does not control. Three of the five lenses did not appear in 2018.
Is the “GPT wrapper” objection still a real concern in 2026?
Yes, in a sharper form. The objection has matured from “you call an API” into “your value is fully reproducible by the model vendor on six months’ notice.” The capability-dependency lens above is how it gets diagnosed.
What if I am pre-product and have no eval to show?
Build a 30-input frozen test set with a written rubric, run it against two frontier models manually. That artifact — even at low pass rate — outperforms a demo, because it shows the founder thinks in the right vocabulary.
How specific should the slide-7 cost numbers be?
Specific enough to be honest, not precise enough to be wrong. Investors do not penalize an order-of-magnitude estimate labeled as one — they penalize false precision. A gross-margin range and one cost-per-task figure are sufficient at seed and Series A.
Does this frame apply to vertical AI startups differently?
Vertical AI startups answer Lens 5 more easily — distribution into a regulated vertical is itself a moat. They score harder on Lens 3 because the test set must be domain-specific. Lead with the vertical-distribution moat; back-fill the eval.
Does this frame apply to agentic-product pitches?
With one addition. Agentic products carry compounded reliability risk — a 95% per-step pass rate compounds to roughly 60% across ten steps. Investors will ask for end-to-end task success, not per-step. Size the test set against full task completion.
Which frontier model should I cite on the deck?
The one your eval set actually passes on. The current frontier in mid-2026 is Claude Opus 4.8, GPT-5, and Gemini 2.5 Pro at the top tier, with Sonnet 4.6 and Haiku 4.5 covering cheaper bands and Llama 3.3 as the open-weights fallback. Know which tier you live in.
What if a VC has not published their AI-evaluation thesis publicly?
Most established AI investors have published something — a memo, essay, podcast, or portfolio post. Read the partner’s writing before the meeting. The five lenses are the union; the firm will weight them differently.
Next step
The slide-7 frame is the deliverable. If you can write it honestly, you are ready to pitch. If you cannot, the gap is upstream — capability not validated, eval not built, or moat not thought through. The next chapters — the idea-validation playbook, how to validate an AI product idea, and the eval-first build playbook — close those gaps in order. The investor meeting is not the validation; the validation is what makes the meeting easy.
Arthur Wandzel