Every non-technical founder with an AI idea hits the same first decision after they pick a partner: do we book a free discovery call, or do we commit to a paid pilot? The two sound like cheap-and-expensive versions of the same thing. They are not. A discovery call is a 60-to-90-minute scope-mapping conversation that produces a memo. A paid pilot is a one-to-two-week engagement with a fixed fee, named deliverables, and a runnable artifact at the end. They solve different problems. Choosing the wrong one wastes two to four weeks of founder time or ten to thirty thousand dollars of cash. This article is the structural comparison and the four-property rubric that points to the right answer.
The framework here builds on the idea validation playbook within the broader idea-to-product manifesto. Companion reading: idea-to-product as a service and what a defensible idea-to-product SOW looks like.
Why this is the wrong question, framed correctly
Founders ask “discovery call or paid pilot?” as if the choice were free versus paid. It is not. A discovery call answers who and what: who the partner is, what they think the problem is, what shape an engagement might take. A paid pilot answers can: can the partner do the work, can your data sustain it, can the eval threshold be met inside a defensible budget. Treating them as substitutes is the same category mistake as treating a job interview and a paid trial week as substitutes — both are valid hiring instruments and they solve different problems.
The market amplifies the confusion. Half the agencies in 2026 brand their two-to-four-week paid scoping engagement as a “discovery phase” and their free sales call as a “discovery call”. This article uses the cleaner pair: discovery call for the free 60-to-90-minute conversation, and paid pilot for the one-to-two-week paid engagement that produces a runnable artifact. The dollar-cost gap matters less than the artifact gap, and the artifact gap is what determines which one is right for your moment.
What a discovery call actually is
A discovery call is a structured 60-to-90-minute conversation between the founder and a senior member of the candidate partner team — typically the principal who would lead the engagement, not a salesperson. It is free, scheduled within a week of first contact, and produces three artifacts when run well:
- A scoping memo — one to two pages, written by the partner within 48 hours. It restates the idea in the partner’s words, names the three to five capabilities the build would require, and flags risks the partner sees from outside.
- A preliminary effort range — not a quote. A weeks-and-team-shape sketch: “this looks like an eight-to-twelve-week engagement with a senior plus a mid, roughly $100K-$160K depending on the data-access work.” The range is wide on purpose.
- A go/no-go on partner-fit — the founder’s read on whether they want to work with this person, and the partner’s read on whether the engagement is one they would take. Either side can decline without consequence.
What a discovery call does not produce: working code, an eval baseline, a tested data assumption, a fixed-fee proposal, or evidence of partner capability beyond the conversation itself. It confirms partner identity and engagement shape. It does not confirm partner ability or idea feasibility.
The discovery call is the right artifact when the founder has multiple candidate partners and needs to compare them, or when the founder is still defining what they would even hire someone to do. A discovery call run badly is a vendor sales pitch with a memo template. A discovery call run well is closer to a structured interview where the founder asks most of the questions and the partner does most of the listening.
What a paid pilot actually is
A paid pilot is a one-to-two-week engagement with a fixed fee — typically $5K to $25K in 2026, depending on data complexity and seniority — that produces three artifacts:
- A runnable prototype — not a demo, not a deck. Code that runs end-to-end on the founder’s actual data or a representative sample, even if narrowly scoped.
- An eval baseline — a measured pass rate on a small (50-150) eval set built during the pilot, with the methodology documented. Without an eval, a pilot is theater.
- A defensible Statement of Work for the main engagement — a fixed-fee, milestone-billed SOW that the pilot’s evidence justifies. Not a guess. A SOW whose effort estimate, eval threshold, team shape, and timeline all trace to what the pilot found.
What a paid pilot does not produce: a finished product, a hardened deployment, full data integrations, or a guarantee that the full engagement will succeed. A pilot tests one critical assumption end-to-end at small scale — usually the riskiest capability. A pilot is a falsification instrument; its job is to surface evidence that lets both sides write a defensible SOW or walk away cleanly.
The paid pilot is the right artifact when the founder already trusts the partner and the open question is feasibility — capability, data, or cost. It is the wrong artifact when the open question is partner-fit, because paid pilots imply exclusivity for the duration. Running parallel paid pilots with three candidate partners is a sign of broken process, not thorough diligence. The structural argument for the pilot, made in depth in why your AI agency should run a paid pilot before the main contract, is that the main-contract-first model is broken in 2026: both sides learn in week eight what they should have learned in week one. A pilot collapses that loop into two weeks of paid, falsifiable work.
The four-property decision rubric
The right modality is determined by four properties of the founder’s situation. Score each one low or high.
Property 1: Capability uncertainty. How uncertain are you that current frontier models can do the core job?
- Low — known-solvable pattern (RAG over documents, structured extraction from semi-structured text, a routine agent loop with two or three tools). Plenty of public reference implementations exist.
- High — edge of model competence (long-horizon planning, novel reasoning, regulated-domain output requiring tight calibration). No close public reference exists.
Property 2: Partner-fit uncertainty. How uncertain are you that this specific partner is the right one?
- Low — existing relationship, strong referral from a trusted operator, or you have done one or more discovery calls and the partner is in the final-two.
- High — you are early in partner-shopping; you have a long list and have spoken to nobody on it, or you have spoken to one person and want to compare against the field.
Property 3: Runway. How much cash do you have, and how much can you allocate to procurement instruments before the main engagement?
- High — $300K+ pre-seed; comfortably absorb $20K-$30K of pilot spend across two candidate partners without affecting the main build budget.
- Low — $150K or less; every dollar of pilot spend reduces the main-build budget materially. Carta’s 2025 startup compensation data puts the median pre-seed raise at roughly $1.2M for technical-founder teams and $400K-$700K for non-technical-founder teams.
Property 4: Urgency. How soon do you need to be in market with a working product?
- Low — nine or more months of runway, no investor-imposed milestone in the next six months.
- High — investor milestone in three to six months, a customer waiting, a competitor moving, or a personal runway forcing a decision.
Capability and partner-fit are the primary axes; runway and urgency are the modulators that compress or extend the recommendation.
The two-by-two procurement matrix
Score Property 1 and Property 2. The cell tells you which artifact to start with.
| Partner-fit LOW (you trust the partner) | Partner-fit HIGH (you are still shopping) | |
|---|---|---|
| Capability LOW (you believe models can do it) | Skip both → go to SOW | Discovery call with two to three partners |
| Capability HIGH (capability is the risk) | Paid pilot with this partner | Discovery call first, then paid pilot with the finalist |
Top-left. You trust the partner and the capability works. Neither artifact is strictly required; go directly to SOW negotiation. The partner may still propose a paid scoping week to write the PRD — that is a different artifact from a falsification pilot.
Top-right. You believe the capability is workable; the open question is which partner. Run two to three discovery calls in parallel within a two-week window, compare memos and effort ranges, pick a partner, go to SOW. Paid pilots here are wasteful — you are paying to learn what discovery calls and the partner’s past work already proved.
Bottom-left. You trust the partner; the capability is the risk. Skip the discovery call and run a paid pilot. A 90-minute call cannot test capability. Spend $10K-$20K on a two-week pilot and let the evidence decide.
Bottom-right. Both axes uncertain. Run discovery calls first to pick a finalist, then a paid pilot with that finalist. This is the hybrid sequence and the most disciplined path for founders with adequate runway.
Low runway pushes toward the cheaper artifact in any tied cell. High urgency pushes toward starting the pilot immediately even when partner-fit is uncertain, because the cost of a wrong-partner pilot is bounded ($10K-$20K and two weeks) while the cost of analysis paralysis is the build window closing.
The hybrid sequence
For the disciplined non-technical founder in the bottom-right cell — capability uncertain, partner-fit uncertain — the dominant 2026 pattern is the three-step sequence:
Step 1: discovery calls with three to five candidate partners (two weeks total). Source partners from a deliberate process — referrals from operators you trust, the partner roundup in the best idea-to-product partners for solo founders, and one or two cold introductions if the warm list is thin. Run each call as a structured interview. Write a one-paragraph internal note within an hour of each call. By the end of week two you have three to five memos and a clear top-two.
Step 2: paid pilot with the finalist (two weeks). Pick one partner — the one whose memo was sharpest and whose preliminary range felt the most honest. Sign a one-page pilot agreement: $10K-$20K fixed, two weeks, one named deliverable, one named eval, one named senior engineer on the partner side, kill clause on both sides. Tell the runner-up the timeline; do not run parallel pilots.
Step 3: main engagement signed off the pilot’s evidence (week five onward). The pilot produced a runnable artifact, an eval baseline, and a defensible SOW. If the pilot succeeded on the named threshold (typically 70-85% pass on the small eval set), sign the main SOW. If it failed cleanly — eval missed by a wide margin, or the capability is genuinely beyond current models — you walk, having spent $10K-$20K to avoid signing a $150K SOW for an unbuildable product.
The hybrid sequence costs five to six weeks of calendar time and $10K-$20K of cash. For a founder with nine-plus months of runway, that is the right insurance. For a founder with three months and a fixed milestone, the sequence compresses: one week of two or three discovery calls, one week of pilot, immediate main engagement. The deeper enterprise-scale variant of Step 2 is the AI agency discovery week.
When each modality is wrong
Discovery call is wrong when:
- You already trust the partner and the open question is capability. A 90-minute conversation cannot test whether models behave on your data. Booking a call here is procurement theater.
- You are running it as a sales screening when what you need is a paid scoping engagement. If the real question is “can I afford the main build”, a preliminary range is not precise enough.
- You have done four already. After three to five calls, the marginal call’s information value is near zero. The next move is a pilot with the finalist.
Paid pilot is wrong when:
- You are still partner-shopping. Paid pilots imply exclusivity. Running them in parallel is expensive theater and signals broken process.
- The capability is a known-solvable pattern with strong public reference implementations. Paying $15K to confirm RAG works on PDFs is paying for what the public literature already established.
- Your runway is below $200K total. At that tier, a $15K pilot is 7.5% of the budget; the same money buys an extra week of main-build work.
- The partner has a recent, public, eval-protected case study in a near-neighbor of your problem. Their existing artifacts already prove what a pilot would prove.
A clean discovery call is sometimes followed by an immediate “no” from either side. That is the call working correctly. A clean paid pilot is sometimes followed by a “no” with the eval baseline in hand. That is the pilot working correctly. The procurement decision was wrong only when the founder paid for the wrong artifact and did not learn what they needed.
A worked example
Marcus is a 14-year veteran of commercial real-estate underwriting at two mid-market lenders. He has watched analysts spend four to six hours per loan reading 80-to-200-page Offering Memoranda, extracting line items, and reconciling them against the lender’s underwriting template. He believes an AI agent could compress that to 30 minutes per loan with a human-in-the-loop checkpoint. He has $480K pre-seed, no engineering background, no co-founder.
He runs the rubric.
- Capability uncertainty. High. OM extraction is structured-extraction-from-unstructured-text, which models do well in general, but the lender’s template has 67 line items, some of which are derived (cap-rate computation, DSCR with stressed assumptions). Plausible but not certain at the precision the lender requires.
- Partner-fit uncertainty. High. He has talked to one partner on a referral and has a list of four more candidates from the partner roundup plus two from a YC networking call.
- Runway. High. $480K is comfortable; absorbing $20K-$30K of procurement spend does not break the main-build math.
- Urgency. Medium. No investor milestone for nine months; one angel has expressed interest in a Series Seed at month 12 conditional on a working product.
The matrix points to bottom-right (Capability HIGH + Partner-fit HIGH). The hybrid sequence is the recommended path. Marcus executes:
- Week 1-2: Five discovery calls. Two produce sharp memos with effort ranges between $130K and $180K and explicit flags on the 67-line-item precision risk. One of those two names a specific pilot proposal — extract 8 of the hardest line items from 12 representative OMs, measure precision and recall against analyst-reconciled ground truth.
- Week 3-4: Paid pilot, $14K fixed, two weeks. The pilot extracts 8 line items from 12 OMs. The eval set is 12 × 8 = 96 measurements scored against analyst ground truth. Result: 89% precision on 6 items, 71% on the seventh (DSCR with stress), 54% on the eighth (cap-rate adjustments for tenant-credit risk). The pilot surfaced an undiscussed risk — three OMs were photographic PDFs requiring an OCR pre-processing step the original effort range did not include.
- Week 5: SOW negotiated on the pilot’s evidence. Main engagement adds a one-week OCR-pre-processing milestone, ringfences cap-rate-adjustments to a Phase 2, and sets the primary eval threshold at 85% precision across the remaining 7 line items. Total: $158K, 10 weeks, eval-bound. Marcus signs.
Total procurement spend: $14K. Total elapsed time: five weeks. Result: a main engagement Marcus believes in, sized to his real capability question, with the riskiest line item ringfenced off the critical path. The same founder, without the sequence, would either have signed a $150K SOW that under-scoped OCR and over-scoped cap-rate, or run five parallel paid pilots and spent $70K to learn what one well-chosen pilot revealed for $14K.
Cost-of-wrong-choice math
The procurement decision has bounded but real downside on both sides of the mismatch.
Wrong-direction error 1: chose discovery call when paid pilot was correct. The founder spends three weeks doing five discovery calls, picks a partner, signs a $150K SOW, and discovers in week six of the main engagement that the capability assumption is wrong. Cost: $60K-$120K of partial main-engagement spend plus four-to-six weeks of calendar burn plus relationship damage.
Wrong-direction error 2: chose paid pilot when discovery call was correct. The founder runs a $15K pilot with a partner they could have screened in 90 minutes. The pilot succeeds (because the capability is known-solvable), but the founder discovers the partner’s communication style or stack preferences are misaligned. Cost: $15K-$25K plus two weeks plus moderate relationship friction.
The asymmetry is structural. Error 1 is materially more expensive than Error 2. The rubric is intentionally biased toward the paid pilot whenever capability uncertainty is high, because the downside of not paying for falsification when capability is uncertain is the dominant procurement risk in AI builds in 2026. McKinsey’s 2025 state-of-AI puts the fraction of organizations capturing meaningful EBIT impact from AI at roughly one in six, with capability-misjudgment at scoping time as a dominant failure mode. The pilot does not eliminate that risk; it converts the discovery of the misjudgment from week six of a $150K engagement into week two of a $15K engagement. That is the trade the rubric is built around.
Frequently Asked Questions
What does a discovery call typically cover?
A well-run 60-to-90-minute discovery call covers four things: the founder’s understanding of the problem in their own words, the founder’s understanding of who the user is and what success looks like, three to five questions the founder cannot yet answer well (data access, model latency, regulatory constraints), and a partner-led discussion of the engagement shape with a preliminary effort range. It produces a one-to-two-page scoping memo within 48 hours and a go/no-go on partner-fit from both sides.
How long should a paid pilot be?
One to two weeks for AI MVP work in 2026. Anything shorter cannot produce a meaningful eval; anything longer is a scoping engagement masquerading as a pilot. Budget $5K-$25K depending on data complexity and seniority. Deliverable: one runnable artifact, one measured eval baseline, one defensible main-engagement SOW. If the partner proposes four-plus weeks at $40K-plus, that is not a pilot — it is a paid scoping phase, which is a third (legitimate) procurement instrument with different uses.
Should I run paid pilots in parallel with multiple partners?
Almost never. Parallel paid pilots split your attention, signal a broken process, and triple the cash spend without tripling the information return. The disciplined path is sequential: discovery calls in parallel (no exclusivity implied), then a single paid pilot with the finalist. The rare exception is a high-runway founder with a very large eventual build ($500K+ SOW) where parallel pilots with two finalists may be justified to compare execution quality on the same problem statement.
What if the partner refuses to do a paid pilot?
That is information. Some partners refuse pilots as a matter of policy — usually because their margin model requires the full engagement. Others refuse because they fear a kill clause they cannot recover from. Either reason is legitimate from the partner’s side and disqualifying from your side if capability uncertainty is high. The exception: an established partner with multiple recent eval-protected case studies in your problem-neighborhood may legitimately offer the case studies as evidence in lieu of a pilot.
Is a “free discovery phase” the same as a discovery call?
No, and the terminology collision is the source of much market confusion. A “discovery phase” is usually a paid two-to-four-week engagement (typically $15K-$40K) that produces a PRD, an architecture decision record, and a SOW. A “discovery call” is the free 60-to-90-minute scoping conversation. Some agencies use “discovery call” to mean “free sales pitch” and “discovery phase” to mean “paid scoping engagement”; others reverse the convention. When you ask a partner what they mean by either term, you are testing their procurement maturity.
How is a paid pilot different from a paid scoping engagement?
Both are paid. The artifact differs. A paid pilot produces a runnable prototype and an eval baseline — code that runs end-to-end on real data, measured against a real threshold. A paid scoping engagement produces a PRD, an ADR, and a SOW — documents, not code. Pilots test capability. Scoping engagements test specification. The choice is determined by Property 1: if capability uncertainty is high, run a pilot; if specification uncertainty is high (you are not sure what to build, even if you trust models can build something), run a scoping engagement.
Can I do the hybrid sequence with a budget under $20K total?
Yes, with compression. Three discovery calls instead of five (one week, not two). A scoped one-week pilot at $6K-$10K instead of two weeks at $15K-$20K. The compressed sequence costs $6K-$10K total and produces less evidence than the full sequence, but more than either modality alone. Founders with very low runway should consider this compressed shape rather than skipping the pilot entirely — the 2026 capability-misjudgment risk is high enough that some falsification work is materially better than none.
How do I know if the discovery call was actually well-run?
Three test questions, applied 48 hours later. Did the memo restate your idea in the partner’s words, not yours? Did it name a risk you had not surfaced? Did the preliminary effort range fall within a defensible 50% band rather than a meaningless 5x band? A well-run call passes all three. A badly run one fails the first (“the memo is your words back at you”) or the second (“no risks named, only opportunities”) or the third (“$50K to $500K”).
Closing
The first decision in an idea-to-product engagement is not who to hire — it is which procurement instrument to use first. A free discovery call and a paid pilot are not cheap-and-expensive versions of the same thing. They produce different artifacts, answer different questions, and have different correct moments. The four-property rubric — capability uncertainty, partner-fit uncertainty, runway, urgency — points to one of three answers: discovery call only, paid pilot only, or the hybrid sequence. For most non-technical founders with a real AI idea and adequate runway, the answer is the hybrid sequence, costing five to six weeks and $10K-$20K, returning a defensible main-engagement SOW the founder believes in.
If you are at the point of choosing between a discovery call and a paid pilot for your idea, we run both modalities at SFAI Labs and can help you think through which one your situation actually calls for. We bring the rubric; you bring the idea, the runway, and the timeline. We tell you which artifact to start with — and if it is the discovery call, we run it; if it is the pilot, we scope it.
Arthur Wandzel