Every non-technical founder hiring an AI partner faces the same first question on day one: a conversation, a small paid project, or a full build? Most agency websites blur the three. They are not the same instrument and they do not solve the same problem. A discovery call is a free 60-to-90-minute conversation. A paid pilot is a one-to-two-week engagement that produces runnable code and an eval baseline. A full engagement is a six-to-twelve-week build that ships an MVP. Picking the wrong starting shape costs two to four weeks of calendar or ten to thirty thousand dollars. This explainer names the three shapes precisely, gives the default sequence, and names the two legitimate cases for skipping a step.
It builds on the founder-AI-partner operating manual and the broader idea-to-product manifesto. Companion reading: what is an AI development partnership, how a 12-week engagement actually works, and the deeper rubric in discovery call vs paid pilot.
Why the three shapes get confused
Half the AI agencies in 2026 brand their two-to-four-week paid scoping engagement as a “discovery phase” and their unpaid sales call as a “discovery call”. The other half use “pilot” to mean anything from a one-hour demo to an eight-week proof-of-concept. The vocabulary evolved slower than the procurement instruments.
A discovery call is the free, scoped conversation. A paid pilot is the small, fixed-fee engagement that produces a runnable artifact. A full engagement is the months-long build that ships an MVP. They differ in fee, duration, founder commitment, and — most importantly — in the falsifiable artifact they leave behind. If the partner cannot name the artifact each one produces, you are reading their sales deck. McKinsey’s The state of AI in 2025 reports 78% of organisations now use AI in at least one business function — demand that pulled hundreds of new boutique partners into the market and produced the vocabulary mess.
The discovery call, defined
A discovery call is a 60-to-90-minute structured conversation between the founder and a senior member of the candidate partner team — the principal who would lead the engagement, not a salesperson. It is free, scheduled within a week of first contact, and produces three artifacts when run well:
- A scoping memo, one to two pages, written by the partner within 48 hours. It restates the idea in the partner’s words, names three to five required capabilities, and flags risks the partner sees from outside.
- A preliminary effort range — not a quote. A weeks-and-team-shape sketch: “an eight-to-twelve-week engagement with a senior plus a mid, roughly $120K-$180K depending on data-access work.”
- A mutual go/no-go on partner-fit. Either side can decline without cost.
What a discovery call does not produce: working code, an eval baseline, a tested data assumption, or evidence of partner capability beyond the conversation. It confirms partner identity and engagement shape; it does not confirm partner ability or idea feasibility. It is the right artifact when the founder is still comparing partners or still defining what they would hire someone to do.
The paid pilot, defined
A paid pilot is a one-to-two-week engagement with a fixed fee — typically $10K to $25K in 2026 — that produces three artifacts:
- A runnable prototype — not a demo or a deck. Code that runs end-to-end on the founder’s actual data (or a representative sample), narrowly scoped but real.
- An eval baseline — a measured pass rate on a small (50-to-150 item) eval set built during the pilot. Without an eval, a pilot is theater. Anthropic’s public Building effective agents guidance treats the eval as the load-bearing artifact of any AI system.
- A defensible Statement of Work for the main engagement — a fixed-fee, milestone-billed SOW whose effort estimate, eval threshold, team shape, and timeline trace to what the pilot found.
A pilot tests one critical assumption end-to-end at small scale — usually the riskiest capability. It is a falsification instrument: its job is to surface evidence that lets both sides write a defensible SOW or walk away cleanly. It is the right artifact when the founder trusts the partner and the open question is feasibility (capability, data, or cost). It is the wrong artifact when the open question is partner-fit, because pilots imply exclusivity. The structural argument is made in depth in why your AI agency should run a paid pilot before the main contract.
The full engagement, defined
A full engagement is a six-to-twelve-week fixed-fee build, typically $130K to $200K for a single AI MVP in 2026, that produces six artifacts:
- A versioned PRD with explicit non-goals.
- A versioned eval set — 50 to 200 items with a launch-threshold pass rate the team commits to.
- A production-ready prototype running end-to-end on real data, hardened against the top failure modes the eval surfaces.
- A deployed MVP under real authentication, real billing if relevant, and real logs.
- A two-week on-call window while the founder onboards a long-term operator.
- A handoff package — repo access, runbook, eval suite, cost dashboard, and a one-page operating note.
A six-week engagement ends before the eval threshold can be defended; a six-month engagement is a team build wearing the costume of an MVP. The defensible shape in 2026 is twelve weeks; the chronology is in how an idea-to-product engagement actually works, week-by-week. The full engagement is the right artifact when the question has shifted from “can this work?” to “what does the shipped product look like?”
Side-by-side comparison
| Discovery call | Paid pilot | Full engagement | |
|---|---|---|---|
| Fee | Free | $10K–$25K | $130K–$200K |
| Duration | 60–90 min + 48-hour memo | 1–2 weeks | 6–12 weeks |
| Founder time | 2 hours | 4–8 hours | 30–60 hours |
| Output artifact | Scoping memo + effort range | Runnable prototype + eval baseline + defensible SOW | Deployed MVP + eval suite + handoff |
| Falsifiability test | Did the memo name three capabilities and three risks? | Did the eval pass-rate get measured against a real threshold? | Is the MVP serving real users under real auth, with real logs? |
| Answers the question | Who is this partner, and what shape would the build take? | Can the partner actually do the work, and is the data sufficient? | What does the shipped product look like, and what does it cost to operate? |
| Right when | You are still partner-shopping | You trust the partner; capability or data is uncertain | Capability is known; you are ready to ship |
The default sequence
The default sequence is discovery call → paid pilot → full engagement. Each step de-risks the next: the discovery call resolves partner-fit, the pilot resolves capability and data feasibility, the full engagement resolves shipping. Total founder cost: roughly $15K-$25K of pilot fees, four to six weeks of calendar before the main contract starts, and twenty hours of founder time across the first two stages.
Carta’s State of Private Markets 2025 puts the pre-seed median raise for non-technical-founder teams at $400K-$700K; in that range, a $20K pilot is two to three percent of the round and a $150K full engagement is twenty to thirty percent. Spending three percent to de-risk thirty percent is sound math.
When to skip the pilot
Skip the paid pilot and go straight from discovery to full engagement in three cases:
- The capability is well-understood. RAG over a defined corpus, structured extraction from semi-structured text, a routine agent loop with two or three named tools — these are not capability-uncertain problems in 2026. Public reference implementations exist; the eval baseline is predictable within five points.
- The partner has a near-identical track record. If the partner has shipped two or more production systems with the same model class, data shape, and user-facing surface, a pilot tells you nothing they cannot tell you in writing. Ask for the eval suites and post-launch dashboards from prior engagements.
- The budget is too constrained for a real pilot. A pilot under $10K is a demo with a fee. If the founder cannot allocate $15K-$25K, the honest move is a longer discovery call with two memos, or a smaller-scoped full engagement (six weeks, $75K-$95K) that builds the eval set in week one.
Skip only when at least two of the three conditions hold.
When to skip the discovery call
Skip the discovery call and go straight to a paid pilot in two cases:
- A strong referral and a clear written brief exist. If a trusted operator has worked with the partner and you have a one-pager describing the idea, the data, and the success metric, the discovery call adds little. Use the first hour of the pilot as the kickoff conversation.
- You have already run a discovery call with this partner on a prior project. The partner-fit question is resolved. Re-open with a pilot scoped to the new question.
The cost of skipping the discovery call when you should not have is $15K-$25K paid to a partner you discover mid-work to be the wrong fit. The cost of holding a discovery call you did not need is two hours. Skip only when the partner-fit signal is already strong.
Cost of starting with the wrong shape
The wrong-instrument cost shows up in three patterns:
- Full engagement when you should have piloted. Weeks four through eight of a twelve-week build get spent discovering data problems or capability gaps a $20K pilot would have surfaced in week one. Expected loss: $40K-$80K of engagement budget rebuilt as rework or descoped features.
- Pilot when you should have done a discovery call. $15K-$25K paid to a partner who turns out mid-pilot to be the wrong fit, plus a two-week calendar block.
- Stopping at the discovery call when you should have piloted. Signing the full engagement on the strength of a memo, then discovering capability gaps in week three. Same loss as pattern one.
Starting too far up the sequence costs ten to fifty thousand dollars and weeks of calendar; staying too far down costs only an unnecessary fee. Bias toward the smaller instrument when the next-step signal is mixed.
Two worked examples
Founder A — technical co-founder, $1.4M raised, structured-extraction product over a known document corpus. Capability is well-understood; the team can read the partner’s code. Sequence: discovery call (40 minutes, partner-fit confirmed quickly) → straight to a ten-week, $145K full engagement. Pilot skipped on capability grounds. Founder time to commitment: 8 hours.
Founder B — non-technical founder, $550K raised, multi-step agent for a regulated workflow. Capability uncertainty is high; partner-fit uncertainty is moderate. Sequence: discovery calls with three partners (six hours) → pick the strongest two → run a $20K paid pilot with the top pick (two weeks, eval baseline measured at 64% on hard items) → use the pilot to write a defensible SOW for a twelve-week, $175K full engagement. Founder time to commitment: 28 hours.
Both end up in a full engagement. The difference is the procurement instruments they spent on the way: $0 for Founder A, $20K for Founder B. Both decisions match the founder-side uncertainty profile.
FAQ
Is the discovery call really free?
Yes, for any reputable partner. The partner treats the call as their unpaid customer-acquisition step. If a partner charges for a first conversation, they are mispricing their sales channel — be cautious.
What is the difference between a paid pilot and a proof-of-concept (POC)?
In 2026 the terms are roughly interchangeable. The caveat: vendors sometimes use POC to mean “we will build something in four to six weeks for $40K-$60K”. That longer shape is closer to a small full engagement. If the engagement is over three weeks, it is not a pilot — it is a constrained build.
Can I run paid pilots with two partners in parallel?
Generally no. Pilots imply the partner stops chasing other work for two weeks; parallel pilots signal you are not committed, and both partners deprioritise. The narrow exception is a structured two-pilot tournament with a clear winner-takes-the-full-engagement commitment.
How much should the discovery-call memo say?
One to two pages. It should restate your idea in the partner’s words, name three to five required capabilities, flag three to five risks, and give an effort range as a band. If the memo does not arrive within 48 hours, the partner is either disorganised or treating the call as routine.
What happens if the pilot eval baseline misses the launch threshold?
That is the point of the pilot. Two defensible responses: (a) the partner scopes the eval-failure work into the full engagement (more weeks, more fee, defensible SOW); or (b) both sides walk away cleanly, and the pilot fee is the cost of the discovery. The wrong outcome is declaring the pilot a success despite the gap.
Can I do a full engagement without a discovery call or a pilot?
Yes, but only with a strong referral, a partner with a near-identical track record, and a written brief you can defend. Even then, week one of the full engagement should run like a discovery call — same memo, same scoping artifact.
Who runs the discovery call from the partner side?
The principal who would lead your engagement. Not a salesperson, not an account manager. If the first conversation is with anyone other than the operator, the memo will be a sales document.
How does this sequence fit into a 12-week engagement?
The discovery call and pilot are the first three weeks of activity before the full engagement starts; the 12-week clock starts at the signed SOW. Some partners credit 50-100% of the pilot fee against the first invoice if you proceed within four weeks — ask for it explicitly.
What if the partner refuses to run a free discovery call?
Walk. A partner who charges for a first conversation is either a top-five-percent expert with strong inbound demand (rare, and they will say so) or has miscalibrated their sales funnel.
What is the SFAI Labs default?
Default to discovery call. Move to paid pilot when capability or data uncertainty is high. Move to full engagement when capability and partner-fit are both resolved. Spend two to three percent of the round on procurement instruments to de-risk thirty percent on the build.
Ready to start? A free 60-to-90-minute discovery call with SFAI Labs produces a written scoping memo within 48 hours and a defensible effort range. Book a discovery call.
Arthur Wandzel