A first call with an AI scoping vendor is a 45-minute capability probe of the vendor, not a sales meeting. The vendor uses the time to qualify the budget and propose a follow-up. The founder’s job is the inverse: extract enough signal in 45 minutes to decide whether the vendor is worth $8K to $15K for a two-day scoping workshop within two weeks. That decision is made on eleven questions, each with a known-good answer shape and a known hand-wave shape — and on two walk-away signals that legitimately end the call inside 20 minutes.
Print this, take it into a call this afternoon, and score the vendor in real time. Companion reading: the idea-to-product manifesto, the eval-first build playbook, the 2-day workshop operating manual, the scoping services worksheet, and anatomy of a great AI agency kickoff.
What the first call is actually for
McKinsey and BCG report the same uncomfortable number: roughly seven in ten enterprise AI pilots never reach production. The largest single cause is not model quality. It is procurement of the wrong vendor for the wrong shape of work. A founder who walks in confused about what they are buying tends to leave with a $150K build proposal anchored to a slide-deck artefact list.
The first call has one job: determine whether the vendor’s senior engineering layer is real, eval-literate, and willing to write a no-go recommendation. Three small disqualifiers, each dispositive. The eleven questions below probe those three properties. The answer keys cluster on operational detail no sales rep memorises.
How to set up the call so the questions work
Twenty minutes of pre-work converts the call from a sales meeting into a diagnostic.
Write a one-paragraph product description. Three to five sentences naming the user, the workflow today, the failure mode that triggered the project, and what unambiguous success looks like. Send it 24 hours ahead. If the vendor cannot reference specifics from that paragraph by minute five, the senior engineer is not on the call.
Insist a senior engineer is on the call. Not a technical co-founder, not a solutions architect. A named senior engineer who will be in the room for the paid workshop, on the live call, for at least the back half.
Take the call in 45 minutes, not 60. Sales-stage rapport expands to fill any agenda. None of the eleven questions require an engineering vocabulary to ask.
The eleven questions
Each question is structured identically: the prompt, why it matters, what good sounds like, what hand-wave sounds like. Score pass, partial, or fail. Six or more passes is a green light to book the paid workshop. Three or more fails is an end-the-call signal.
Question 1 — “Who from your team would be in the room for the paid workshop?”
Why it matters. Workshops fail when staffed with facilitators and strategists. Real artefacts ship only when a senior engineer is in the room for the full duration and codes during the session. Senior engineering capacity is the binding constraint on AI delivery quality (per Stack Overflow’s 2024 Developer Survey).
What good sounds like. A named individual, tenure at the firm, prior experience with the model layer this build will use, and confirmation they would be in the room for all 16 hours.
What hand-wave sounds like. “We’d assign based on availability.” “Our team approach pairs a senior engineer with a project lead.” Anything that splits coding hours from the senior engineer’s calendar is a fail.
Question 2 — “How do you decide whether to recommend a no-go on a workshop?”
Why it matters. A vendor whose business model cannot tolerate a no-go cannot deliver an honest workshop. The procurement protection you are paying for is a written recommendation that says do not build.
What good sounds like. A specific frame for no-go: cost-per-query above 60% of expected ARPU, eval pass rate below a named threshold on the highest-severity task, or a domain today’s frontier models cannot reach. A claim of past no-go memos, ideally with a redacted example sent after the call.
What hand-wave sounds like. “We almost always find a path forward.” “We’d reframe the scope rather than say no.” Sales-pipeline language, not diagnostic.
Question 3 — “What does your eval set construction process look like in week one?”
Why it matters. AI features must be scoped against evals, not user stories. A vendor who treats evals as a marketing word produces documents with no acceptance criterion. The most diagnostic question on the list — non-technical founders need only hear whether specific terms surface.
What good sounds like. The vendor references inter-rater reliability, rubric design, severity-weighting per task, 30–50 starter cases expanding to 80–150 for a longer engagement, and a domain expert as the second grader. Anthropic and OpenAI’s public evaluation patterns are familiar reference points.
What hand-wave sounds like. “We run tests against the model.” “Our QA process catches issues before launch.” If the word rubric does not surface in three minutes, the vendor does not do evals.
Question 4 — “What would the five named files we get at the end of the workshop be?”
Why it matters. Named-file lists are the procurement contract. A workshop SOW that uses prose synonyms — “a roadmap”, “alignment on capabilities” — preserves vendor optionality on what ships.
What good sounds like. Five filenames with extensions: capability-probe.md, task-taxonomy.md, eval-set-v0.csv, cost-per-query.xlsx, go-no-go-memo.md. A one-sentence description of each. Confirmation you own the repository.
What hand-wave sounds like. “You’d get a comprehensive scoping document.” “Format depends on what we find.” If extensions are negotiated post-payment, the artefacts are negotiable post-payment.
Question 5 — “What’s a recent project where your cost-per-query model led to a different architecture than originally planned?”
Why it matters. Cost-per-query decides whether the product is shippable at gross margin. A vendor who only models it post-launch cannot protect you from the most common 2026 AI failure: a product that works but cannot make margin at scale.
What good sounds like. A specific story: “We were going to call Claude Opus 4.8 on every turn, modelled 60 cents per active user against a $5 ARPU, and routed to Sonnet 4.6 for the 80% of turns that did not need the bigger model.” Vendors named, sensitivity-tested.
What hand-wave sounds like. “We optimise for cost where we can.” Generic language without numbers means the vendor has not done this work.
Question 6 — “What’s the smallest architecture sketch you’d commit to at the end of the workshop?”
Why it matters. The architecture artefact must be small enough to be honest. A vendor who promises a full system diagram in two days is inflating output. A one-page list of named components — model call layer, fallback path, retrieval scaffolding if needed, logging requirement — is the honest 2-day artefact.
What good sounds like. A one-page document with five to eight named components, each with a one-line description and a note on whether it is in scope for v1. Honesty about what is not yet committed.
What hand-wave sounds like. “We’d deliver a full system design.” A real senior engineer knows what fits inside two days.
Question 7 — “How does the workshop output translate into a build proposal — and is the build a separate decision?”
Why it matters. A vendor whose workshop is structurally a sales meeting for their build practice has the wrong incentive at every fork. The five artefacts must be portable: any downstream builder should grade against the v0 eval set and the go-no-go memo.
What good sounds like. “The artefacts are yours. We’d be glad to bid on the build, and so should two other vendors. The eval set is the spec.”
What hand-wave sounds like. “Our workshop and build practices are integrated.” “We’d discount the workshop fee against a signed build.” Structural coupling is a fail.
Question 8 — “Who would grade the eval set with your engineer — and how do you handle disagreement between graders?”
Why it matters. A rubric scored by a single engineer produces an engineer’s view of correctness, not an operator’s. The domain expert in the room is the highest-impact founder commitment. Inter-rater reliability is the technical term — you should hear the vendor use it.
What good sounds like. “Your domain expert grades half the outputs, our engineer grades the other half, and we double-grade 10–15 cases to compute inter-rater agreement. Below 70% IRR means the rubric needs sharpening — itself a finding.”
What hand-wave sounds like. “Our team handles the scoring.” “Your team can review outputs after.” Single-grader scoring is not eval design.
Question 9 — “What’s the smallest engagement you’d sell us — and what does it not include?”
Why it matters. A vendor who cannot describe their smallest unit of sale is incentivised to upsell. The honest 2026 band for a 2-day workshop is $8K to $15K (typically $12K). A floor of $40K is a discovery week, not a workshop.
What good sounds like. “The 2-day workshop is $12K. It does not include a runnable prototype, a data audit, or a build proposal — those are the discovery week at $25K–$40K.” Clear pricing, clear exclusions, clear naming of the larger product.
What hand-wave sounds like. “Pricing depends on scope.” “Most engagements we run are in the $50K to $100K range.” A vendor who cannot quote the smallest unit at minute 30 is positioning for a larger sale.
Question 10 — “Show me one ground-truth example from a past project that the model failed on — and what you did about it.”
Why it matters. Eval-literate vendors maintain a mental library of failure modes on real data. A vendor who can rattle one off has done the work. The highest-signal probe on the list.
What good sounds like. A redacted example: “On a contract-review feature, the model failed on three cases with unusual indemnity language. We added a fallback to human review for indemnity sections.”
What hand-wave sounds like. “Models fail in lots of ways.” “We have processes for handling edge cases.” Abstraction means absence.
Question 11 — “What happens if your senior engineer is sick on Tuesday morning of the workshop?”
Why it matters. Operational continuity is procurement protection. A two-person bench with a named backup is the credible answer. A one-person bench means you absorb the vendor’s risk — and surfaces whether the firm hires senior engineers or sells access to a freelance pool.
What good sounds like. “Our backup for Maria is David — same tenure, same background. If Maria is out, we reschedule by 48 hours so David can read in. If both are out, we refund the deposit and re-book.”
What hand-wave sounds like. “We’d cross-staff.” “It’s never happened.” Without named individuals, the bench is theoretical.
Two walk-away signals
End the call inside the 45 minutes if either signal fires.
Walk-away signal 1: the senior engineer is not on the call. If 25 minutes have passed and the only voices are a sales lead and a project manager, the questions cannot produce real answers. Close politely and schedule a follow-up; if the vendor cannot produce that engineer inside a week, the firm’s senior engineering capacity is not real.
Walk-away signal 2: the word rubric is structurally absent after Question 3. If the vendor has not used rubric, eval set, or inter-rater by minute 20, the firm does not do eval-first work. Their output will be a user-story document and a roadmap deck — not procurement protection. Close the call: “Thanks — I’m focused on eval-bound scoping, and we may not be a fit.” Same pattern in stop scoping AI features in user stories, scope them in evals.
The 24 hours after the call
Within 24 hours, write a one-page memo.
Section 1 — Pass/partial/fail on the 11 questions. Six passes is green-light. Three fails is hard stop. Two to five fails warrants a follow-up call with the senior engineer and a re-score.
Section 2 — The cost band quoted. 2026 workshop pricing falls into three honest bands: $8K–$15K (2-day workshop), $25K–$40K (5-day discovery week), or $50K–$120K (multi-week scoping with prototype). Pricing between bands without a clear rationale is anchoring on what the vendor thinks you will pay. Full breakdown: scoping services worksheet.
Section 3 — The named individuals. Write down the senior engineer and the backup. The SOW must reference both by name. The 2-day workshop operating manual details the engineering layer.
Run the diagnostic on two more vendors before signing. If a vendor passes six of eleven and quotes a band consistent with the work, book the workshop. The next operational document is the kickoff — see anatomy of a great AI agency kickoff for Day 1.
Ready for the first call?
A 30-minute idea review with SFAI Labs is structured exactly like the diagnostic above — senior engineer in the room, eval-bound questions, written follow-up. No deck, no rapport-building, no upsell. If the call passes your scoring, book the workshop. If not, you walk away with a sharper diagnostic for the next vendor.
Book a 30-minute idea review →
FAQ
How long should the first call with an AI scoping vendor be?
Forty-five minutes. Sixty expands to rapport-building, which dilutes the diagnostic. Thirty is too short for the eleven questions plus answer keys.
What if a vendor refuses to put a senior engineer on the first call?
That is the first walk-away signal. A vendor unable to staff a senior engineer on a 45-minute call cannot staff one for a 16-hour paid workshop. Run the diagnostic on a different vendor.
How many vendors should I run this diagnostic on before signing?
Three is the minimum. One is observation, two is comparison, three is decision. Three 45-minute calls is two hours of total time against a $12K decision.
Should I ask about pricing on the first call?
Yes, directly, at minute 30 (Question 9). A vendor who refuses to quote their smallest unit of sale is positioning for an upsell.
The senior engineer is impressive but the pricing is double the typical band — now what?
Two honest possibilities. Either the vendor is selling a 5-day discovery week ($25K–$40K) and miscommunicated the unit, or anchoring on willingness-to-pay. Ask: “What does the $25K cover that the $12K version does not?” A clear answer keeps the option open.
Can I send the eleven questions to the vendor before the call?
You can, but the answers degrade. A vendor coached on the list sands the rough edges off the hand-wave phrases. Reserve the questions for the live call.
What if the vendor wants to run a paid discovery before answering most of these?
That is itself an answer — fail on Question 9. A vendor who cannot describe named files, eval process, or pricing band without a paid engagement is selling discovery as the diagnostic.
How does this first call differ from a paid pilot?
The first call is the procurement screen; the paid pilot is the work. The first call decides whether to commit $8K–$15K to a 2-day workshop. A paid pilot is a fixed-fee build of a narrow scope with eval-bound acceptance, often $25K–$60K. Decision tree: discovery call vs. paid pilot.
Arthur Wandzel