By the time you are about to sign with an AI MVP partner, marketing materials are useless. The decks have been read, the case studies polished, the discovery call went well. The only signal left between you and a $150K-and-twelve-weeks commitment is uncensored testimony from a founder who already wrote the same cheque.
Most founders waste that signal. They ask “how was working with them?” and “would you hire them again?” — questions designed to get a polite yes — then sign because nobody said no. That is not a reference check; that is permission-seeking dressed as diligence.
This is a script. Nine questions, in order, with verbatim wording, the reason each matters, what a good answer sounds like, and what a hand-wave sounds like. Read it onto the call. Take notes verbatim. Run it twice — two references minimum, three for any engagement above six figures.
Decision Scope
This article is an editorial decision framework, not legal, financial, security, or accounting advice. Treat numeric examples as illustrative planning heuristics; validate against your own contracts and budgets before acting. For the week-by-week founder-side context, see the founder–AI partner operating manual.
Before the call: the frame that gets honest answers
Open every call with the same thirty-second framing:
“I am evaluating Partner X for an AI MVP build of similar shape and budget. I do not need a glowing review. I will not repeat anything you say back to them. What I need is the operational truth — what shipped, what broke, what the bill actually was, and what you would do differently. Where you tell me they were great, that is signal. Where you hesitate, that is also signal. I am taking notes; please be specific.”
That paragraph improves answer quality more than any single question on the list. References want to be honest; they need permission and a frame that is not “would you recommend this vendor.”
Order matters: eval discipline first because it is the highest-information question for an AI build; behavioural tests in the middle; the counterfactual last so it does not anchor the conversation.
Question 1: Eval discipline
“Walk me through how they measured whether the agent was actually working.”
Eval discipline is the strongest single procurement signal for an AI MVP partner. A reference who cannot describe the eval methodology from memory worked with a partner that did not have one — they had testing, which is not the same thing. See the eval-first build playbook for the underlying practice.
Good: “They built a Promptfoo suite, around 200 cases — happy path, regulatory edges, adversarial prompts our security team submitted, and a hand-curated set of real user phrasings from beta. Pass threshold was 91% on the regulatory subset, negotiated in week three. CI ran it on every PR. We caught a model-drift regression in month four because of that pipeline.” Named framework, case count, categories, threshold, workflow, real catch.
Hand-wave: “Yeah, they tested everything thoroughly before launch. The QA was really solid.” Testing is not evals. “Solid QA” is what a reference says when the eval suite did not exist — only smoke tests written the week before launch.
Question 2: On-time delivery
“Did the original timeline hold? If not, what slipped and how was it handled?”
Schedule slip in an AI MVP is common — McKinsey’s 2024 state-of-AI work and BCG’s value-realisation studies both put production AI delivery slip rates well above traditional software. The question is not whether the timeline slipped; it is how the partner handled it. Honest renegotiation is signal. Silent overrun is anti-signal.
Good: “Original SOW was twelve weeks. We hit week eight and their lead engineer ran a forty-minute review — said the eval set was tighter than the original spec implied and they wanted a two-week extension to hit threshold. They gave us three options: ship at the original threshold, extend by two weeks at no cost, or descope a secondary flow. We chose extend. Final delivery was week fourteen.” Specific renegotiation, multiple options, honest trade-off.
Hand-wave: “Honestly? It mostly held. There were a couple of weeks at the end where things were a little tight, but they got there.” Polite language for an engagement that went two months over without a written renegotiation.
Question 3: Scope-creep handling
“When you asked for something outside the original scope, what was the actual conversation?”
Scope-creep handling is a behavioural reveal — most contracts have a change-order clause; few founders ever see it executed. How the partner responds to the first out-of-scope ask tells you whether the engagement is governed by judgement and real paperwork or by yes-saying that ends in invoice surprises.
Good: “We asked for a second agent flow in week five — completely outside the SOW. Their CTO joined a call the next morning, said it was a two-week add, gave us a fixed change-order quote, and was explicit it would push the eval work or launch by one week. We took the trade-off. The change order arrived within forty-eight hours.” Named process, named owner, named trade-off, written paperwork.
Hand-wave: “Oh, they were so flexible. Whatever we asked, they just made it happen.” A partner that absorbs unlimited scope without paperwork is either eating margin silently — which means eval work or documentation is being descoped — or building toward a final-invoice surprise.
Question 4: Handoff quality
“If you had to onboard a new engineer to the system tomorrow, how long would it take them to be productive?”
Handoff quality determines whether you own an asset or a black box at week twelve. Documentation, code legibility, eval-suite readability, and runbook completeness collapse into one observable proxy: engineer-onboarding time. Three days means the system was built to be operated. Three weeks means it was built to be billable.
Good: “We onboarded a new senior engineer in month four. She was running an eval review and writing prompt changes by day three. The README pointed her at four files, the eval set was self-documenting, and the partner’s lead engineer did a one-hour video walkthrough we still use.” Specific time, specific artefacts, reusable walkthrough.
Hand-wave: “Honestly we have not had to bring anyone new in. They documented everything, I think.” “I think” is the answer. If the founder cannot point at one file or one walkthrough, the handoff was theatrical.
Question 5: On-call honoring
“When something broke at 11pm on a Friday, what actually happened?”
On-call is where SOWs and reality diverge most cleanly. Every SOW has a support clause; the operational question is whether the partner answered the page or whether it bounced to a Monday-morning ticket. Most production AI systems hit a failure inside the first ninety days. “Never came up” is rarely the truth.
Good: “It happened twice in the first ninety days. First time, the retrieval index had stale embeddings — we paged at 10:30pm, their senior engineer was online inside fifteen minutes, root-cause inside an hour, hotfix shipped by 1am. Second time, a model provider rate-limited us, they escalated to the provider, workaround live before morning.” Named time, named engineer, named root cause, named fix.
Hand-wave: “It honestly never really came up. The system was super stable.” Press once: “and were you the one fixing it, or them?” That follow-up is usually where the truth surfaces.
Question 6: IP terms in practice
“Did you ever need to invoke any of the IP clauses or weights-ownership terms? How did they respond?”
Most founders read the IP clauses, sign, and never look at them again. The interesting question is what happens when the founder needs to invoke them — switching providers, exporting training data, or moving code out of the partner’s repo. See the IP and weights conversation every founder should have for the day-one context.
Good: “We forked the repo to our own GitHub org in month six — it was an option in the contract. They ran the migration with us in an afternoon, transferred the eval data, signed the receipt the same week. No friction.” Specific event, specific timeline, specific paperwork.
Hand-wave: “We never had to invoke them, I don’t think. The terms were fine.” “Fine” terms that have never been invoked are theoretical terms. Press: “If you had wanted to switch model providers tomorrow, would you have had the weights and the eval set to do it?”
Question 7: Post-launch support
“Six months after launch, how often do you actually talk to them?”
Most AI MVP engagements are sold as build-and-handoff, but production AI systems do not stay stable on their own — model providers update, data shifts, edge cases emerge. See the 30-day post-launch period explained for the structural side.
Good: “Monthly check-in call, async Slack any time. Their CTO flagged a Claude Opus 4.8 update to us two weeks before it shipped — said our prompt structure would need a tweak, gave us the recalibration eval, offered a four-hour engagement to handle it. We took it.” Named cadence, named channel, proactive flag on a real model update.
Hand-wave: “They are super responsive if I email them.” Reactive only. The partner is not watching the system; the founder is. That may match your expectations, but it is not what the SOW probably implied.
Question 8: Hidden costs
“What did you end up paying for that was not in the original SOW?”
The SOW is the floor, not the ceiling. Founders routinely find inference bills, third-party API quotas, observability tooling, and post-launch retainer hours sitting outside the original quote.
Good: “Inference came to about $4,200 a month at steady state, on our own Anthropic account — they did not mark it up. We added Helicone for observability — that was $200 a month. The eval-suite expansion in month two was a written $8K change order. Apart from that, the SOW held.” Named line items, named dollar figures, named change orders. No mark-up on tokens is a strong signal — see the AI agency manifesto on why no-token-arbitrage is a category test.
Hand-wave: “It was within budget, more or less. I would have to check the invoices.” Press: “Roughly, were you 10% over, 30%, more?”
Question 9: Would you sign again
“Knowing what you know now, would you hire them again? And if so, for what kind of project specifically?”
Generic ref-check guides ask “would you hire them again” as a yes-or-no — which gets a yes-or-no, usually yes. The sharper version is the second half: “for what kind of project specifically?” That forces the reference to be honest about fit, not just quality.
Good: “Yes — for a focused agentic build, six to twelve weeks, where eval discipline is the main value. I would not hire them for a multi-month strategy engagement or a pure UI-heavy product where the AI is a small piece. They are sharpest when the agent is the product.” Specific fit, specific anti-fit.
Hand-wave: “Yes, definitely. They are great. I would recommend them to anyone.” Press: “What would you not hire them for?” If they cannot name one anti-fit, the recommendation is decorative.
How to actually book the call
The partner books these calls — that is part of the procurement protocol. If they refuse, refuse to sign.
- Two minimum, three above $100K. Two gives you a sanity check; three lets you weight any outlier.
- Different domains, different sizes. Ask for references from at least two adjacent verticals or two engagement sizes.
- Thirty minutes per call. Block thirty, end at twenty-seven, leave three for the reference to ask you anything.
- Audio only, no partner rep on the line. If the partner asks to join, decline politely.
- Note template: one column per question, one row for verbatim quote, one row for signal/anti-signal, one row for follow-up.
- One-line thank-you afterward. Reference calls are a small founder community; you are building a network as well as diligence.
For pre-call vetting, see the founder-friendly AI partner checklist and what to ask in a discovery call with an AI MVP partner. For the broader 11-question ToFu sibling of this script, see the AI agency reference call.
What to do with red flags
One hand-wave on one question is data, not a verdict. Two across different questions is a pattern. Three is a procurement decision.
The two questions you cannot afford a hand-wave on are eval discipline and on-call honoring — the operational floor for an AI MVP partner. If a reference cannot describe the eval suite, or names a single after-hours incident the partner missed, that is not a soft signal; that is a no-go.
The other questions are weighted by your own constraints. A founder with a hard regulatory deadline weights on-time delivery higher; a founder with sensitive data weights IP-in-practice higher.
If two of nine answers are red flags, do one more call. If three of nine across both references are red flags, walk. There are other partners.
If you have run the script, the calls are clean, and the diligence is done — book the next step. We run a thirty-minute idea review where we will read your reference notes alongside the partner shortlist and tell you, honestly, whether the build is shaped the way the partners you are evaluating actually ship. Book the idea review. For the broader context behind this vetting workflow, see the idea-to-product manifesto.
FAQ
How many references should I ask for?
Two minimum, three above $100K. Ask for at least one reference whose engagement is now six or more months post-launch — that one gives you the post-launch-support truth recent references cannot.
What if the partner refuses to share references?
Refuse to sign. A partner that cannot produce two willing reference clients is either too new — fine, but ask for a discounted paid pilot instead of a full engagement — or is hiding something.
Can I find references the partner did not introduce me to?
Yes, and you should. LinkedIn searches on the partner’s engineers surface past clients in their work history. A Google search of "powered by" + partnername often turns up case studies the partner did not flag. A quiet outreach to one founder the partner did not pick is the highest-information call on the list.
Should I do reference calls before or after the SOW is signed?
Before. Reference calls after signing are confirmation theatre. Block them between the discovery call and the SOW; treat them as the final gate before paper, not as comfort after it.
How long should each call be?
Thirty minutes, ending at twenty-seven. Nine questions plus opener plus closer fits cleanly into the half-hour if you keep the conversation moving.
What if every reference is glowing?
Either the partner is genuinely excellent, the references are coached, or you asked the wrong questions. The script above is designed to make coaching visible — references who hand-wave on eval discipline, on-call, and hidden costs are giving you the answer either way. If three references all sound the same scripted notes, ask the partner for one reference whose engagement was difficult. Their response to that ask is also signal.
What does an outdated AI model reference in a case study tell me?
Stale model names — older GPT-4 variants, Claude 3.5, Gemini 1.5 — mean the partner is not refreshing their public proof. Current SOTA references are GPT-5, Claude Opus 4.8 and Sonnet 4.6, and Gemini 2.5. A soft signal that the partner’s bench is not current.
Should I ask the reference about price?
Yes — question 8 is the price question, and it works because it asks about variance, not the headline number. The headline number is on the SOW. The variance is in the invoices.
What is the single most important question on the list?
Question 1 — eval discipline. It tests the partner’s actual engineering practice. A reference who describes the eval suite back to you from memory has worked with a partner that builds AI properly. A hand-wave on it alone is enough to put the engagement at risk.
How do I weight the answers across two or three calls?
Build a 9-by-3 matrix — nine questions, one column per reference, signal or hand-wave per cell. Any question where two of three references hand-wave is a structural issue with the partner. Compare across references; the partner is the constant.
Arthur Wandzel