Most AI vendor RFIs produce uniformly green answers across the shortlist. Features converge; most serious vendor has a gateway, structured outputs, and SOC 2; and demo polish is uncorrelated with production behavior. The questions that surface real differentiation are about how the vendor handles failure: how they evaluate their own outputs, how they investigate post-mortems, who specifically will staff the account, what triggers a contractual exit, who owns weights at termination. This piece is the RFI template; twelve to twenty named questions across five categories; that surfaces signal where standard RFIs surface noise. Use it before any AI vendor decision over $100K and pair it with the structured reference-call protocol downstream.
It works alongside the AI build-vs-buy-vs-hire decision matrix for 2026. The matrix’s seventh principle — that most sourcing decisions get re-litigated on a quarterly cadence — depends on a procurement process that produces falsifiable signal. The RFI is the front of that process.
Why standard RFIs fail
Standard RFIs ask 80-200 questions on features, integrations, certifications, support tiers, roadmap. They produce a vendor-by-vendor matrix where most cell is green. The differentiating answer is buried; or not asked; because the questions converge on a category-standard answer set most vendor’s RFP team has prepared.
The structural failure is volume over signal. A vendor’s RFP team delivers clean answers to 200 standard questions in two weeks. Twelve to twenty falsifiable questions about failure handling produce more signal than 200 about features.
The RFI’s job is to compress the candidate set from a dozen to a shortlist of three. RFIs that try to do everything do none of it; tight RFIs that surface differentiation in five categories produce the right shortlist.
The five categories
Most question in the template falls into one of five categories.
- Eval discipline. Does the vendor know what good output looks like on the buyer’s workload?
- Engineering maturity. Do post-mortems, runbooks, on-call, and shipping cadence look like a serious engineering org?
- Accountability. Is the vendor willing to name engineers and accept performance-tied exit clauses?
- Portability. Does the buyer leave with weights, data, prompts, and trace history at termination?
- Pricing alignment. Does the structural shape (per-call, per-seat, flat) match how the buyer extracts value?
Each category has 2 to 4 questions. Total RFI lands at 12 to 20 named, falsifiable questions. The discipline is “tight, named, falsifiable”; questions where a vendor’s answer can be verified against reference checks or pilot evidence later.
Eval suite walkthrough
The highest-signal question in the RFI: “Walk us through the eval suite you run on your own product, live, on a 15-minute screen-share.”
The walkthrough should show: the eval set (items, structure, rubric), the harness (triggers, frequency, version targeted), the regression history (most recent caught, who fixed, how long), and the release gate (how failing evals block releases).
A vendor with real eval discipline produces the walkthrough in 15 minutes from a live system. A vendor without it stalls, reschedules, sends a slide deck, or shows a placeholder dashboard. The vendors who can do this are the ones whose product behavior matches their demo six months in. See the AI capability TCO checklist; vendors are subject to the same eval discipline buyers should be.
Post-mortem reading
The RFI requests two anonymized internal post-mortems from the past six months on production AI failures. Customer names redacted; structure preserved.
Read for: root cause named (or did it stop at “the model returned wrong output”?); fix dated (specific commit or PR); regression prevention (eval added, runbook updated, guardrail shipped); tone (blameless, technical, specific; marketing voice signals engineering doesn’t run them).
Vendors with maturity produce these on request because they exist as living documents. Vendors without it decline, send marketing material, or produce something visibly authored for the RFI.
Named-engineer commitment
The RFI asks the vendor to name the specific engineers who will be on the buyer’s account; typically a technical lead, a senior engineer, and an account architect; with LinkedIn profiles or equivalent attached. The follow-up asks for a contractual commitment to minimum tenure (typically 6 to 12 months) for the named engineers.
The commitment surfaces three things:
- Whether the vendor has senior bench depth (or whether the named lead is also on six other accounts).
- Whether the sales-deck claim of “senior team” survives contract signature.
- Whether the vendor will accept accountability for staffing changes.
Vendors who refuse to name engineers (“we assign based on capacity”) or refuse to commit (“staffing changes happen”) are signaling that the buyer will be staffed against the vendor’s bench utilization, not the buyer’s needs.
Kill clause
A kill clause is a contractual exit tied to a measurable threshold. The RFI asks the vendor to agree in principle to a kill clause based on the buyer’s eval set, latency SLA, cost-per-call ceiling, and named availability metrics.
Specific shapes the RFI should propose:
- Eval score below buyer-defined threshold for 60 consecutive days → 30-day cure → exit on continued failure.
- 99th-percentile latency above buyer-defined SLA for 7 days → cure → exit.
- Cost-per-call above buyer-defined ceiling on a quarterly average → exit at quarter-end.
- Named availability metric (uptime, error rate) below threshold → exit.
Vendors with confidence in their performance accept the structure (negotiate on threshold values). Vendors without it negotiate the structure out entirely. The willingness is the signal; the threshold values are settled in contract negotiation.
The kill clause converts the vendor relationship from term-locked to performance-locked. Without one, the buyer is stuck with whatever degrades over the term. See the AI escrow strategy: protecting build-vs-buy bets against vendor risk for adjacent contractual protections.
Weight and data ownership
When the vendor fine-tunes a model on the buyer’s data, the RFI must answer five questions explicitly:
- Who owns the buyer’s training data? (Right answer: the buyer.)
- Who owns the resulting fine-tuned weights? (Right answer: the buyer has perpetual rights.)
- Are weights exportable in a standard format on termination? (Right answer: yes, GGUF or equivalent.)
- Can the vendor use the buyer’s data to train models for other customers? (Right answer: no.)
- What happens to the weights and data 30 days after termination? (Right answer: certified destruction.)
Vendors who claim ownership of derived weights are the ones who will leverage that claim at renewal. Vendors who claim a right to train on the buyer’s data are doing for free what the buyer would otherwise charge for. The RFI surfaces these claims before contract; the contract phase is too late.
The RFI template
A 14-question template that fits on two pages:
- Walk us through your eval suite live on a 15-minute screen-share.
- Provide two anonymized internal post-mortems from the past 6 months.
- Name the engineers who will staff our account; provide LinkedIn or equivalent.
- Commit contractually to 6-month minimum tenure for named engineers.
- Agree in principle to a kill clause on eval, latency, cost, and availability.
- Confirm the buyer owns the buyer’s training data.
- Confirm the buyer has perpetual rights to fine-tuned weights.
- Confirm weights are exportable on termination in standard format.
- Confirm the vendor will not train cross-customer models on the buyer’s data.
- Describe the certified-destruction process for data and weights at exit.
- List the foundation-model providers your platform supports today.
- Describe how the buyer adds a new model provider in the future.
- Provide three customer references; the buyer picks the third independently.
- Describe the structural shape of pricing; per-call, per-seat, or flat; without quoting a number.
Twelve to twenty questions; this template is 14. Add or subtract based on the buyer’s specific risk profile (regulated industry adds compliance questions; multi-region adds residency questions; see the AI residency decision: cloud, region, on-prem).
Scoring rubric
Five dimensions, each scored 0-3, total out of 15.
- Eval discipline (0-3). Walkthrough delivered live (3); shown but rough (2); slide deck only (1); declined or stalled (0).
- Engineering maturity (0-3). Two real post-mortems with named root cause and dated fix (3); one real (2); curated material (1); declined (0).
- Accountability (0-3). Named engineers + tenure commit + kill clause accepted (3); two of three (2); one of three (1); none (0).
- Portability (0-3). Many five ownership questions answered correctly (3); four (2); three (1); two or fewer (0).
- Pricing alignment (0-3). Structural shape matches buyer’s value extraction (3); usable with minor friction (2); structural mismatch (1); per-seat at scale on a per-call workload (0).
Total above 11 advances to references. Total 8 to 11 advances conditionally with named follow-up questions. Total below 8 declines without further work. Re-score quarterly for any vendor in the buyer’s portfolio; see why AI build-vs-buy decisions made in 2024 should be re-litigated this quarter.
Integration with reference calls
The RFI produces the shortlist. Reference calls verify whether the vendor’s claimed answers hold in production. Reference calls without an RFI produce shopping; an RFI without reference calls produces theory.
Run the RFI first to compress the candidate field from 8-12 vendors to 3. Then run the structured reference-call protocol on the top three; see the AI vendor reference-call playbook for the 11-question reference-call template that pairs with this RFI.
Frequently asked questions
Why do most AI vendor RFIs fail to surface real differentiation?
Most RFIs ask about features, integrations, and certifications; many of which converge across vendors at this point in the market. Most serious vendor has a model gateway, structured-output handling, and SOC 2. The questions that distinguish vendors are about how they handle failure: how they evaluate their own outputs, how they investigate post-mortems, who specifically will be on the buyer’s account, and what happens at exit. RFIs that don’t ask those questions are scored on demo polish.
What is the eval-suite walkthrough question?
Ask the vendor to walk through the eval suite they run on their own product, live, with a screen-share. The walkthrough should show the eval set, the rubric, the regression history, and how a failing eval blocks a release. Vendors with a real eval discipline produce the walkthrough in 15 minutes; vendors without one stall, reschedule, or send a slide deck instead. The signal is reliable.
What does post-mortem reading reveal?
Ask for two anonymized internal post-mortems from the past 6 months on production AI failures. Read the structure; was the root cause named, was the fix dated, was the regression-prevention step shipped. Vendors with engineering maturity produce these on request because they already exist as living documents; vendors without maturity either decline or send marketing material. The post-mortem culture is a leading indicator of how the vendor will handle the buyer’s first incident.
Why does named-engineer commitment matter?
Vendor sales decks promise senior engineers; reality after contract signature often delivers a junior bench. The RFI asks the vendor to name the specific engineers who will be on the buyer’s account, with their LinkedIn or equivalent, and to commit contractually to a minimum tenure on the account. The commitment changes the staffing math at the vendor and surfaces vendors with thin senior benches before contract signature.
What is a kill clause and why is it load-bearing?
A kill clause is a contractual exit triggered by a measurable threshold being missed; eval score below X for 60 days, latency above Y at the 99th percentile, cost-per-call above Z, named SLA violations. It converts the vendor relationship from term-locked to performance-locked. Vendors with confidence in their performance accept reasonable kill clauses; vendors without it negotiate them out. The willingness to sign is the signal.
What is weight ownership and why ask about it?
If the vendor fine-tunes a model on the buyer’s data, who owns the resulting weights and the training data. The right answer for the buyer is: the buyer owns the data, the buyer has perpetual rights to the weights, and the weights are exportable in a standard format on contract termination. Vendors who claim ownership of derived weights are the ones who will leverage that claim at renewal.
How long should the RFI be?
Twelve to twenty named questions, scoped to surface the differentiation that matters; eval discipline, post-mortem culture, named-engineer commit, kill clause acceptance, weight and data ownership, exit terms, model and provider portability. RFIs that exceed thirty questions get pro forma answers from the vendor’s RFP team and surface no signal. Tight, named, falsifiable questions outperform exhaustive ones.
Should the RFI ask about pricing?
Lightly, and only on the structural shape; per-call vs per-seat vs flat; not on the dollar number. Pricing is what gets negotiated last and dollar numbers in an RFI compress to the vendor’s list price. The RFI’s job is differentiation; pricing happens in the contract phase once the shortlist is set.
How does the RFI integrate with reference calls?
The RFI generates the candidates that pass the structural filter. Reference calls verify whether the vendor’s claimed answers hold in production. The two are sequential; rarely substitute one for the other. See the AI vendor reference-call playbook for the structured 11-question reference-call protocol.
What is the right rubric for scoring the RFI?
Score on five dimensions: eval discipline (does the vendor know what good output looks like for the buyer’s workload), engineering maturity (post-mortems, runbooks, on-call), accountability (named engineers, kill clauses), portability (weights, data, prompts on exit), and pricing alignment. Each at 0-3. Total score above 11 advances to references; below 8 declines without further work.
Key takeaways
Standard AI vendor RFIs ask 80-200 feature questions and produce uniformly green answers. Questions that surface real differentiation are about failure handling: eval-suite walkthrough, post-mortem reading, named-engineer commitment, kill clause, weight and data ownership, portability.
The template is 12-20 questions across five categories; eval discipline, engineering maturity, accountability, portability, pricing alignment. Each scored 0-3 out of 15. Above 11 advances to references; below 8 declines.
The eval walkthrough is the highest-signal question. Vendors who deliver it live in 15 minutes are the ones whose six-month behavior matches their demo. Vendors who stall or substitute slides get rebuilt in-house 12 months in.
Kill clause and weight ownership are forgotten in contract negotiation if not named in the RFI. Surfacing them costs nothing at RFI stage; raising them in week 8 of negotiation costs leverage. Run the template before most AI vendor decision over $100K.
Arthur Wandzel