Home About Who We Are Team Services Startups Businesses Enterprise Case Studies Industries Commercial Real Estate Blog Guides Contact Connect with Us
Back to Guides
Enterprise Software 13 min read

Pre-build AI feasibility studies: when they're worth $5K

Pre-build AI feasibility studies: when they're worth $5K

A $5K pre-build AI feasibility study buys four hours of a senior engineer, a 10-input capability probe, a 3-frontier comparison, and a one-page go-or-no-go memo — and it is the highest-ROI line item in an AI MVP budget when the build at stake is $50K or more. The math is brutally simple. A founder who skips the probe and discovers in week four of a $150K build that the current frontier cannot do the task has burned thirty times the price of the study. A founder who pays $5K up front, gets a “kill” or “pivot” memo on Friday, and walks back to the drawing board has paid one-thirtieth of that cost to avoid the same outcome. That asymmetry is the procurement case.

This is the buyer-side operating manual for the $5K probe. Companion reading: the idea-to-product manifesto, the eval-first build playbook, the DIY AI feasibility check, and the next-product-up AI scoping workshop.

Decision scope

A procurement guide for one line item: the paid feasibility study. Assumes the founder has decided feasibility matters and is choosing between DIY ($0) and a $5K senior engineer probe. It does not re-argue the case for feasibility — the eval-first build playbook does that. Dollar figures are 2026 US market bands, not contractual quotes.

Why a $5K probe exists as a category

The $5K probe sits in a specific gap on the founder procurement ladder. Below it: the DIY feasibility check (three chat consoles, ten test inputs, a Saturday afternoon). Above it: the $12K two-day workshop (eval set, cost model, architecture sketch). The $5K probe answers one narrow question — can the current frontier reliably do this task, and which model is best at it.

It is not a strategy engagement. It is not architecture. It is a senior engineer running ten test inputs against three frontier consoles, scoring outputs against a three-tier rubric, and writing a one-page recommendation by end of day.

The category exists because two conditions hold in 2026. First, frontier model capabilities are uneven and quarter-by-quarter volatile — what GPT-5 handled in March may shift in June after a Claude Opus 4.8 or Gemini 2.5 Pro point release. Second, non-technical founders cannot reliably distinguish a “the model can do this” prompt from a “this is the easiest case the model will ever see” prompt. The probe answers both for the price of a domain plus a year of hosting.

What’s in the $5K box: four named artifacts

A real $5K feasibility study ships four named files. If your vendor cannot list these in the proposal, the engagement is theatre.

Artifact What it contains
capability-probe.md The AI task stated in capability terms — extraction, classification, generation, planning, multi-turn dialog, tool-use, or multi-modal — with the worst-link identified for chained capabilities.
test-set.csv Ten real anonymized input examples from the founder’s data, distributed 3 easy / 4 medium / 3 hard, with expected output labeled in advance.
three-frontier-comparison.md Per-input scores against Claude Opus 4.8, GPT-5, and Gemini 2.5 Pro using the useful / wrong-recoverable / wrong-dangerous rubric, plus an overall winner and tie-breaker notes.
go-no-go-memo.pdf A one-page signed recommendation: proceed, pivot, or kill, with trigger conditions cited, the recommended build model, and residual risks the founder must accept.

Anything else is filler. A real probe does not ship a 40-slide deck, an architecture diagram, a roadmap, or a “phase 2 proposal.” Each is a separate purchase at a different price point.

The hour budget: where the 4 hours go

Four hours of senior engineer time is the load-bearing constraint. The hour budget is the structural reason the price equilibrium is $5K.

Hour Work
Hour 1 Founder briefing. Capture product description, sample data, domain context. Name the capability. Build the test set with the founder.
Hour 2 Run the 10-input test against all three frontier consoles. Score each output. Take screenshots of dangerous failures.
Hour 3 Analyze the score table. Identify the cross-model pattern. Determine the outcome — proceed, pivot, kill.
Hour 4 Write the one-page memo. Founder readout call. Sign and send.

Economics: a senior engineer at $400 fully-loaded hourly bills $1,600 for four hours. Add 40% for memo writing and overhead, and labor cost is roughly $2,250. The remaining margin ($2,750) covers agency overhead, sales cost, and the 24-hour reservation premium. At $5K, neither side is getting a steal — which is what the equilibrium price looks like.

The three outcomes: proceed, pivot, kill

A real feasibility study produces one of three outcomes. The market bias is to bury the second two, because “proceed” sells the next engagement.

Proceed. Eight or more of the ten test inputs scored useful on at least one frontier model. Zero or one dangerous failures across all three models on easy and medium inputs. Memo recommends model X for the build, names the residual risks, and suggests the next purchase up the ladder.

Pivot. The frontier cannot do the task as defined, but it can do an adjacent task that may still solve the founder’s user problem. Example: the model cannot extract every clause from a 60-page contract, but it can surface the three highest-risk clauses for human review. The memo names the adjacent task and recommends a follow-on probe.

Kill. Two or more dangerous failures on easy inputs across all three frontier models, or fewer than four useful outputs total. The frontier cannot do this task today with chat-console access. Some come back as proceed in twelve months — models improve quarter by quarter — but today’s build is not viable. The memo recommends a calendar reminder for re-evaluation.

A vendor whose proposal has no path to “kill” is not selling a feasibility study. They are selling a permission slip.

When the $5K study is worth it

Three conditions, all must hold:

  1. The MVP at stake is $50K or more. Below that, the asymmetry collapses. A $20K MVP saved by a $5K study has a 4× ROI — still good, but inside the band where founder-DIY may have done the same work for free.
  2. The founder cannot run the 3-frontier comparison themselves. Running the probe well requires writing clean prompts for three model families, distinguishing “bad prompt” from “bad model,” and scoring against a three-tier rubric without rationalizing. Founders with engineering backgrounds can skip the paid version.
  3. The founder’s time costs more than $1,250 per hour. Doing four hours of probe work themselves at that rate matches the purchase price. Founders mid-fundraise, mid-launch, or mid-hiring may be in this band.

Two out of three is a close call. One out of three means do the DIY check and save the $5K for the build.

When to skip it (and what to buy instead)

You have an engineering co-founder. Run the DIY feasibility check on a Saturday. The output is functionally identical, and the co-founder learns the failure modes firsthand.

The MVP at stake is under $25K. Build a 2-week throwaway prototype. At sub-$25K, the build is the probe.

The task is well-trodden. Skip if the capability is documented working in production at well-known products. Email classification, extraction from clean PDFs, short-form English generation — these have shipped at scale for years.

You need more than capability confirmation. If you need a full eval set, cost-per-query model, or architecture sketch, buy the 2-day workshop instead. The probe deliberately stops at capability.

You are still in idea-validation mode. If you have not talked to ten target users, you are too early. The probe answers “can the model do this,” not “should it be doing this.” User research is the prerequisite, not the substitute.

What a good study reads like

A real go-no-go memo is one page, signed, dated, and survives a CTO reviewing it during diligence. Specific markers:

Named models with version numbers. “Tested against Claude Opus 4.8 (June 2026 release), GPT-5 (xhigh setting), and Gemini 2.5 Pro.” Not “tested against leading LLMs.”

Numeric scores per input. A table with ten rows, four columns (input, score per model), and an aggregate. Not “the model performed well overall.”

Screenshots of dangerous failures. Verbatim outputs where the model invented a fact, fabricated a clause, or confidently mislabeled. The dangerous failures are the most valuable artifact in the study.

A go-or-no-go in the first sentence. “Recommend proceed with Claude Opus 4.8 for v0; pivot from per-clause to highest-risk-clause extraction.” Not “the technology shows promise.”

Residual risk named in dollars or scenarios. “Residual risk: model behavior on Spanish-language contracts untested — recommend $3K supplementary probe before adding non-English support.” Not “additional risks may exist.”

Vendor diagnostic questions

Five questions for any vendor selling a “pre-build AI feasibility study”:

  1. Who specifically runs the four hours — name and title? Acceptable: a named senior engineer with 4+ years of LLM work. Unacceptable: “our team”, “a strategist”, “an account lead and an engineer”.
  2. What files ship by Friday — exact filenames? Acceptable: capability probe, test set CSV, frontier comparison, go-no-go memo. Unacceptable: “a final report”, “a slide deck”, “a recommendation document”.
  3. Tell me about the last time you recommended kill. Acceptable: a specific prior engagement. Unacceptable: “we always find a path forward”, “we never recommend kill”.
  4. What’s the smallest task you have ever probed? Acceptable: a narrow capability (e.g., “extracting the effective date from a Master Services Agreement”). Unacceptable: vague generalities.
  5. What’s the 2026 version of each frontier model you’ll test? Acceptable: named versions (Claude Opus 4.8, GPT-5, Gemini 2.5 Pro). Unacceptable: “the leading LLMs”.

Most vendors fail question 3 and question 5 — the first because they cannot afford the reputational cost of recommending kill, the second because they have not actually been in the chat consoles recently.

Where the $5K study sits in the procurement ladder

Product Price Who runs it What walks out
DIY feasibility check $0 Founder, 4 hours Same artifacts, lower-quality scoring
$5K pre-build feasibility study $5K Senior engineer, 4 hours Capability probe + memo
2-day scoping workshop $8K–$15K Senior engineer, 2 days Probe + eval set + cost model + architecture sketch
5-day discovery week $20K–$30K Senior engineer + PM, 5 days Workshop output + domain interviews + signed roadmap
2-week classical scoping $35K–$60K Multi-person team, 2 weeks Detailed SOW + full architecture + sprint plan

Procurement principle: buy the smallest product that answers your open questions. The $5K probe collapses 60% of those answers; the 2-day workshop closes the rest. Skipping straight to a 2-week engagement is over-buying — and over-buying eats runway just as efficiently as a wrong build does.

The reverse — skipping the probe to “save money” — is the failure mode behind the runaway AI project anti-pattern and most cost lines on the AI project TCO checklist. The cost discipline of an AI MVP is set on the day the founder decides whether to do a probe. Everything downstream — the $250K MVP budget, the cost-per-query model, the model selection — compounds against that decision.

The $5K is not a tax. It is a forcing function that protects the next $150K from being burned on a question that has a four-hour answer.

FAQ

How is the $5K probe different from a free DIY feasibility check?

The work product is similar; the seniority of the operator is the difference. A non-technical founder running the DIY check produces lower-quality test inputs, prompts, and scoring because they have not done this fifty times. For a technical founder the DIY check is genuine signal; for everyone else it is mostly noise. The $5K version is the same artifacts produced by someone with reps.

Why $5K and not $1K or $25K?

A senior engineer at $400 fully-loaded hourly costs $1,600 for the probe, plus $650 in memo and overhead. Below $3K the engagement cannot afford the right seniority. Above $10K the buyer is paying for theatre — a deck, a binder, a closing presentation that adds no signal. $5K is the equilibrium.

Can a $5K probe really kill a $150K project?

Yes, and it is the highest-ROI use of the $5K. Roughly one in four probes returns “kill” or “pivot.” The dollar value of a single avoided $150K wrong-build covers thirty probes. The avoided 6-week schedule slip — founder opportunity cost, runway, team morale — is rarely smaller than the build itself.

What if the probe says “proceed” and the build still fails?

The probe is bounded scope. It answers whether the frontier can handle the capability — not whether the data pipeline will work, users will adopt, or unit economics will hold. A “proceed” memo names those residual risks. A build that fails on data engineering after a clean probe is a risk the probe deliberately did not cover.

Can I do the probe myself and pay someone to write the memo?

In theory yes, but the value is the memo signature. A memo signed by the founder carries the founder’s bias. A memo signed by a senior engineer with no equity in the outcome is what procurement, investors, and co-founders will accept.

Should I run probes against open-weight models like Llama 3.3?

Not in the $5K version. The probe asks whether the closed frontier can do the task — that sets the ceiling. Open-weight re-enters as a cost lever during the build phase, not the feasibility phase.

What happens between the memo and the build?

A “proceed” memo is typically followed by either the 2-day scoping workshop (if the founder needs an eval set and cost model before signing a build proposal) or a direct $50K–$150K MVP build (if the team already has eval discipline in-house). About 70% of probes flow to the workshop next.

What if the vendor wants to scope the probe larger?

Red flag. The probe is deliberately small. A vendor proposing a “$15K extended feasibility study” is selling the 2-day workshop under a different name, or padding the engagement to justify a junior engineer’s time. If the vendor cannot ship the four artifacts in four hours, they are not a feasibility-study vendor.

Ready to get one done? Book an idea review and we will scope a $5K pre-build feasibility study, or point you to the right product up the ladder.

Last Updated: Jul 16, 2026

AW

Arthur Wandzel

SFAI Labs helps companies build AI-powered products that work. We focus on practical solutions, not hype.

See how companies like yours are using AI

  • AI strategy aligned to business outcomes
  • From proof-of-concept to production in weeks
  • Trusted by enterprise teams across industries
Get in Touch →
No commitment · Free consultation

Related articles