A portfolio screenshot is not evidence of prior work, and a one-paragraph case study with a logo on it is marketing, not engineering. The standard “review the portfolio, check references” loop fails because it does not name what a buyer should actually request. This article replaces that loop with six concrete artifacts a founder can request inside a week, with a script for what to ask, what an evasive answer signals, and a substitute when the artifact genuinely cannot be shared. The goal is to convert an idea-to-product partner’s marketing portfolio into something a non-technical buyer can defend before signing a $100K to $300K engagement.
This procedure extends the idea validation playbook within the broader idea-to-product manifesto. It pairs with the solo-founder partner guide, the non-technical buyer’s partner-pick rubric, and the case-study skepticism in decoding the AI agency case study. This piece is the verification procedure run on artifacts.
Table of Contents
- Why the Portfolio Review Fails as Diligence
- Artifact 1 — The Eval Rubric From the Prior Project
- Artifact 2 — Deployed URL or Proof of Shipped vs. Prototype
- Artifact 3 — Handoff Artifact (Runbook or Post-Mortem)
- Artifact 4 — Reference Call With a Named Client
- Artifact 5 — Cost-vs-Scope Honesty
- Artifact 6 — Post-Launch On-Call Window
- A Five-Day Procedure to Run the Six Requests
- Frequently Asked Questions
- Closing
Why the Portfolio Review Fails as Diligence
The conventional portfolio review assumes the buyer can look at past projects and infer capability. That worked when the artifact was a website design or marketing campaign — both visible at face value. It does not work for an AI MVP, where almost everything that determines whether the build shipped, held quality, or held cost is invisible in a screenshot.
The AI agency category since 2023 produced a wave of “case studies” that are demos labelled as launches — a working prototype screenshotted and posted as a shipped product. McKinsey’s State of AI 2025 reported 78 percent of organizations using AI with value concentrated in a small fraction of deployments, and 2024–2025 Gartner CIO surveys put AI pilots stalling before production at roughly 85 percent. Most projects in any AI agency’s portfolio are statistically pilots — and the portfolio page rarely says so. The artifacts that distinguish a real build from a demo — eval rubric, deployed URL with live users, handoff runbook, post-launch on-call window — exist as documents and operational records, not screenshots. A buyer who does not ask for them does not get them.
A portfolio is a hypothesis about prior work. The six artifacts below test the hypothesis. For each: the request, what a stall signals, and the substitute when the artifact is genuinely unavailable.
Artifact 1 — The Eval Rubric From the Prior Project
What to ask for. “Send me the eval rubric from a prior engagement similar in scope to ours, with the client name redacted — test cases, scoring method, pass-rate threshold, and whether the threshold was contractually attached.” Strong: arrives within 48 hours, names cases, scoring method, and threshold. Weak: a generic QA deck, reassurance about “thorough testing,” or a marketing-shaped quality narrative.
What a stall signals. Either the prior engagement had no eval rubric — quality was negotiated post-hoc rather than contracted up-front — or the rubric exists but the partner is uncomfortable showing the shape of their quality bar. Both signal an engagement sold without a contracted quality definition. A partner running eval-first builds has redacted rubrics ready to share; the rubric is the load-bearing artifact of their model.
The substitute when unavailable. A current eval rubric template the partner uses on new engagements. It does not prove the prior project ran with one, but it proves the operating discipline. If neither a redacted rubric nor a template exists, the partner has not run eval-first builds.
Artifact 2 — Deployed URL or Proof of Shipped vs. Prototype
What to ask for. “For the three portfolio entries closest to our scope, send me the live URL the system runs at today, the current monthly active user count, and the date of the last production commit.” Strong: three URLs that load, three user counts in the hundreds-to-thousands range, three recent commit dates. Weak: “the client has the URL behind login,” “we cannot share user counts,” “the project transitioned to client maintenance.”
What a stall signals. A “transitioned to client maintenance” answer across all three usually means the system is dormant, abandoned, or never ran on real traffic. A partner whose clients legitimately gate URLs behind login can still produce evidence of liveness: a public marketing site for the product, a launch press release, a current support page, a visible commit history. A partner who cannot produce any external evidence for any portfolio entry has portfolio entries that did not graduate from prototype.
The substitute when unavailable. A recorded 15-minute screen-share with the partner’s engineering lead, demonstrating one prior project running on its production environment, logged in as a real user, with monitoring dashboards visible. A partner with shipped work produces this inside a day; a partner without it deflects.
Artifact 3 — Handoff Artifact (Runbook or Post-Mortem)
What to ask for. “Send me a redacted handoff document from a prior engagement — operational runbook, deployment guide, or post-mortem. Twenty to forty pages.” Strong: a structured document with deployment topology, secrets management, monitoring, known issues, on-call procedure, model and infrastructure cost per query, and a hardening backlog. Weak: a screenshot-laden slide deck labelled “handoff summary,” a one-page checklist, or “we handle handoff verbally.”
What a stall signals. A handoff artifact is the document that makes an engagement transferrable. A partner who does not produce one structurally cannot hand off — every prior project remains operationally dependent on them, sometimes as a deliberate retention strategy. A post-mortem is a stronger signal than a runbook because it requires naming what went wrong. A partner whose telling produces only “successful launches” has either zero prior engagements or zero learning loop.
The substitute when unavailable. A handoff document template the partner uses today, plus a written commitment in the SOW that a complete version will be produced for your engagement and reviewed at the hardening milestone.
Artifact 4 — Reference Call With a Named Client
What to ask for. “Three references — names, titles, companies, calendar links. PMs or product owners, not executive sponsors. I want to call them this week.” Strong: three named contacts arrive within 24 to 48 hours with scheduling links. Weak: “we will pass the request to past clients,” anonymized references, executive sponsors only, or a stall past 72 hours.
What a stall signals. A partner who has done good work has clients who will say so on the record, by name. Anonymous references — “a Fortune 500 financial services client” — are unverifiable and over-represent failed engagements where the client refused to be associated. An executive-sponsor-only reference list signals that operational contacts are unwilling to vouch, usually because they experienced the engagement differently than the executive who signed. The full reference-call procedure is in the AI agency reference call: 11 questions that surface real client outcomes.
The substitute when unavailable. Two named PM references plus one written attributable case study with named client, project lead, and quoted contact. One reference total is below the diligence floor.
Artifact 5 — Cost-vs-Scope Honesty
What to ask for. “For three prior engagements, send me the contracted scope, contracted price, delivered scope, and final price billed. Redact client names.” Strong: three projects, all four numbers each, plus a paragraph on each explaining why the delta exists. Weak: “we always ship on scope and on budget,” refusal to discuss prior pricing, or vague ranges without specific projects.
What a stall signals. A partner unwilling to discuss the gap between contracted and delivered cost has nothing to compare against your engagement. The honest version of every engagement has variance — scope grows, milestones slip, a model upgrade re-prices inference. A partner claiming zero variance across all prior projects is either lying or has not run enough projects to encounter the realistic frequency of change orders.
The substance is not “did variance exist” — variance is normal. The substance is whether the partner has language for talking about variance honestly. A partner who can name a project that came in 20 percent over budget, explain why, and explain what they changed afterwards has a learning loop. A partner who cannot does not.
The substitute when unavailable. A written breakdown of the partner’s three most common change-order categories from the last twelve months, with the typical magnitude of each. This swaps a project-level disclosure for a pattern-level disclosure. Neither one is a pass.
Artifact 6 — Post-Launch On-Call Window
What to ask for. “For two prior engagements, describe the post-launch on-call window — duration, hours covered, response time committed, number of incidents handled, and whether any exceeded the response SLA.” Strong: 30 days of weekday business-hours coverage, four-hour response SLA, five incidents handled, zero SLA breaches. Weak: “we always support our clients after launch,” no specifics, no formalized structure.
What a stall signals. The post-launch window is where AI products fail most predictably — quality drift, cost spikes from a retrieval-side bug, integration breaks when a third-party API updates. A partner without a formal on-call window has either not seen these failure modes or has chosen not to take operational responsibility at handoff.
A “30-day on-call window with zero incidents on every engagement” is statistically improbable past two or three engagements. A partner reporting zero incidents either is not measuring what counts as an incident or is not operating the system. Two to six incidents per 30-day window across engagements is the honest range.
The substitute when unavailable. A written commitment in your SOW for a specific post-launch window with all five parameters — duration, hours, response SLA, escalation path, transition to maintenance retainer or in-house engineer. The future commitment partially substitutes past evidence only if the partner is willing to write it with specific numbers. Past evidence absent and future commitment refused is a pass.
A Five-Day Procedure to Run the Six Requests
The six artifacts collapse into a one-week diligence sprint that runs in parallel with the rest of the founder’s evaluation.
Day 1. Send the partner a single email listing all six requests with specific deadlines. “Eval rubric and handoff artifact by Wednesday EOD. Three named PM references with calendar links by Thursday EOD. Cost-vs-scope and on-call records by Friday EOD. Deployed URLs and liveness evidence by Friday EOD.” A partner who can answer is operating at the cadence required to run your build. A partner who needs a two-week extension is not.
Days 2 and 3. Score each artifact: strong (present, specific, dated, attributable), workable (present but partial), weak (absent, substitute offered), pass (absent, no substitute). Six strong is credible. Four strong with the rest workable is signable with a paid proposal review. Three or fewer is a pass.
Day 4. Run the three reference calls. The portfolio that survives the artifact requests but does not survive the reference calls is a marketing portfolio with poor client experience underneath.
Day 5. Compile a one-page summary — six scores, three reference-call notes, cost-vs-scope finding, on-call evidence. This is the document you defend the engagement against twelve months from now if the build runs into trouble. The buyer who runs the procedure has roughly the same evidentiary base a technical buyer would have produced through code review, without writing code.
Expected operational cost: six to ten hours of buyer time across the week. Expected return: the difference between signing a $150K engagement on provable prior work and signing on a screenshot. The asymmetry is why the procedure exists.
Frequently Asked Questions
How is this different from “check the portfolio and talk to references”? The standard advice names two artifacts and ignores the four that actually predict engagement quality — eval rubric, deployed URL evidence, handoff artifact, cost-vs-scope honesty, and on-call evidence. This procedure makes the full six-artifact request explicit, with a stall interpretation and substitute for each.
What is the single most predictive artifact? The eval rubric from a prior project. A partner who can produce a real, contracted eval rubric has run the engagement model that ships AI products at quality. No other artifact compensates.
How long should this take? Eight days total — five for artifact requests and reference calls, two to three to read the artifacts. Founders running in parallel with their broader evaluation typically close the diligence inside two weeks.
What if the partner cannot share client names due to NDAs? Plausible for one or two engagements, implausible across a whole portfolio. A blanket “all our work is under NDA” usually means the work is very limited in volume or so recent it has not finished.
Does the procedure apply across different partner specializations? The six artifacts apply identically across AI MVP partners, traditional dev shops, and product studios. Calibrate the eval-rubric threshold to how AI-load-bearing your build is.
How do I evaluate a partner with limited prior work? A partner with two or fewer engagements cannot pass on artifacts alone. Compensate with a paid feasibility engagement ($5K to $15K) before the larger build, plus named senior engineers in the SOW with founder approval on swaps.
What does this procedure cost in vendor goodwill? Less than founders fear. Strong partners welcome the requests because their portfolios survive them. Weak partners stall — which is the signal the buyer is looking for.
What if a partner fails on one or two artifacts but is strong elsewhere? Acceptable if the failures are on artifacts 4, 5, or 6. The first three — eval rubric, deployed URL evidence, handoff artifact — are non-negotiable.
Does SFAI Labs go through this procedure with founders? Yes. The first conversation is a 30-minute idea review where the founder walks through the six-artifact request, and we name which artifacts we can produce. Book the idea review here.
Closing
A portfolio is a hypothesis the partner asks the buyer to accept. The six artifacts test the hypothesis: the eval rubric tests whether quality was contracted, the deployed URL evidence tests whether the prior project graduated from prototype, the handoff artifact tests whether the engagement was transferrable, the reference call tests whether the client experienced the work the way the case study claims, the cost-vs-scope disclosure tests operational visibility, and the on-call window tests whether the partner took responsibility past handoff.
Run together inside one week, the six artifacts convert a marketing portfolio into a defensible evidentiary base. A buyer who signs a $150K engagement on that base is signing on evidence. A buyer who signs on a screenshot is signing on hope. The procedure is the difference, and it is doable by any non-technical founder with five days and the willingness to send a single email asking for six specific things.
Arthur Wandzel