Most off-track AI MVP engagements are recoverable. The founder mistake that turns a recoverable engagement into a dead one is jumping from move 0 — silent worry — to move 4 — kill the contract — without running the lighter moves in between. The five moves below are a graduated escalation ladder. Each costs more trust capital than the last; each tells the founder something specific about whether the engagement is salvageable. The point of the ladder is that a lighter move is always available before the heaviest, so the kill conversation only happens after the cheaper ones have been tried and failed.
The founder-AI-partner operating manual covers the weekly cadence that prevents most off-track engagements. The idea-to-product manifesto frames the broader operating model. This piece is the companion for when the cadence has been running and the engagement is still drifting.
Table of Contents
- Thesis: the ladder protects the founder from the kill conversation
- Move 1: request the eval data
- Move 2: call a milestone reset
- Move 3: ask for senior engineer substitution
- Move 4: invoke kill-clause notice
- Move 5: bring in a third-party reviewer
- The order matters: a 4-week escalation calendar
- What each move costs the partner — the mirror view
- Decision Scope
- Frequently Asked Questions
- Closing
Thesis: the ladder protects the founder from the kill conversation
A six-to-twelve-week AI MVP engagement that is off track in week 3 is recoverable in week 4 in roughly four-out-of-five cases — observationally, in the engagements we have seen. The same engagement, untouched until week 6, is recoverable in roughly one-in-five. The difference is rarely engineering quality; it is whether the founder ran the lighter moves while a reset was still cheap, or waited for the relationship to break and tried to kill the contract under emotional pressure.
The McKinsey 2024 State of AI survey reports roughly 30% of organisations cite “lack of clear strategy” as the top adoption barrier. At MVP scale, that rarely means the SOW is unclear — it means the eval rubric drifted, the scope grew quietly, or a particular engineer is not a fit, and nobody raised it as a formal escalation. The moves below are the escalation language a founder can use without legal counsel, without a procurement function, and without a co-founder. They are graduated, reversible until move 4, and each one buys information about whether the next is needed.
The moves are professional escalations, not punitive ones. A competent partner welcomes the first three because the partner is also looking for a clean reset and rarely raises one unilaterally. The kill conversation in move 4 is the first move that costs trust capital. Move 5 — the third-party reviewer — often saves the engagement precisely because the partner ships harder under external review even when the review finds nothing serious.
Move 1: request the eval data
When. Week 2 onward, the moment the engagement feels more opaque than the kickoff promised. Default: bring the request to the Wednesday eval review.
How. One sentence in Slack: “Can you share the current eval suite — the rubric, the held-out set, the pass/fail thresholds, and last week’s run results? I want to read the numbers myself.” The request is low-friction and within scope. A founder is paying for the eval suite; the suite is a deliverable.
Expected response — competent partner. Within 48 hours, a shared dashboard with the rubric, held-out set, thresholds, and the last two weeks of run results. A short walkthrough offered by the technical lead. No defensiveness — the partner is glad the founder asked.
Expected response — evasive partner. Delay, a Loom video instead of the data, a verbal summary, or a reframe — “the demo is the better way to evaluate this”. Demo as a substitute for eval data usually means there is no rubric, or the rubric exists but is not being run weekly.
What the response tells you. Eval data shared cleanly: the engagement may be off track on goals, but not on practice. Move to move 2 if the data shows rubric drift, or stay at move 0 if it shows on-track work the founder had stopped seeing. Eval data not shared: the engagement has a deeper rubric problem, and move 2 follows within the week.
Cost. Zero trust capital. Asking for eval data is standard founder behaviour in 2026. The eval rubric template every non-technical founder should ask for is the checklist for what a competent answer looks like.
Move 2: call a milestone reset
When. Week 3 or 4. Use after move 1 if eval data shows rubric drift, or directly if a milestone has slipped and the partner has not raised a formal reset.
How. A scheduled 45-minute Friday call with the technical lead and project manager. Stated agenda: “Reset the next milestone. Revised scope, revised eval rubric, revised timeline, revised budget. Documented in the SOW.” Not a vibe conversation — a documented amendment.
Expected response — competent partner. Relief, often. The partner has been wanting the conversation but could not raise it without looking inflexible. They show up with a draft revised scope, a rubric proposal, and a timeline impact estimate. The meeting ends in a signed Notion amendment.
Expected response — evasive partner. Pushback — “we don’t need a reset, we’ll catch up” — or agreement followed by no documented amendment. The verbal-only reset is the most common evasion. If the partner cannot produce a written amendment within five business days, treat that as evasion and move to move 3.
What the response tells you. A clean reset is recoverable engagement — the partner is shipping, the goalposts just need to move. A failed reset means the operating layer is broken and engineering throughput cannot fix it. Move 3 follows within a week.
Cost. Low trust capital. A reset is a healthy conversation that founders typically delay too long. The weekly founder-partner cadence names the Friday scope standup as the recurring slot where small resets happen routinely; move 2 is the big version of that meeting.
Move 3: ask for senior engineer substitution
When. Week 4 or 5, after moves 1 and 2. Trigger: the team is shipping volume but the founder cannot read the technical decisions as competent, or the eval review consistently surfaces the same engineer making the same questionable calls.
How. Direct conversation with the partner’s principal — not the project manager. One sentence: “I would like a senior engineer substitution on [model selection / retrieval / eval design]. The current owner is not a fit for this scope.” Frame it as a fit problem with the scope, not a complaint about the engineer.
Expected response — competent partner. A swap proposed within five business days. Sometimes internal — a more senior team member takes the specific scope. Sometimes escalatory — a principal personally takes the scope for two to three weeks until the area is stable. Both are healthy.
Expected response — evasive partner. Defence of the engineer — “they are our most experienced on this stack” — or a stall with no revert. The stall is a signal that the partner’s bench is thin and the swap is not possible without a partnership-level conversation, which is move 4.
What the response tells you. A clean swap is recoverable engagement at the team-fit level — the partner is healthy; one person was not a fit, and the bench can fix it. A blocked swap means the partner does not have the bench; move 4 follows.
Cost. Moderate. The conversation raises the question of whether the partner has the bench. Most founders skip move 3, jumping from move 2 to move 4 because they read the problem as “the partner” when it was “this engineer”. The piece on the 6 anti-patterns we see in every failed AI agency engagement is the partner-side mirror.
Move 4: invoke kill-clause notice
When. Week 5 or 6, after moves 1, 2, and 3 have run and failed. Trigger: eval data not shared, reset verbal-only, senior-eng substitution blocked, engagement still off track.
How. Written notice — email, not Slack — to the partner principal, referencing the kill clause in the SOW. One paragraph: the moves run, the responses observed, the kill clause activated, the notice period started, the handoff and code escrow terms invoked per contract. Copy a lawyer if one is on retainer. The notice does not commit the founder to terminating; it activates the notice period, typically 14 to 30 days.
Expected response — competent partner. Immediate principal-level engagement. Within 24 hours, the principal proposes a recovery plan or, if they read the engagement as unrecoverable, an orderly wind-down with code escrow and handoff. The notice period is used to either restore the engagement or deliver a clean handoff. Both are valid outcomes.
Expected response — evasive partner. Delay, defensive contract interpretation, attempts to renegotiate the notice period mid-clock. A partner that defends the contract instead of defending the work has already left the engagement.
What the response tells you. A competent principal response often saves the engagement at the eleventh hour — the kill notice forces the partnership-level conversation that moves 1, 2, and 3 could not surface. An evasive response confirms the kill and shifts the conversation to handoff terms. Either way, the founder now has a 14-to-30-day window with explicit contract standing.
Cost. High. Trust capital that does not return. Even if the engagement recovers, the relationship after a kill notice is transactional, not partnership-shaped, for the remainder of the engagement. Use move 4 only after moves 1, 2, and 3 have been run honestly and documented. The AI agency exit clause every founder should negotiate is the BoFu reference for the contract language move 4 invokes.
Move 5: bring in a third-party reviewer
When. Week 4 or 5 — typically in parallel with move 3, or as an alternative to move 4 when the founder is unsure whether the engagement is unrecoverable. Also useful pre-handoff if move 4 has been invoked.
How. Hire an independent senior AI engineer or a small review firm for a 5-to-10-day technical audit. Scope: code, eval suite, architecture decisions, model selection rationale, deploy pipeline. Deliverable: written report with a competence verdict (on-track, drifting, or compromised) and three to five specific recovery moves. Budget: roughly five-to-fifteen thousand dollars; lower for code-only, higher with eval design audit.
Expected response — competent partner. Welcome it. A confident partner is glad to be reviewed because the review usually validates the work and resets the founder’s confidence. Full code, eval, and decision-log access without negotiation.
Expected response — evasive partner. Resistance — “this violates the partnership spirit” or “the reviewer will not have full context”. Resistance to an external technical audit on a fixed-price MVP is the strongest single signal that the partner has something to hide. A confident partner cannot afford resistance; it reads worse than any finding.
What the response tells you. A welcomed review usually saves the engagement — the report either validates the work or surfaces specific recovery moves the founder can use as moves 2 or 3 with a written technical basis. A resisted review is move 4 territory.
Cost. Moderate. The review fee plus a small trust-capital cost. The benefit is a senior technical opinion the founder does not have in-house — the structural reason most founders cannot self-diagnose engagement health. The decision tree for non-technical founders managing an AI build names the moments where external input is needed.
The order matters: a 4-week escalation calendar
The five moves are not interchangeable. The order is calibrated so each move buys information about whether the next is needed. Skipping moves is the most common founder error.
| Week | Move | Rationale |
|---|---|---|
| Week 2 | Move 1 — request eval data | Cheapest move; baselines whether rubric exists |
| Week 3 | Move 2 — milestone reset | If move 1 surfaced drift; documents the reset |
| Week 4 | Move 3 — senior-eng substitution | If reset was verbal-only or rubric drift continued |
| Week 4–5 | Move 5 — third-party reviewer | In parallel with move 3 if engagement still ambiguous |
| Week 5–6 | Move 4 — kill-clause notice | Only after moves 1, 2, 3 ran and failed |
A founder who runs moves 1, 2, and 3 honestly recovers the engagement in roughly seven-in-ten off-track cases — observationally. A founder who jumps from move 0 to move 4 recovers roughly one-in-ten, because move 4 is a binary conversation and the partner usually defends. The ladder works because each move is reversible until move 4. The 5 founder anti-patterns that delay AI MVPs covers the founder behaviours that turn a recoverable engagement off-track in the first place.
What each move costs the partner — the mirror view
A competent partner reads each move differently. Knowing the partner-side read sharpens the founder’s calibration on which move to use.
- Move 1 — eval data request. Reads as: “the founder is engaging professionally; I am glad they asked.” Costs the partner nothing if the eval suite exists; costs significantly if it does not.
- Move 2 — milestone reset. Reads as: “this is a normal mid-engagement conversation.” Costs small trust capital if the reset is needed.
- Move 3 — senior-eng substitution. Reads as: “a fit fix, not a partnership fix; I can solve this internally.” Costs moderate trust capital and bench depth.
- Move 4 — kill-clause notice. Reads as: “formal jeopardy; principal-level.” Costs the partnership relationship even if the engagement recovers.
- Move 5 — third-party reviewer. Reads as: “second opinion; I will look strong if I welcome it.” Costs small trust capital but signals founder seriousness — often constructive.
The anatomy of a runaway AI project is the cost-side mirror. Together, the two pieces give the founder both the action ladder and the cost-side diagnostic.
Decision Scope
This article is an editorial playbook, not legal advice. Move 4 invokes specific contract language; if the founder has a lawyer on retainer, copy them on the notice. If the kill clause language is ambiguous in the SOW, retain a lawyer before invoking move 4. The ladder is calibrated for fixed-price idea-to-product engagements in the six-to-twelve-week range with an eval-gated definition of done — adjust the cadence proportionally for longer engagements.
Frequently Asked Questions
When should a founder run move 1 — the eval data request?
Week 2 onward, the moment the engagement feels more opaque than the kickoff promised. A competent partner shares the eval rubric, held-out set, and run results within 48 hours.
What does it mean if the partner offers a demo instead of the eval data?
Demo as a substitute for eval data usually means there is no rubric, or the rubric exists but is not being run weekly. Move 2 (milestone reset) follows within the week.
How is a milestone reset different from a normal scope conversation?
A milestone reset is a documented amendment to the SOW with revised scope, eval rubric, timeline, and budget. A verbal-only reset feels collaborative and changes nothing operationally. If the partner cannot produce a written amendment within five business days, treat that as evasion.
Is asking for a senior-eng substitution unfair to the engineer?
No, if framed as a fit problem with the scope rather than a complaint about the engineer. The partner’s principal reads the request as a fit fix they can solve internally. The most common founder error is skipping move 3 and treating an individual-fit problem as a partner-fit problem.
What is the typical notice period after invoking the kill clause?
Typically 14 to 30 days, depending on the SOW. Invoking the notice does not commit the founder to terminating; it activates the notice period, during which the engagement either restores under principal oversight or delivers an orderly handoff with code escrow.
How much should a third-party AI project review cost?
Roughly five-to-fifteen thousand dollars for a 5-to-10-day review by an independent senior engineer. Lower for code-only, higher with eval design audit. Small relative to the kill-decision it supports either direction.
Can move 5 be run instead of move 4?
Often yes. Move 5 frequently saves an engagement at the same week move 4 would have been invoked. A confident partner welcomes the review, and the report either validates the work or surfaces specific recovery moves. Resistance to a third-party review is itself the signal that move 4 was the right move.
What if the engagement does not have a kill clause in the SOW?
Move 4 is much weaker without one. Most idea-to-product SOWs in 2026 include a 14-or-30-day kill clause as standard. If yours does not, retain a lawyer before invoking move 4 — the conversation becomes a renegotiation rather than a clean clause invocation. The lesson for the next engagement is to negotiate the kill clause up front; the cluster’s BoFu sibling on the exit clause covers the language.
Should the founder run move 5 (third-party review) every engagement as a baseline?
No. A baseline third-party review on a healthy engagement reads as low-confidence behaviour and damages partner trust unnecessarily. Reserve move 5 for engagements where moves 1, 2, or 3 surfaced ambiguity that the founder cannot resolve in-house.
Closing
Most off-track AI MVP engagements that end in a kill conversation could have been recovered if moves 1, 2, and 3 had been run two weeks earlier. The ladder is calibrated so the lightest move costs nothing and the heaviest costs the partnership. A founder running moves 1, 2, and 3 honestly will recover most off-track engagements; a founder jumping from silent worry to kill notice rarely will. The point is not to use every move — it is to know which move is right this week and to use the lighter ones first.
If you are pre-kickoff, the AI MVP Scoping Worksheet makes the eval rubric, milestone definitions, and kill clause explicit before week 1 — most of move 1 and move 4 become unnecessary when the worksheet is filled in honestly. If you are mid-engagement, the founder-AI-partner operating manual is the weekly cadence that catches drift in week 3 instead of week 6. The founder briefing pattern is the lighter-than-move-1 ritual for surfacing engagement health weekly without ever needing the escalation ladder.
Arthur Wandzel is the founder of SFAI Labs, a forward-deployed AI development studio in San Francisco. He has run, observed, or been brought in to review enough off-track AI MVP engagements to name the five moves above as the ladder that separates recoverable engagements from dead ones.
Arthur Wandzel