Most non-engineer founders sign the AI MVP contract on Monday and have no plan for what they personally do on Tuesday. They have a partner, an SOW, a Slack channel, and a vague intention to “stay close to the build.” Twelve weeks later, the product ships either off-target (because the founder disappeared) or behind schedule (because the founder over-meddled). Both outcomes are downstream of the same missing artifact: a founder-side calendar of what to verify, what to do, and what to decide each week. This playbook is that calendar — a 12-week table with one row per week, the four founder failure modes that the calendar replaces, and a 1-page printable that fits over the founder’s desk on day one.
Table of Contents
- Why 90 days, and why the founder needs their own playbook
- The week-by-week table — 12 weeks, 3 columns
- Phase one — Weeks 0 to 2: scope, kickoff, eval baseline
- Phase two — Weeks 3 to 6: build cadence
- Phase three — Weeks 7 to 9: hardening and shipping
- Phase four — Weeks 10 to 12: post-launch and handoff
- The four founder failures that the playbook prevents
- The 1-page printable
- Frequently Asked Questions
- Closing
Why 90 days, and why the founder needs their own playbook
The 90-day window is not arbitrary. It is the cadence that Michael Watkins documented in The First 90 Days — the period long enough to produce a real artifact, short enough that the operator cannot coast, and aligned with how most fixed-price AI MVP engagements are scoped (a 6-to-12-week build plus a 30-day post-launch window). McKinsey’s State of AI surveys repeatedly show that the operational variance between high-performing and low-performing AI deployments sits in cadence and governance, not in model selection. Bain’s 2024 generative-AI survey reaches the same conclusion from the enterprise side. The pattern translates cleanly to founder-vendor relationships: the variance between two engagements with identical SOWs and identical staffing is the founder’s operating rhythm.
The founder-AI-partner operating manual describes that rhythm in narrative form. This playbook is the operational table: one row per week, three columns per row, and an explicit list of the four founder failure modes the rhythm prevents. It sits between that manual and the more specific guides on individual partner choices. It is the framework the founder operates after signing the SOW, before the first stand-up.
The structure is downstream of the idea-to-product manifesto: a founder who does not write code can still own the product, but only by owning a small, mechanical set of weekly artifacts, actions, and decisions. The playbook is the mechanism.
The week-by-week table — 12 weeks, 3 columns
Every week in the 90-day window has three things the founder owns: an artifact the partner produces and the founder verifies, an action the founder takes that week, and a decision the founder makes that the partner cannot make on their behalf. The discipline is to read the table once, put the meetings on a Google Calendar, and operate the calendar. The personality of the founder is irrelevant once the calendar is on the wall.
The week-by-week table below covers the standard 12-week shape. Variants — a 6-week sprint, a 16-week regulated build — adjust the row count but not the column structure. The companion piece on how an idea-to-product engagement works week by week covers the partner side of the same calendar. This piece is the founder side.
Phase one — Weeks 0 to 2: scope, kickoff, eval baseline
| Week | Artifact (partner produces, founder verifies) | Founder action | Founder decision owned |
|---|---|---|---|
| Week 0 (pre-kickoff) | Signed SOW, eval-set seed list, founder-onboarding doc | Confirm Slack channel, calendar holds, intro to internal stakeholders | Single point of contact on founder side (you, or one named designate) |
| Week 1 | PRD draft v0.1, capability inventory, named eval-set v1 | 60-minute kickoff meeting; weekly cadence locked | Top 3 user jobs the MVP must do (cut everything else) |
| Week 2 | Architecture sketch, model selection memo, first eval baseline numbers | 30-minute eval-review meeting; 30-minute scope-decision stand-up | Eval threshold for go-live (the “good enough” bar) |
Week 0 is the week most founders skip. The SOW is signed and the founder assumes the partner will run the engagement. In reality, week 0 is when the founder owes three things: a single named point of contact (the founder, not a delegated VP), a Slack channel that the partner can post into without permission gymnastics, and an introduction to any internal stakeholder whose opinion the founder will later quote. Skipping week 0 is the most common cause of week-3 surprises.
Weeks 1 and 2 produce the PRD, the eval set, and the eval baseline — the three artifacts the rest of the build depends on. The companion piece on what the first 14 days should produce walks the artifact list in detail. The founder’s job in this phase is not to write the PRD or the eval set. It is to verify that the partner’s drafts are honest about scope, and to make the one decision that only the founder can make: the eval threshold for go-live. The threshold is the contract for “done.”
Phase two — Weeks 3 to 6: build cadence
| Week | Artifact (partner produces, founder verifies) | Founder action | Founder decision owned |
|---|---|---|---|
| Week 3 | First working demo, eval pass/fail report, cost-per-call estimate | Watch a 10-minute Loom of the demo; 30-minute weekly review | Whether to broaden or narrow the eval set |
| Week 4 | Updated demo, eval scores by category, change-order register | Weekly review; sign-off on any change orders | Approve or refuse the first change order |
| Week 5 | Mid-point readout — what is on-track, what is at-risk, what is being deferred | 60-minute mid-point review with the partner lead | Which deferred items are post-launch, which are scope-cut |
| Week 6 | Pre-hardening demo, eval set finalised, integration list | Weekly review; founder-side user testing on 3 named users | Whether to launch on schedule or extend by one week |
Weeks 3 to 6 are the build half of the engagement. The founder’s calendar in this phase is mechanical: a 10-minute async Loom on Monday, a 30-minute live review on Friday, and the discipline to leave engineering decisions to the engineers. The companion role-of-evals piece describes the eval review as the founder’s primary instrument. Eval-scores by category are the only weekly metric the founder needs; commit counts, line counts, and ticket counts are partner-side metrics that should not be on the founder’s dashboard.
The mid-point readout in week 5 is the decision point. By week 5 the partner can name what is on-track, what is at-risk, and what is being deferred. Deferral is the most common failure mode for founders — they say “yes, defer it” without naming whether the deferral is post-launch (it ships later) or scope-cut (it never ships). The founder owns that distinction.
Phase three — Weeks 7 to 9: hardening and shipping
| Week | Artifact (partner produces, founder verifies) | Founder action | Founder decision owned |
|---|---|---|---|
| Week 7 | Hardening list (security, error handling, observability), pen-test plan | Weekly review; legal/compliance check-in | Production environment owner (founder or partner) |
| Week 8 | Launch checklist, pricing-page copy, on-call runbook | 60-minute launch-prep meeting | Pricing for the first 30 days (free, paid, gated waitlist) |
| Week 9 | Final eval run, go-live decision memo, customer-comms draft | Go-live meeting; founder sends launch comms to first cohort | Go / no-go for production launch |
Weeks 7 to 9 collapse two questions into one: is the product good enough to ship, and is the founder ready to charge for it. The good-enough question is answered by the eval set the founder defined in week 2 — the threshold is the threshold, not a feeling. The ready-to-charge question is owned exclusively by the founder. The companion hand-off process piece covers the artifact list the partner delivers at launch.
A frequent founder failure at week 9 is to defer pricing — to launch as “free for now” and decide pricing later. The deferral is almost always the wrong call. A free launch produces user behaviour that does not predict paid behaviour, and the founder loses the first 30 days of pricing signal. The decision the founder owns at week 9 is pricing — even a wrong price is more useful than no price.
Phase four — Weeks 10 to 12: post-launch and handoff
| Week | Artifact (partner produces, founder verifies) | Founder action | Founder decision owned |
|---|---|---|---|
| Week 10 | Production eval scores, real-user behaviour report, top-5 surprises | Weekly review; 3 customer calls (founder runs them) | Which surprise becomes the next iteration |
| Week 11 | Knowledge-transfer pack, model-migration plan, ongoing-cost forecast | Knowledge-transfer session with founder-side engineer (or future hire) | Ongoing engagement model (retainer, on-call, exit) |
| Week 12 | Final handoff, exit memo, 30-60-90-day forward roadmap | 60-minute retrospective with partner team | Renew, scope-down, or exit |
Weeks 10 to 12 are the 30-day post-launch window that most founders treat as an afterthought. The window is when real-user traffic meets the eval set, and the eval set either holds up or reveals which production scenarios the founder under-scoped. The founder’s job is to run three customer calls a week — the partner cannot run those calls, because the customers will not give a vendor the same signal they give the founder.
The week-12 decision is the engagement-shape decision. The companion exit-clause piece walks the contract mechanics. The founder owns one of three calls: renew the engagement on a retainer, scope down to an on-call window, or exit. There is no fourth option, and “we will figure it out next month” is the failure mode that produces silent over-spend.
The four founder failures that the playbook prevents
Founders fail at four mechanically distinct things. Each has a diagnostic and a remedy.
1. Under-engaging. The founder signs the SOW, misses week 0, attends one stand-up in week 2, and shows up at week 9 with surprise notes. The eval set does not match the product the founder wanted, but it is too late to rewrite it. The diagnostic is missed meetings — three missed weekly reviews in a row is the signal. The remedy is the 30-minute weekly review as a non-negotiable calendar block. The inside-the-AI-agency-standup piece makes the case that 30 minutes a week is the smallest possible founder commitment that prevents 30 days of rework.
2. Over-engaging. The founder reads every commit, sends Slack messages about variable names, second-guesses architecture in week 3, and produces a partner team that spends more time defending decisions than making them. The diagnostic is more than five founder messages per day in the engineering channel. The remedy is to channel founder input through one weekly meeting and one async Loom — engineering decisions stay with engineers, product decisions stay with the founder.
3. No-decisions. The founder attends every meeting, listens, takes notes, and never decides. By week 6 the partner is making product decisions by default because the founder will not. The diagnostic is the partner asking a binary question and getting a paragraph in response. The remedy is the table above — the founder is responsible for one named decision per week, and the decision is written down in the weekly review notes.
4. Scope drift. The founder verbally adds a feature in week 4 that was not in the SOW, the partner absorbs it as a “small change,” and by week 8 the engagement is two weeks behind. The diagnostic is any deviation from the eval-bound SOW that does not produce a written change order. The remedy is the change-order register (column three in the table above) and the discipline to refuse the absorbed favor. The cluster sibling hybrid playbook covers the related question of which 30 percent the founder keeps in-house so that scope drift on the partner side is contained.
The 1-page printable
The mechanical version of the playbook is a single page the founder pins above their desk on day one. The page lists 12 rows — one per week — with three columns: artifact, action, decision. Below the table are the four failure-mode names, each with a one-line diagnostic. The page closes with three numbers: the eval threshold for go-live, the cost-per-call ceiling, and the date of the week-12 retrospective. A founder who can read the page in 90 seconds on a Monday morning has the operating rhythm they need.
For the printable companion to this playbook — the 12-week table, the four failure-mode diagnostics, and a single-page founder dashboard — download the AI MVP scoping worksheet. The worksheet maps each weekly row to a partner-side artifact and a founder-side action that can be reviewed in under 30 minutes.
Frequently Asked Questions
What if my engagement is 6 weeks, not 12? Collapse phases two and three. The artifact-action-decision columns do not change; the row count compresses from twelve to six. A 6-week sprint loses the mid-point readout (week 5) and the formal hardening week (week 7) — the founder absorbs those into the weekly review.
Can I delegate the founder side of the playbook? Only the action column. The artifact-verification column can be co-owned with a founder-side engineer once one exists. The decision column cannot be delegated — by definition it is the founder’s decision, and a partner who takes founder decisions on the founder’s behalf is the partner failing the engagement.
What if my partner refuses to produce the weekly artifacts? That is the signal to terminate inside the exit clause. The companion exit-clause piece walks the contract mechanics. A partner that will not produce weekly eval-pass/fail reports and a change-order register is not a partner the founder can operate. The playbook is the founder’s instrument; if the partner refuses to populate it, the engagement is structurally broken.
Is the 1-page printable a replacement for the SOW? No. The SOW is the legal artifact. The 1-page printable is the operating artifact. The SOW changes only through written change orders; the printable does not change at all — it is the calendar the founder operates against, week after week.
What is the eval threshold for go-live, and how do I set it? The eval threshold is the score on the eval set the founder defined in week 2, above which the product is “good enough to ship.” Typical thresholds are 80 to 95 percent on a faithfulness or task-completion rubric, depending on the use case. The founder sets it in week 2 with the partner’s input, but the founder owns the number. The piece on the role of evals in your weekly partner relationship walks the threshold-setting conversation.
How do I run the three customer calls a week in weeks 10 to 12? 30 minutes per call, founder-only, no partner in the room. Ask three questions: what did you try to do, where did it fail, what would you pay for that does not exist. Write the answers in the same Google Doc the partner reads. The partner reads them and updates the production eval set in week 11.
What if the partner wants to skip week 0? A partner that wants to skip the pre-kickoff week is a partner that does not understand the cost of an unprepared kickoff. Insist on it. The week-0 artifacts (SOW, eval-set seed list, founder-onboarding doc) are non-negotiable, and a partner that refuses to produce them is signaling that the engagement will run hot.
Can I run this playbook with a freelance engineer instead of a partner? The structure transfers, but the eval-set discipline usually does not. Most freelance engineers will not author an eval set in week 1. If the founder is committed to a freelance route, the founder owns the eval-set authoring as well, which roughly doubles the founder-side workload. The piece on the founder’s role in an AI MVP build covers the role split.
What is the most common founder mistake in the first 14 days? Over-scoping the MVP. Founders bring lists of 15 user jobs to the kickoff and expect the partner to ship all of them in six weeks. The remedy is the week-1 decision: top three user jobs, cut the rest. The companion piece on what the first 14 days should produce walks the scoping artifacts.
How do I know if the playbook is working? Two signals. First, the partner can name what was on-track and what was at-risk at the week-5 mid-point readout without hedging. Second, the founder can name the week-12 decision (renew, scope-down, exit) by week 10, not week 12. If both signals are green, the operating rhythm is doing its job.
Closing
The 90-day playbook is not a personality trait. It is a calendar of artifacts, actions, and decisions that a non-engineer founder can run from a Google Calendar with named meeting blocks. The mechanical version of the rhythm beats every charismatic alternative — because the four founder failures (under-engaging, over-engaging, no-decisions, scope drift) are not character flaws. They are the predictable consequence of a missing schedule. The schedule above replaces all four. Pin the 1-page printable above the desk, hold the meetings, make the decisions, and ship the product the founder will be proud to defend on the day it launches.
Arthur Wandzel