Home About Who We Are Team Services Startups Businesses Enterprise Case Studies Industries Commercial Real Estate Blog Guides Contact Connect with Us
All Commercial Real Estate guides
Real Estate 17 min read

Anatomy of a successful automation pilot

Anatomy of a successful automation pilot

A successful automation pilot is not the one where the AI works — it is the one that returns a decision you can act on, cheaply, before you commit real money to a build. For a small commercial real estate firm, the pilot is a decision instrument, not a launch. Its job is to tell you, in a few weeks and for a few thousand dollars, whether a custom automation is worth tens of thousands more — or whether the honest answer is a subscription, a prompt library, or nothing at all. Most pilots fail not because the technology is weak but because they were never designed to answer that question. This is the anatomy of one that is.

Most pilot advice online is written for a company with a data team, a project manager, and a change-management budget. You have none of those. You are a principal or an operations director at a firm of four to twenty people, the deal data is confidential, and the person running this pilot is the same person who will live with whatever it produces. That changes the design of the whole thing. A pilot at your firm has to be small enough to run without hiring anyone, honest enough to survive a real workload, and structured to end in a clear yes or no. If you have not yet decided whether custom software is even the right category of answer, the buy-versus-build playbook is the step before this one; this piece assumes you are ready to test a specific idea.

Why Most Pilots Fail Before They Start

The failure rate is not a rumor. MIT’s Project NANDA, in its 2025 State of AI in Business study, found that roughly 95% of enterprise generative-AI pilots produced no measurable impact on the bottom line. Gartner has projected that around 30% of generative-AI projects are abandoned after the proof-of-concept stage. Those numbers get quoted as evidence that the technology is overhyped. They are better read as evidence that most pilots are badly designed — run as demos, judged on vibes, and never tied to a number anyone recorded before the work began.

A demo shows the tool working on a clean example the vendor chose. A pilot tests whether the tool works on your worst inputs, at your real volume, at a cost you can live with, and produces a result you can compare against how you do the job today. The difference between the two is the difference between the 5% and the 95%. The seven parts below are what a pilot has that a demo does not.

Part 1: One Workflow, Chosen for the Right Reason

A successful pilot automates exactly one workflow, and it is almost never the most exciting one. The right candidate is repetitive, high-volume, measurable, and painful — a task someone at your firm does the same way dozens of times a month and can describe in a sentence. Lease abstraction, rent-roll standardization, first-pass deal screening, and CAM reconciliation all qualify. “AI strategy” and “an assistant that helps with everything” do not; they cannot be tested because they cannot be bounded.

Pick the workflow where you can already picture the input and the output. “Read inbound lease PDFs and extract twelve named terms into our existing Excel abstract template” is a pilot. “Modernize our document operations” is a wish. The narrower the scope, the faster and cheaper the pilot, and the cleaner the eventual decision. If you are tempted to automate three things at once to save time, you have designed a project, not a pilot — and a project is exactly the commitment the pilot exists to de-risk.

Part 2: A Baseline Measured Before You Build Anything

This is the part almost every failed pilot skips, and its absence is why so many end in an argument instead of a decision. Before a single prompt is written, record how the workflow performs today: how long it takes per item, how often it produces an error, and what it costs in loaded hours. If your analyst spends nine hours a month standardizing rent rolls and catches an error in roughly one file in ten, write that down. That is the number the pilot has to beat.

Without a baseline, “the AI is pretty good” is unfalsifiable and “the AI made a mistake” is disqualifying, and neither is true. Every automation makes mistakes; so does every analyst. The only question that matters is whether the automated workflow is faster, cheaper, and at least as accurate as the manual one you run now — and you cannot answer it if you never measured the manual one. Spend the first few days of the pilot measuring the present. It is the cheapest and most-skipped step in the anatomy.

Part 3: A Test Set Built From Your Ugliest Documents

A pilot that runs only on clean, representative inputs is a demo with extra steps. The test set for a real pilot is assembled deliberately from your firm’s worst material: the scanned lease with the handwritten rider, the rent roll a property manager built in a spreadsheet with merged cells, the deal package that arrives as forty pages of mixed PDFs. Those edge cases are where automations break, and they are the ones you deal with in production, so they belong in the test.

Hand the builder — or the tool, if you are testing one yourself — twenty to fifty real documents that span your actual range, including the outliers. If the automation holds up on the ugly ones, it will breeze through the clean ones. If it only works on the tidy examples, you have learned something valuable and cheap: the workflow is not yet automatable at your firm, and you just found that out for the price of a pilot instead of the price of a build. This is also the honest test that separates a custom build from the duct-tape approach of a spreadsheet wired to a chatbot — the messy inputs are exactly where the simple version tends to fall over.

Part 4: A Human-in-the-Loop Acceptance Bar

A language-model automation is probabilistic: right most of the time, wrong some of the time. That is not a defect to be shocked by; it is a property to be designed around. A successful pilot defines, in advance, what “good enough to accept” means and keeps a human in the loop where the stakes require it. The acceptance bar has three parts: a stated accuracy threshold on the test set, a rule that every extracted figure links back to its source page or cell, and a defined review step for anything the model flags as low-confidence.

“To my satisfaction” is not an acceptance bar; it is a moving target. “The automation extracts all twelve lease terms with at least 95% field-level accuracy on the fifty-document test set, flags any field it is unsure of, and links every value to the page it came from” is one. That sentence tells you exactly when the pilot has passed and exactly what the human still has to do — which, for a firm handling confidential deal data with no IT department, is the safeguard that keeps a fast tool from becoming a liability. Speed with a verification step still beats manual re-keying by a wide margin.

Part 5: A Run-Cost Ceiling Declared Upfront

Every pilot has a build cost and a run cost, and the second one is where firms get surprised later. Before you start, get an estimate of what the automation will cost to operate each month at your real volume — model usage, any hosting, and maintenance. At small-firm volumes, model and infrastructure costs for a single scoped workflow typically land in the tens to low hundreds of dollars a month, usually less than the per-seat subscription the automation would replace. The pilot should confirm that number against actual usage, not leave it as a guess.

Declaring the ceiling upfront does two things. It stops a pilot from quietly validating an automation that is technically impressive and economically absurd, and it gives you a real figure to fold into the three-year cost comparison that decides whether the build pays for itself. A pilot that proves the workflow works but never measured what it costs to run has answered half the question and left you to discover the other half after you have paid for the build.

Part 6: A Time Box and a Go/Kill Gate

A pilot has an end date, decided before it begins. Two to four weeks is enough to test a single workflow against a real document set; a pilot that has run for three months has stopped being a test and become an unmanaged project. The time box protects you from the most expensive failure mode in the NANDA and Gartner numbers — the pilot that never quite finishes, never quite fails, and never quite ships, absorbing money and attention while producing no decision.

At the end of the box sits a go/kill gate: a scheduled review where you compare the pilot’s results against the baseline from Part 2 and the acceptance bar from Part 4, and you make one of three calls. Go — the automation beat the baseline, cleared the bar, and costs less to run than it saves; commission the build. Kill — it did not, and you stop. Or iterate once — the results are close and one specific fix would get them over the bar, so you run one more short cycle and then decide for real. The gate is a calendar event with named criteria, not a feeling you arrive at eventually.

Part 7: A Verdict That Can Legitimately Be “Don’t Build”

The hardest part of a good pilot to accept is that a “no” is a win. A pilot exists to buy a decision cheaply, and a pilot that ends in “do not build this — the workflow is too messy, the volume is too low, or an off-the-shelf tool already does it well enough” has done its job perfectly. You spent a few thousand dollars to avoid spending fifty thousand on the wrong thing. That is the highest return a pilot can produce, and the vendors running the pilot as a sales motion will never frame it that way, because for them only “go” counts.

Judge the pilot by the quality of the decision it produced, not by whether it green-lit a build. Some of the most valuable pilots at a small firm end with the automation shelved and a cheaper answer adopted instead — a prompt library the team runs by hand, a subscription that turned out to fit, or a decision to wait until volume grows. A partner who scopes the pilot so that a kill is a possible and respectable outcome is a partner you can trust with the build; one who has pre-decided that the pilot will succeed is selling, and the pilot is theater. This judgment is part of the wider case for how a lean firm out-operates larger competitors, covered in the small-firm AI manifesto.

The Pilot on One Page

A successful automation pilot at a small CRE firm has seven parts, and you can check for all of them before you authorize a dollar:

  1. One workflow — repetitive, measurable, describable in a sentence.
  2. A baseline — today’s time, error rate, and cost, recorded before any building.
  3. A test set — twenty to fifty of your real, ugly documents, edge cases included.
  4. An acceptance bar — a stated accuracy threshold, source-linking, and a human review step.
  5. A run-cost ceiling — the monthly operating cost estimated and then confirmed against real usage.
  6. A time box and a go/kill gate — two to four weeks, ending in a scheduled decision with named criteria.
  7. A permitted “no” — a design where “don’t build” is a legitimate, valuable outcome.

Miss the baseline and you cannot judge the result. Miss the test set and you validated a demo. Miss the run-cost ceiling and the build surprises you later. Miss the kill option and you never ran a pilot — you ran a pre-approved rollout with a trial period bolted on. When the pilot returns a “go,” the next document to get right is the contract that defines the build, which is its own section-by-section anatomy.

Frequently Asked Questions

What makes an automation pilot successful?

A successful automation pilot returns a clear, defensible build-or-don’t-build decision at low cost and low risk — not a working demo. It automates one bounded workflow, measures a baseline before any building starts, tests the automation on the firm’s real and messy documents, defines in advance what accuracy counts as passing, confirms the monthly run cost, and ends on a scheduled date with a go, kill, or one-more-iteration verdict. Success is the quality of the decision it produces, which sometimes means proving you should not build at all.

How long should an automation pilot take?

Two to four weeks is enough to test a single workflow against a real document set at a small firm. The pilot should have an end date decided before it starts. A pilot that runs for several months has stopped being a time-boxed test and become an unmanaged project — the exact failure mode that consumes budget without ever producing a decision. If the results are close but not conclusive at the deadline, run one more short, specific iteration and then decide.

What workflow should we pilot first?

Choose the workflow that is repetitive, high-volume, measurable, and painful — one someone at your firm does the same way many times a month and can describe in a sentence. Lease abstraction, rent-roll standardization, first-pass deal screening, and CAM reconciliation are strong candidates. Avoid broad, unbounded ideas like “an AI assistant for everything,” because a scope with no edges cannot be tested and cannot pass or fail.

Why do so many AI pilots fail?

MIT’s Project NANDA found that around 95% of enterprise generative-AI pilots showed no measurable bottom-line impact in 2025, and Gartner has projected roughly 30% of generative-AI projects are abandoned after the proof-of-concept stage. The common cause is design, not technology: pilots run as demos on clean data, judged without a baseline, with no accuracy bar and no end date. A pilot built with those parts is far more likely to land on the right side of those statistics.

Do we need a baseline before running a pilot?

Yes, and skipping it is the most common reason pilots end in disagreement instead of a decision. Record how the workflow performs today — time per item, error rate, and loaded cost — before any building begins. Without that number, “the AI is pretty good” cannot be verified and “the AI made a mistake” wrongly disqualifies a tool that is still faster and more accurate than the manual process. The manual baseline is the only fair thing to judge the automation against.

What documents should we use to test the pilot?

Your worst ones. Assemble twenty to fifty real documents that span your actual range, deliberately including the edge cases — scanned leases with handwritten riders, rent rolls with merged cells, mixed-format deal packages. Automations break on messy inputs, and those are the inputs you handle in production. A pilot that only runs on clean, representative examples has tested a demo and will disappoint you the first week it meets real work.

How much does an automation pilot cost?

Market ranges, not a single price. A scoped pilot is a small fraction of a full build — the build itself for a single-workflow custom automation generally starts in the tens of thousands, and the pilot is designed to be far cheaper than that so a “don’t build” verdict is affordable. The point of the pilot’s low cost is exactly that math: a few thousand dollars spent testing can save tens of thousands spent building the wrong thing. Confirm both the pilot cost and the ongoing run cost before you start.

What are the run costs of a custom automation?

Three recurring lines: model usage, hosting or infrastructure, and maintenance. At small-firm volumes, model and infrastructure for a single workflow typically total tens to low hundreds of dollars a month — often less than the per-seat subscription the automation replaces. Maintenance is the line a firm with no IT department cannot skip. A pilot should confirm these numbers against real usage so the run cost is a measured figure, not a guess you discover after the build.

Is it a failure if the pilot says “don’t build”?

No — it is one of the best outcomes a pilot can produce. The pilot exists to buy a decision cheaply, and a well-run pilot that concludes “the workflow is too messy, the volume is too low, or an off-the-shelf tool already does this well enough” has saved you from an expensive mistake for a small fee. Judge a pilot by the quality of its decision, not by whether it approved a build. A partner who allows a kill as a real outcome is one you can trust; one who has pre-decided success is selling.

Should a vendor run our pilot, or should we?

You should own the pilot’s design and criteria even if a builder runs the technical work, because a vendor-owned pilot is a sales motion in which only a “go” counts. Keep control of the workflow choice, the baseline, the test set, the acceptance bar, and the go/kill gate. A trustworthy partner will welcome that structure and will scope the pilot so a “don’t build” is a respectable result. If a vendor resists a defined kill criterion, that resistance is itself a finding.

Where to Start

A pilot is only as good as the workflow it tests, and most firms want to run one before they have decided which of their workflows is actually worth automating — and in what order. That inventory is the real first step: a ranked read on which tasks would pay for a custom build, which are better served by a subscription or a prompt library, and where the honest answer is to wait. That is what a free AI-readiness assessment produces — a working session that maps your firm’s workflows, flags the ones where a pilot is worth running, and returns a plan with real cost ranges. Book a free AI-readiness assessment if you want that map before you commission a single pilot. If a build is not the right spend, the assessment will say so, and you will still leave with the plan you needed to design a pilot that returns a decision you can act on.

Last Updated: Aug 14, 2026

AW

Arthur Wandzel

SFAI Labs helps companies build AI-powered products that work. We focus on practical solutions, not hype.

Make your firm fluent in AI — then automate what works

  • Hands-on training applied to LOIs, lease summaries, and market write-ups
  • Automation across documents, deals, communications, and back office
  • Built for 4–20-person firms with no IT department

Related articles