Document AI is sold on a demo run against a clean lease. Your firm runs on scanned PDFs, faxed amendments, and hand-annotated estoppels of wildly uneven quality — and that gap is where a $30,000 decision goes wrong. A small commercial real estate firm cannot absorb a bad software purchase or maintain a broken build, so the buying decision has to be disciplined in ways the vendor listicles never teach. These ten rules are that discipline. Each one is a testable criterion you can apply before you sign — what to run through the tool, what to demand from the vendor, and what should disqualify a product outright. They are ordered roughly by how often we watch small firms get burned when they skip them.
Rule 1: Test on your worst PDFs, not the demo’s best
Every vendor demo runs on a clean, native-text lease that reads perfectly. Your document stack does not look like that. It is scanned amendments, a faxed estoppel, a purchase agreement someone annotated in pen, a rotated page from a 1990s ground lease. Accuracy claims — Prophia, for instance, markets figures as high as 99% on standard fields — are real for clean fields on clean documents, and they fall on the messy ones. The buying decision lives entirely in that gap.
The test: pull ten real documents that represent your actual mix, deliberately including your worst scans and least standard leases. Run them through the trial. Then have someone who knows leases check every extracted value against the source. A tool that nails base rent and misreads a co-tenancy trigger on a faxed page has not passed — and the polished demo will never show you that failure. This one test, which costs a day, is the single most decisive thing a buyer does. The reading layer under every tool is the same foundation we cover in our guide to turning lease stacks into structured data; the point here is to prove it works on your stack before money changes hands.
Rule 2: Buy grounding, not just answers
A chat window returns text. A tool worth buying returns text linked to its exact location in the source document — click the extracted rent figure and jump to the clause it came from. This feature has an unglamorous name, grounding or source-linking, and it is the most important thing you can buy, because it turns verification from an hour of hunting through a PDF into a few seconds of confirming a highlighted clause.
The test: ask the vendor to show you a review screen, not an output screen. Extract a value, click it, and confirm it takes you to the source. If a reviewer cannot trace every number back to a page and clause in one click, the tool has quietly handed your team the slowest part of the job. A product without grounding is selling you answers you still have to find yourself.
Rule 3: Ask who fixes it when it breaks
This is the rule that decides most purchases and disqualifies most custom builds for a firm your size. Document AI breaks — not loudly, but silently. A landlord changes its lease template, a new scan defeats last quarter’s extraction, a model update shifts behavior, and suddenly a renewal date is wrong with full confidence behind it. Someone has to notice, diagnose, and fix that. In a firm with no IT department, that someone does not exist.
A purpose-built product folds this into the subscription: the vendor’s team maintains the extraction and carries the accuracy claim. A custom pipeline hands the entire maintenance obligation to you. That trade is the heart of the buy-versus-build question, and we lay out the full framework in our piece on off-the-shelf document AI versus custom pipelines. The buying test is blunt: ask the vendor exactly what happens when a new document type starts failing, and if you are considering a build, name the person on your staff who will own it. If you cannot name them, you are not ready to build.
Rule 4: Separate the sticker price from the maintenance tail
The number on the quote is not the cost. For a subscription product, the sticker price and the total cost are close, because maintenance is the vendor’s problem. For a custom build, they are worlds apart: a competent developer or agency delivers a working pipeline inside a normal budget — a scoped custom automation project runs roughly $25,000 to $150,000 in the current market — but the number small firms forget is the standing cost of keeping it accurate as documents and models drift.
The test: for every option, write down two numbers — what it costs to acquire and what it costs to keep working for two years. A product’s second number is its subscription. A build’s second number is a person, internal or retained, and it is the line item that turns a cheap build into an expensive one. Price the maintenance tail explicitly, or it will price itself later.
Rule 5: Require confidence scores you can route on
Extraction accuracy is never uniform. The value of a good tool is not that it is always right — nothing is — but that it tells you when it is unsure, so a human looks at the doubtful 10% instead of re-checking all of it. That is what a confidence score is for: it turns “verify everything” into “verify what the model flagged.”
The test: ask whether the tool exposes per-field confidence and whether you can set a threshold that routes low-confidence extractions to human review. A number on a dashboard you cannot act on is decoration. A score wired into a review queue is a workflow. For a lean team, the difference is whether the tool saves hours or just relocates them.
Rule 6: Keep the money clauses human-reviewed by design
The market sells full automation as the goal. For CRE documents it is the wrong goal, and a tool that pushes you toward it should worry you. Accuracy holds up on parties, dates, and base rent; it drops on the clauses where money actually lives — co-tenancy, exclusives, unusual escalations, and renewal options. Those are exactly the clauses a firm cannot afford to get wrong, and they are exactly where extraction is least reliable.
The test: confirm the tool is built to keep a human in the loop on the clauses that matter, not to remove them. A permanent verification step on money clauses is not a temporary crutch you automate away — it is the correct design for a document type where a single missed exclusive can cost a tenant relationship. Where general chat tools and purpose-built products each break on this is the subject of our comparison of ChatGPT versus purpose-built document AI for lease review. Buy the tool that makes verification fast, not the one that promises you will never verify again.
Rule 7: Confirm the data terms before you upload a rent roll
Your documents contain rent rolls, tenant financials, and material under NDA. Before any of it goes into a tool, you have to know what happens to it. The safe pattern is a business or enterprise tier where your inputs are not used to train the model by default, backed by real security posture. On the general-purpose side, ChatGPT Business and Enterprise and Claude Team and Enterprise both contractually exclude your data from training and hold SOC 2 Type II certification; on the proptech side, ask the vendor for the equivalent.
The test: read the terms of the specific plan you intend to buy — not the marketing page, the data-processing terms — and confirm three things: inputs are excluded from training, the vendor holds SOC 2 Type II or equivalent, and you have a rule for anonymizing NDA material where possible. Terms change, and a firm handling confidential deal data cannot assume the free-tier default is the plan default. This is due diligence, not paranoia.
Rule 8: Match the tool to your volume, not the brand
The right tool is a function of how many documents you actually process, not which product demos best. Under roughly ten leases a month, a thin workflow — a saved prompt in ChatGPT, Claude, or Gemini that returns your template fields, with a person verifying the important clauses — is frequently the honest answer, and no one sells it to you because there is no subscription to sell. In the dozens per month, a purpose-built product earns its price through the structured data store, audit trails, and time saved. In the hundreds of repeating documents, a build starts to compete on unit cost.
The test: count your real monthly volume before you shop, and be suspicious of any tool priced for a portfolio far larger than yours. Paying enterprise pricing for a handful of leases a month is the most common overspend we see. Volume, not brand preference, is the variable that should decide.
Rule 9: Own your data’s exit
A tool that reads your leases into a structured store is holding your data. The question no one asks during a demo is how it comes back out. If your abstracts, critical dates, and clause-level terms live only inside the vendor’s platform and cannot be exported cleanly, you have not bought a tool — you have rented a hostage situation, and the switching cost compounds every month you stay.
The test: before signing, confirm you can export your full structured data in a usable format — CSV, Excel, or an API — on demand and on exit. A vendor confident in its product will make this easy; a vendor that makes it hard is telling you something. This matters more the further you get from a general document type, which is one reason to weigh how specialized your needs really are, as we discuss in our look at legal-grade contract AI versus CRE-focused document tools.
Rule 10: Get your people fluent before the tool, not after
Every rule above depends on a person who can judge the output. Rule 1’s ten-document test needs a reviewer who knows what a correct abstract looks like. Rule 6’s human-in-the-loop needs someone who can spot a misread exclusive. A firm that buys document AI before its people are fluent has bought a machine no one can quality-check — and it will trust the wrong outputs for exactly as long as it takes for one to cause a problem.
The test: before the tool, make sure at least two people on your team can read an AI-generated lease abstract and confidently say where it is right and where it is guessing. That fluency is cheap to build — market-rate training workshops run roughly $2,000 to $15,000 — and it is the capability that makes every other rule enforceable. The broader argument for why this fluency is the real edge a small firm holds over larger competitors runs through the small-firm AI playbook. Buy the capability first; buy the tool second.
The rules as a buying scorecard
Run every candidate tool through the ten rules as a pass/fail scorecard before you commit a dollar.
| # | Rule | The disqualifying answer |
|---|---|---|
| 1 | Test on your worst PDFs | Only demoed on clean, native-text leases |
| 2 | Buy grounding | No click-to-source verification |
| 3 | Who fixes it | No clear maintenance owner (product or person) |
| 4 | Price the maintenance tail | Two-year cost never calculated |
| 5 | Actionable confidence scores | Scores you cannot route on |
| 6 | Human-reviewed money clauses | Full automation pushed as the goal |
| 7 | Confirm data terms | Inputs used for training; no SOC 2 or equivalent |
| 8 | Match to volume | Priced for a portfolio far larger than yours |
| 9 | Own your data’s exit | No clean export on demand |
| 10 | Fluency first | No one who can judge the output |
A tool does not have to score a perfect ten, but the answers to rules 3, 6, and 7 are non-negotiable for a small CRE firm: an unmaintainable tool, one that removes human review from money clauses, or one with unsafe data terms should be disqualified regardless of how well it demos on everything else.
Frequently asked questions
What is the most important thing to test before buying document AI?
Run ten of your own real documents through the tool, deliberately including your worst scans, faxed amendments, and least standard leases, then have someone who knows leases check every extracted value against the source. Vendor accuracy claims are measured on clean documents and clean fields; your buying decision lives on the messy PDFs and the money clauses. A tool that gets base rent right but misreads a co-tenancy trigger has not passed, and the demo will never show you that.
What is grounding and why does it matter when buying document AI?
Grounding, also called source-linking, means every extracted value links back to its exact location in the source document, so a reviewer clicks a number and jumps straight to the clause it came from. It matters because verification is the slowest part of using document AI, and grounding turns an hour of hunting through a PDF into seconds of confirming a highlighted clause. A tool without it hands the hardest work back to your team.
Should a small CRE firm build its own document AI or buy a product?
For most 4-to-20-person firms, buy. A custom pipeline hands you a permanent maintenance obligation — keeping extraction accurate as lease templates and models drift — that a firm with no IT department cannot staff. Building makes sense only under specific conditions: high monthly volume, a document type no vendor handles well, or a strategic need to own the data layer. Absent those, and absent a named person to maintain it, a product whose vendor carries the maintenance is the right call.
How much does document AI cost for a commercial real estate firm?
Purpose-built products are priced as subscriptions, usually per lease or per portfolio, so the sticker price and the total cost are close. A custom automation project runs roughly $25,000 to $150,000 in the current market depending on complexity, but the number firms forget is the maintenance tail — the standing cost of a person, internal or retained, who keeps it accurate after launch. Price both the acquisition cost and the two-year cost of keeping the tool working before you decide.
Is our confidential deal data safe with document AI tools?
It can be, but you have to verify the specific plan. The safe pattern is a business or enterprise tier where your inputs are not used to train the model by default, backed by SOC 2 Type II or equivalent certification. ChatGPT Business and Enterprise and Claude Team and Enterprise both contractually exclude your data from training and hold SOC 2 Type II; ask any proptech vendor for the equivalent. Read the data-processing terms of the exact plan you buy, not the marketing page, and set a rule to anonymize NDA material where you can.
How accurate is AI lease abstraction?
Vendors commonly report accuracy above 95%, and some market figures as high as 99%, on standard fields such as parties, dates, and base rent — and that holds for the easy fields on clean documents. Accuracy drops on non-standard clauses such as co-tenancy, exclusives, and unusual escalations, which is exactly where the financial risk sits, and it drops further on poor scans. Treat headline numbers as vendor claims about the easy cases, and keep a human verifying the money clauses against the source.
How many leases per month justify buying a dedicated tool?
Under roughly ten leases a month, a thin workflow built on ChatGPT, Claude, or Gemini with human verification is usually enough. In the dozens per month, a purpose-built product earns its subscription through a structured data store, audit trails, and time saved. In the hundreds of repeating documents, a custom build starts to compete on unit cost. Count your real volume before you shop, and be wary of paying portfolio-scale pricing for a handful of documents.
What should disqualify a document AI tool outright?
Three answers should end the evaluation for a small CRE firm regardless of how well the tool demos otherwise: no clear owner to maintain accuracy when it breaks, a design that removes human review from the money clauses, and data terms that use your inputs for training or lack SOC 2 or equivalent certification. An unmaintainable tool, an over-automated one, or an unsafe one is a liability no feature list offsets.
Do we still need people reviewing the output if the tool is accurate?
Yes. AI is a fast first-pass reader, not a substitute for reading the clause that decides money. Grounding and confidence scores shrink the reviewer’s job to the doubtful and the high-stakes, but the clauses where accuracy drops are the clauses with financial consequences, so a human confirms them against the source before anyone relies on the abstract. A firm also needs at least two people fluent enough to judge the output, or it cannot tell when the tool is guessing.
Where to start
Before you shortlist a single product, get honest about two things: whether your team is fluent enough to judge what any tool produces, and which of your real documents will break it. A free AI-readiness assessment produces that read — a short working session that maps your document mix, your monthly volume, and your workflows, and returns a plain recommendation for whether a thin workflow, an off-the-shelf tool, a build, or a month of fundamentals first is your right next move. Book a free AI-readiness assessment before you commit a dollar to any of these tools.
Arthur Wandzel