Home About Who We Are Team Services Startups Businesses Enterprise Case Studies Industries Commercial Real Estate Blog Guides Contact Connect with Us
All Commercial Real Estate guides
Real Estate 14 min read

Document AI's Hidden Hallucination Problem — and the checks that catch it

Document AI's Hidden Hallucination Problem — and the checks that catch it

A tool reads a 40-page office lease and returns a clean abstract: base rent, escalation, notice window, renewal option. Every field is filled and every field looks right. One of them is a 3% annual escalation the model never found in your document — your lease actually escalates on CPI with a 2% floor and a 5% cap, and the model quietly substituted the market-typical number. The abstract is wrong, it is confident, and nothing about it looks wrong. That is the document AI hallucination problem, and the reason it costs small commercial real estate firms money is not that it happens often but that it hides. This piece covers what a hallucination is in a document-extraction setting, why the dangerous ones are invisible on a casual read, and five checks a 4–20 person firm can run without an engineer to catch them.

What a Document AI Hallucination Actually Is

A hallucination is a value the model presents as fact that appears nowhere in your source document — not a misread of text that exists, but text that does not exist. When ChatGPT, Claude, or a dedicated lease platform extracts a rent escalation your lease never stated, it has not made a transcription error. It has generated a plausible answer because that is what these models do, and a blank is less satisfying to the model than a market-typical number.

This matters because most firms evaluate extraction tools the way they would evaluate a scanner: does it read the page correctly? A scanner that misreads produces garbage you can see. A language model that hallucinates produces clean, format-perfect data that reads exactly like a correct answer. Only one of those failure modes announces itself.

The frequency is low and the stakes are high — the worst combination for a busy reviewer. Leading platforms report 90–97% accuracy on standard commercial lease terms, and roughly 95% once human review is layered on. Flip that around: on a 42-lease stack with a dozen fields each, a 3–5% miss rate is 15 to 25 wrong values, and a handful will be fabrications rather than misreads. The question is never whether they exist, but whether your process surfaces them.

Why the Dangerous Hallucinations Hide

There are two classes of extraction error, and firms that treat them as one thing build the wrong review process. Call them loud errors and quiet errors.

Loud errors are self-catching. OCR mangles a faxed exhibit into $4,S00, a date lands in 1902, a tenant name comes back as symbols. These are the errors vendors show in demos because they are easy to fix — any reviewer catches them in a glance, and format rules flag them automatically. Annoying, not dangerous.

Quiet errors are the hidden hallucinations. The fabricated value is well-formed and close to what an experienced broker would expect. A 3% escalation, a 60-day notice window, a 5-year renewal option — all so common that a reviewer skimming an abstract accepts them without checking the source, precisely because they look normal. The model’s confidence and the value’s plausibility combine to defeat the reviewer’s attention.

The pattern is consistent: fabrication concentrates wherever your lease deviates from the market-standard term. When your lease matches convention, a hallucinated “standard” value happens to be correct. When your lease was negotiated into something unusual — the reason it mattered enough to negotiate — the model reaches for the convention and gets it wrong on the term that carried the most weight in the deal. Research on document question-answering describes this directly: models are frequently “not wrong, but untrue,” confidently supplying answers unsupported by the source rather than admitting the source is silent.

Where Hallucinations Concentrate in a Lease

You do not have to review every field with equal suspicion. Fabrication concentrates in clauses that are long, negotiated, and cross-referenced — the ones that require interpretation rather than lookup. Identification fields and explicit dates are rarely fabricated; the model has no incentive to invent a party name sitting in plain text on page one.

Field class Hallucination risk Why
Party names, premises, square footage Low Stated plainly, once, in fixed locations
Explicit dates (commencement, expiration) Low Unambiguous and easy to ground
Base rent (initial) Low–medium Clear unless restated across amendments
CPI / percentage-rent escalations High Floors, caps, base years, breakpoints invite a “typical” substitute
CAM caps and exclusions High Buried, negotiated, easy to swap for a standard cap
Renewal and notice windows High Model reaches for the common 60/90/180-day convention
Co-tenancy, kick-out, holdover High Conditional logic scattered across sections and amendments

The financial stakes track the risk column. A misidentified escalation cap corrupts the rent projection for the whole remaining term. A missed notice window can trigger holdover rent at 125–150% of the last contractual rate. A CAM misclassification produces billing variances in the thousands of dollars per tenant per year, with disputes that reach litigation running well north of $25,000 a case. The clauses where the model is most likely to fabricate are the clauses where a fabrication is most expensive — which is why a flat “spot-check a few fields” habit is not a review process. Our guide to what a 95% document-AI accuracy claim actually measures breaks down why portfolio averages hide the fields that matter.

The Real-Citation Trap

The standard advice is to require a citation for every field. That advice is right but incomplete. A citation proves the model pointed at a page; it does not prove the page supports the value. The signature hidden hallucination is a real, checkable citation attached to a claim the cited text does not make.

Here is how it happens. You ask for the renewal notice window with a page reference. The model returns “180 days, see page 14.” Page 14 exists and does discuss the renewal option — it says the tenant “may renew upon notice as provided herein” and cross-references a definitions exhibit three amendments away that sets the window at 270 days. The citation is real, the answer is wrong, and a reviewer who confirms only that page 14 mentions renewals has confirmed nothing that matters.

This is why the research on catching model fabrication does not stop at “require a citation.” It grounds the citation — checking that the cited text actually entails the claim — and decomposes the answer so each piece is verified against the source. You do not need that tooling to apply the principle. You need reviewers who treat a page number as the start of a check, not the end, and who read the cited text on the fields that carry real money.

Five Checks That Catch a Hidden Hallucination

Each of the following catches a specific failure, and none requires an engineer. Run them in Excel, Outlook, and a PDF reader — the tools your firm already runs on.

1. Require verbatim quotes, not summaries. Demand that every financial field return the exact quoted clause and its page, with the value derivable from the quote alone. A summary lets the model paper over a fabrication in fluent prose; a verbatim quote forces the invented value into contact with text that either contains it or does not. You review the quote, not the model’s paraphrase of it.

2. Give the model permission to say “not found.” Hallucinations come partly from the model’s reluctance to return a blank. Instruct it explicitly that “not stated in the provided documents” is a correct and preferred answer when a term is absent, and that guessing is a failure. This converts a class of confident fabrications into honest blanks — which is the outcome you want, because a blank gets reviewed and a plausible fabrication does not.

3. Reconcile the numbers arithmetically. Rent schedules, escalations, and CAM figures are math, and math checks itself. Pull the abstracted base rent and escalation into a spreadsheet, project the schedule forward, and compare it against the rent roll. A fabricated escalation rarely reconciles to the independently stated figures, because the model invented one number without adjusting the others. Disagreement between two sources that should match is your loudest signal.

4. Diff two models on the high-stakes stack. Run the same lease through two different models — for example ChatGPT and Claude — and compare the fields. Agreement raises confidence; disagreement flags a field one of them fabricated or misread, which you route to human review. It is the cheapest form of the confidence scoring dedicated systems build in. Reserve it for diligence stacks and material leases, not routine volume.

5. Keep a golden set and re-run it. Pick five leases you know cold — including one with a messy amendment chain and one bad scan — and record the correct value for every field. Run this regression test against any tool before you trust it, and again whenever a vendor ships an update or you change your prompt. It reveals real per-field behavior on documents like yours, which no accuracy percentage can, and it is the discipline that separates firms that can switch tools in a week from firms held hostage by whatever their tool got wrong — a theme in our guide to why most lease-abstraction projects fail.

A Review Protocol for a Firm With No Analysts

The checks above are only useful inside a protocol, and a small firm’s protocol has to survive a 30-day diligence window without a dedicated review team. Tier the work by consequence.

  • Tier 1 — confirm, don’t read. Party names, premises, square footage, commencement and expiration dates. Glance, confirm it is present and sane, move on. These almost never hallucinate.
  • Tier 2 — reconcile. Base rent, escalation schedule, CAM figures. Run check 3; if the arithmetic ties out against the rent roll, accept, and if not, read the source.
  • Tier 3 — read the cited clause. CPI mechanics, caps and floors, renewal and notice windows, co-tenancy, kick-out, holdover. Open the page, read the quoted clause, confirm the value comes from it. This is where check 1 and the real-citation discipline pay off, and where a fabrication costs the most if it slips through.

A two-person team processes a 42-lease stack this way in days rather than weeks, because Tiers 1 and 2 move fast and the slow reading is reserved for the handful of clauses per lease that carry risk. The protocol is the deliverable, not the tool — the process-ownership-first philosophy that runs through the small CRE firm AI playbook and the broader document intelligence playbook this piece sits within.

General Assistant or Dedicated Platform

Hallucination control is where the general-assistant-versus-platform question stops being about features and becomes about verification. A business-tier assistant like ChatGPT or Claude abstracts a clean lease well, but it gives you nothing between the extraction and your trust except the prompt you wrote — you supply the citation requirement, the “not found” permission, and the reconciliation. That is fine for low volume and clean stacks, and it is why the two-model diff matters there.

Dedicated platforms build verification in — confidence flags on low-certainty fields, amendment-chain consolidation so a value cites the latest controlling document, and reconciliation against your system of record. That does not make them hallucination-proof — a confident wrong answer with a confidence flag is still wrong — but it moves work off your reviewers. The honest test is the same for both: run your golden set through the candidate and score it per field. A platform that fabricates on your five known leases is not safer than an assistant that does. Which documents each class of tool can and cannot handle is mapped in our field guide to what document AI can and can’t read.

Frequently Asked Questions

Why are AI hallucinations hard to spot in lease abstraction?

Because the fabricated value is plausible. The model substitutes the convention a reviewer would expect, so a busy person skimming the abstract accepts it without opening the source. Fabrication also concentrates in the clauses your lease negotiated away from standard — the clauses that carried the most money — so a wrong value tends to land on the term where being wrong is most expensive.

Does requiring a citation stop hallucinations?

Not on its own. A citation proves the model pointed at a page; it does not prove the page supports the value. The most dangerous hidden hallucination is a real page reference attached to a claim the cited text never makes. Treat a citation as the start of a check: on high-stakes fields, read the cited clause and confirm the value is derivable from it, not just that the page exists.

How accurate is AI lease abstraction, really?

Leading platforms report 90–97% accuracy on standard commercial lease terms, and about 95% once human review is added. A useful rule of thumb is that AI reliably extracts roughly 80% of a lease’s key information, with the negotiated financial 20% needing verification. Any single percentage without a per-field breakdown is a marketing number, because identification fields sit near 100% and complex escalation clauses drag well below the average.

Which lease fields does AI hallucinate most often?

CPI and percentage-rent escalations with floors, caps, and breakpoints; CAM caps and exclusions; renewal and notice windows; and conditional terms like co-tenancy, kick-out, and holdover scattered across amendments. The common thread is that these require interpretation and often deviate from the market standard. Party names, premises descriptions, and explicitly stated dates are rarely fabricated.

How do I check for hallucinations without an engineer?

Five checks run in the tools you already have. Require verbatim quotes for every financial field. Give the model explicit permission to answer “not found.” Reconcile extracted rent and escalation math against the rent roll in a spreadsheet. Diff two models on high-stakes leases. And keep a five-lease golden set with known-correct answers to re-run whenever a tool or prompt changes. None requires code.

What is the two-model diff and when should I use it?

Run the same lease through two different models — for example ChatGPT and Claude — and compare the extracted fields. Agreement raises confidence; disagreement flags a field one model fabricated or misread, which you review by hand. It is a low-cost version of the confidence scoring platforms build in. Reserve it for diligence stacks and material leases, not routine volume.

Are dedicated platforms safe from hallucinations?

No tool is hallucination-proof, including platforms with confidence flags and amendment-chain consolidation. Those features move verification work off your reviewers, which is real value, but a confident wrong answer can still pass. Judge every tool the same way: run your five known leases through it and score the output field by field. A platform that fabricates on documents you understand is not safer than an assistant that does.

How much does a hidden hallucination actually cost?

It depends on the field. A misidentified escalation cap corrupts the rent projection for the entire remaining term. A missed notice window can trigger holdover rent at 125–150% of the last contractual rate. A CAM misclassification produces billing variances in the thousands of dollars per tenant per year, with disputes reaching litigation running well above $25,000 a case.

Key Takeaways

  • A hallucination is an invented value, not a misread; it comes back clean and confident, which is exactly why a casual review misses it.
  • Separate loud errors (obvious, self-catching) from quiet hallucinations (plausible, well-cited, dangerous) — they need different review, not the same spot-check.
  • Fabrication concentrates in negotiated financial clauses: CPI escalations, CAM caps, notice windows, and conditional terms. Tier your review so those clauses get read, not confirmed.
  • A citation is the start of a check, not proof of correctness — on high-stakes fields, read the cited clause.
  • Five no-code checks catch most hidden hallucinations: verbatim quotes, “not found” permission, arithmetic reconciliation, two-model diff, and a golden-set regression test.

The fastest way to know where your firm stands is to measure it: which documents you process, where your current review misses, and which of these checks your process lacks. Our free AI-readiness assessment walks through that in a single working session.

Last Updated: Aug 8, 2026

DJ

Dirk Jan van Veen, PhD

SFAI Labs helps companies build AI-powered products that work. We focus on practical solutions, not hype.

Turn lease stacks into structured data

  • Lease abstraction with verification steps, not blind trust
  • LOIs, estoppels, and amendments handled the same way
  • Your documents never leave your firm's control

Related articles