The wrong question is whether you can trust AI with legal documents. The right question is how much to trust it on which line — because a base-rent extraction and a co-tenancy trigger deserve completely different levels of faith, and the money in commercial real estate hides in the second kind. A small firm cannot afford either mistake: distrust the tool everywhere and you have bought a machine that saves no time; trust it everywhere and one confident misread of a renewal deadline costs more than a year of the subscription. These ten rules are a calibration, not a warning — each tells you precisely how far to trust an AI reading of a lease, an LOI, or a purchase agreement, and which documents to keep out of the pipeline entirely.
Rule 1: Trust the extraction, verify the interpretation
AI is a strong reader and a weak reasoner, and the trust line runs between those two jobs. Pulling the base rent, the commencement date, or the tenant’s legal name off a page is extraction — mechanical and reliable on clean text. Deciding whether a co-tenancy clause has triggered, whether a CPI escalation with a floor and a cap applies this year, or whether an option was validly exercised is interpretation — and that is where the model guesses.
The distinction matters because vendors sell both under one accuracy number. A tool marketed at 95% is describing extraction of standard fields, not the judgment calls that decide money. When an AI gives you a plain value that sits in one place on the page, accept it after a glance; when it gives you a conclusion that stitched several clauses together, treat that as a draft opinion to confirm, not a fact to file.
Rule 2: Trust nothing you cannot click back to the source
A trustworthy AI output links every value to the exact clause it came from — click the rent figure, land on the paragraph it was read from. This feature, called grounding or source-linking, is the mechanism that makes trust safe instead of blind, because it turns verification from an hour of hunting through a PDF into seconds of confirming a highlighted line.
A chat window that summarizes a lease into a tidy paragraph is quietly dangerous: you cannot tell which sentence came from the document and which the model invented. Grounding is the reading layer every reliable pipeline is built on, the same foundation covered in our guide to turning lease stacks into structured data. If you cannot trace it, do not trust it.
Rule 3: Distrust the confident answer most
The failure mode that costs small firms money is not the AI saying “I don’t know.” It is the AI reporting a renewal option that does not exist, in the same assured tone it uses for a figure it read correctly. This is hallucination — plausible, confident, wrong — a property of how these models work, not a bug a better version retires.
The scale is documented. A 2024 Stanford study found general-purpose large language models hallucinated on 58% to 88% of specific legal queries, and more than 75% of the time when asked about a court’s core holding. A follow-up study of purpose-built legal tools with retrieval grounding found error rates still ran from roughly 17% to 33% — grounding lowers the rate but does not reach zero. The rule: the more certain an answer sounds, the more it deserves a look at the source, because fluency is exactly what a hallucination is good at. We break down the checks that catch these in our piece on document AI’s hidden hallucination problem.
Rule 4: Never let AI close the loop on a money clause or a date
Some outputs get read by a person before anyone relies on them; some get written straight into a rent roll or a critical-date calendar with no human in between. The rule is absolute: every clause that moves money or sets a deadline stays human-reviewed by design. A missed notice deadline can push a tenant into holdover at 125% to 150% of the last contractual rent; a CAM misclassification quietly overbills or underbills for years.
This is not a vote against automation — it is a decision about where the loop closes. Let AI draft the abstract, flag the dates, and pre-fill the rent schedule at full speed, then require a person to confirm the money clauses against the source before the data becomes operational. The clauses where AI accuracy drops most — co-tenancy, exclusives, percentage rent, unusual escalations — are the exact clauses with the largest financial consequence, which is why removing the human there is the one design choice to never accept.
Rule 5: Grade trust by field, not by document
“Review everything” is safe and slow; “trust everything” is fast and reckless. The workable middle is a trust grade per field type. Tenant name, suite number, square footage, base rent, and commencement date are high-trust extractions you confirm at a glance. Renewal mechanics, escalation formulas, co-tenancy, exclusives, assignment rights, and any defined term that references another exhibit are low-trust interpretations you verify against the clause every time.
Grade by field rather than by document because a single lease contains both kinds, and a reviewer who trusts the whole abstract because the easy fields looked right is the one who misses the hard clause that mattered. This is the discipline that separates firms that get value from AI abstraction from those whose projects quietly fail, a pattern we examine in why most lease abstraction projects fail.
Rule 6: Keep the documents AI cannot read out of the pipeline
Trust in the output depends on the tool’s ability to read the input, and some inputs defeat it. A faxed estoppel, a rotated scan of a 1990s ground lease, a purchase agreement annotated in pen — these are where the reading layer itself fails, and a value extracted from a page the tool could not actually parse is a guess dressed as data.
Route those documents to a human first. The tell is often silent: instead of refusing, a tool returns a clean-looking value read from a garbled page. Know which of your documents are high-risk to read — poor scans, handwriting, dense exhibits, anything predating reliable digital origination — and treat any AI reading of them as unverified.
Rule 7: Watch for the amendment that quietly rewrites the deal
An AI reads the document you give it; it does not know the deal’s history unless you assemble that history for it. A second amendment can change a renewal notice period the original lease set, a side letter can waive an escalation, and a tool that abstracts each document in isolation will confidently report a term the deal no longer operates under.
Never rely on a single-document abstract for a value a later document could have changed. Feed the full chain — original lease, every amendment, side letters, estoppels — and require the reviewer to confirm the AI reconciled conflicts in favor of the most recent controlling language. The current state of the deal is precisely the interpretation Rule 1 said never to take on faith.
Rule 8: Read the data terms before you upload a confidential lease
Trusting AI with a legal document also means trusting a vendor with confidential deal data, and that trust has to be verified in the contract, not assumed from the marketing page. The safe pattern is a business or enterprise tier where your inputs are excluded from model training by default and the vendor holds SOC 2 Type II or equivalent certification. ChatGPT Business and Enterprise and Claude Team and Enterprise both contractually exclude customer data from training and carry SOC 2 Type II; ask any proptech vendor for the same in writing.
The mistake small firms make is pasting an NDA-covered lease into a free consumer chat tool because it was convenient. Read the data-processing terms of the exact plan you are on, not the tier above it, and anonymize the most sensitive material where the analysis does not need names.
Rule 9: Require two people who can catch a wrong answer
Every rule above assumes a reviewer who can tell a right abstract from a wrong one, and that capability is the real precondition for trusting AI at all. A firm that deploys document AI before its people are fluent has automated a job no one can quality-check.
Set a floor of at least two people who can read an AI-generated abstract and say confidently where it is right and where it is guessing — two, not one, so the capability survives a vacation and a resignation. That fluency is inexpensive to build, with market-rate training running roughly $2,000 to $15,000, and it makes every other rule enforceable. The broader case for why this fluency is the durable edge a small firm holds over larger competitors runs through the small-firm AI playbook.
Rule 10: Earn trust from your own error log, not the vendor’s demo
Trust should be a number you compute from your own records, not a claim you accept from a sales deck that runs on a document chosen to succeed. Keep a simple log of every correction a reviewer makes: which field, which document type, what the tool said, what was right.
Within a few weeks that log tells you what a headline figure never will: which fields your tool nails, which it misses, and on which document types its output cannot be trusted at all. What “95% accurate” actually means for your stack is a question only your own data answers, a point we develop in decoding document AI accuracy claims. Let the log raise your trust where the tool has earned it and lower it where it has not — that is the difference between trusting AI and hoping it works.
The rules as a trust scorecard
The ten rules resolve into one operating question for every AI output: how far can I trust this, and what has to happen before I rely on it?
| # | Rule | Trust it when |
|---|---|---|
| 1 | Extraction vs. interpretation | It is a plain value in one place; verify anything stitched together |
| 2 | Traceable to source | You can click it back to the clause; distrust anything you cannot |
| 3 | Distrust confidence | Never on tone alone — fluency is what a hallucination is best at |
| 4 | Money and dates | A human has confirmed the clause against the source |
| 5 | Grade by field | The field is high-trust (name, rent, date); verify the low-trust ones every time |
| 6 | Readable input | The tool actually parsed a clean page; route bad scans to a person first |
| 7 | Full amendment chain | The current state was reconciled across every controlling document |
| 8 | Data terms | Inputs are excluded from training and the vendor holds SOC 2 or equivalent |
| 9 | Fluent reviewers | At least two people can judge the output |
| 10 | Your own error log | Your records, not the demo, put the number there |
No output has to clear all ten to be useful, but rules 4, 6, and 8 are non-negotiable: a money clause closed without human review, a value read off an unreadable page, or a confidential lease sent through unsafe data terms is a risk no time savings offset.
Frequently asked questions
Can you trust AI with legal documents at all?
Yes, but conditionally, and the condition is calibration rather than blanket faith. AI is reliable at extracting plain values from clean text and unreliable at interpreting clauses that require stitching several provisions together, so you trust the first kind on sight and verify the second against the source. Use it as a fast first-pass reader, and trust any output only as far as it traces back to a clause a person can confirm.
How accurate is AI at reading leases and legal documents?
On standard fields such as parties, dates, and base rent in clean documents, tools commonly report accuracy above 95%. Accuracy falls on non-standard clauses — co-tenancy, exclusives, percentage rent, unusual escalations — which are exactly the clauses that carry financial risk, and it falls further on poor scans. Treat headline figures as claims about the easy fields, and verify the money clauses against the source regardless of the number a vendor quotes.
What is an AI hallucination in a legal document context?
A hallucination is a confident, plausible, incorrect output — the AI reporting a renewal option that does not exist or reading a rent figure from the wrong clause, in the same assured tone it uses when it is right. A 2024 Stanford study found general-purpose models hallucinated on 58% to 88% of specific legal queries, and grounded legal tools still showed error rates from roughly 17% to 33%. Because it reads as fluently as a correct answer, the confident-sounding output is the one that most deserves a check against the source.
Which parts of a lease should never be automated without human review?
Any clause that moves money or sets a deadline: rent and escalation figures, renewal and notice dates, co-tenancy and exclusive provisions, percentage rent, assignment rights, and CAM classifications. A missed notice deadline can trigger holdover rent at 125% to 150% of the prior rate, and a CAM misclassification overbills or underbills for years. AI can draft and pre-fill these at full speed; a person closes the loop first.
Is it safe to upload confidential leases to ChatGPT or Claude?
It can be, on the right plan. ChatGPT Business and Enterprise and Claude Team and Enterprise contractually exclude your inputs from model training and hold SOC 2 Type II certification; free consumer tiers do not carry the same guarantees. Read the data-processing terms of the exact plan you are on, confirm training exclusion and a recognized security certification in writing, and anonymize the most sensitive material where the analysis does not need names.
How do I check an AI-generated lease abstract quickly?
Use grounding, also called source-linking: a good tool links every extracted value to the exact clause it came from, so a reviewer clicks a figure and lands on the source in seconds instead of hunting through the PDF. Confirm the high-trust fields at a glance, then spend the review time on the low-trust interpretations — escalations, renewal mechanics, co-tenancy. If an output cannot be traced to a source, treat it as unverified and re-derive it by hand.
What documents defeat AI lease abstraction?
Poor scans, faxed estoppels, rotated pages, handwriting, hand-annotated agreements, dense cross-referenced exhibits, and leases predating reliable digital origination. On these the reading layer itself fails, and instead of refusing, a tool often returns a clean-looking value read from a page it could not actually parse. A value extracted from an unreadable page is not data — it is a well-formatted assumption, so route those documents to a person first.
Do I need special software to trust AI with legal documents, or is a checklist enough?
The discipline matters more than the software tier. The same trust rules — verify interpretation, demand traceability, keep humans on money clauses, grade trust by field, log your own errors — apply whether you use ChatGPT or Claude with human verification, a purpose-built abstraction tool, or a custom pipeline. A firm with the discipline and a simple tool beats one with a powerful tool and blind trust.
Where to start
Before you decide how much to trust AI with your legal documents, get honest about two things: whether your team is fluent enough to judge what any tool produces, and which of your real documents will break it. A free AI-readiness assessment produces that read — a short working session that maps your document mix, your monthly volume, and the workflows where AI would touch money clauses, and returns a plain recommendation for where AI is safe to trust today and where it is not. Book a free AI-readiness assessment before you rely on a single AI-generated abstract.
Arthur Wandzel