“AI can read your documents” is true and useless. Your deal folder is not one kind of document. It holds a lease that was typed in Word last year, a first-generation lease that was scanned from a fax in 2003, a rent roll in Excel, a T-12 someone exported to PDF, an estoppel with a signature scrawled across the bottom, a site plan drawn in CAD, and a handful of phone photos of a roof. A tool that extracts every term from the first document flawlessly can quietly invent numbers on the third and has no idea what the sixth even shows. The useful question is not whether AI reads documents. It is which of your documents it reads reliably, which it reads only if a person checks it, and which it cannot be trusted with at all. This is that map.
Reading text is not the same as understanding the deal
Two different jobs hide inside the phrase “read a document.” The first is extraction: pulling stated facts off the page — base rent is $32.50 per square foot, the term expires June 30, 2029, the tenant is Acme LLC. The second is interpretation: judging what those facts mean — whether the co-tenancy clause is unusually tenant-favorable, whether a site plan shows an access easement crossing the parcel, whether a scrawled signature is the right officer’s.
Modern assistants are strong at the first job and unreliable at the second. Given clean text, ChatGPT, Claude, Gemini, or Microsoft Copilot will extract the terms of a lease accurately and fast. Ask the same tool whether that lease is a good deal, and you get a plausible-sounding opinion that no one should underwrite against.
Keep the two jobs separate as you read this guide. The tiers below rate how reliably AI can extract what a document literally says. They do not promise judgment. Even a green-tier document still needs a human to decide what the extracted facts mean for the deal — the same discipline our lease abstraction framework applies field by field.
The three tiers: green, yellow, red
The single variable that decides how well AI reads a CRE document is not how important the document is. It is the document’s form — specifically, whether the words exist as real text the machine can access, or as pixels it has to guess at.
- Green — reliable. Digitally native documents whose text is machine-readable. AI extracts stated terms at high accuracy; a light spot-check is enough.
- Yellow — reliable with verification. Scanned or image-based documents that require optical character recognition (OCR) to turn pixels into text, and documents with dense tables. AI usually gets them right, but the error rate is high enough that a human must confirm the fields that matter.
- Red — human-led. Handwriting, signatures, poor faxes, heavy redlines, drawings, and photographs. These are either not text at all or too degraded to trust. AI can assist, but a person owns the read.
The rest of this guide places your actual document stack into these three buckets and explains the mechanics behind each placement.
Green: documents AI reads reliably
A document is green when it was born digital and its text layer is intact — the words are selectable, not painted on. For these, extraction accuracy on standard content sits in the mid-to-high 90s with a well-built, reusable prompt, and the failures are rare enough that a quick review clears them.
Modern leases and amendments (digitally native). A lease generated in Word or a document platform in the last several years carries a clean text layer. Parties, premises, base rent, escalations, term, and options come out reliably. This is the anchor use case, and it is why abstracting a lease once into structured data works as well as it does.
LOIs, PSAs, and letter agreements (Word-generated). Deal correspondence and contracts produced in a word processor read cleanly. Extracting the economic and timing terms of a letter of intent or purchase agreement is squarely in scope.
Rent rolls and operating spreadsheets (Excel/CSV). When the source is an actual spreadsheet rather than a picture of one, the data is already structured. AI reads it directly and can reconcile it against a lease abstract or reformat it without OCR risk.
Digital offering memoranda and marketing PDFs. An OM exported from design software usually keeps a text layer for its body copy. The prose and the labeled figures extract well; be more careful with numbers that live inside embedded chart images, which cross into yellow.
Tax bills, insurance certificates, and utility statements that arrive as digital PDFs. Increasingly these are issued as native PDFs. When they are, the key figures — assessed value, premium, coverage limits, account totals — extract cleanly.
The green-tier rule of thumb: if you can highlight the text with your cursor and copy it, the machine can read it the same way you can.
Yellow: documents AI reads with a verification pass
Yellow documents are legible but not native. The words exist as an image, so the machine must run OCR to recover them, or the content sits in a dense table whose structure is easy to scramble. Accuracy is good and getting better, but “good” on a document where a single wrong digit costs five figures is not the same as safe. Extract freely; verify the fields that carry money or dates.
Scanned leases and older lease generations. A lease imaged from paper — common for anything more than a decade old, and for assets that changed hands several times — depends entirely on scan quality. A crisp 300-dpi scan reads nearly as well as native text. A faint, skewed, or double-copied scan starts dropping and transposing characters, and rent figures and dates are exactly the characters you cannot afford wrong.
T-12s and operating statements exported to PDF. These are tables of numbers, often with subtotals and multiple columns. OCR generally recovers the values, but table structure — which number belongs to which line item and month — is where AI slips. Confirm that totals foot and that line items land in the right rows.
Estoppel certificates and SNDAs. The typed body reads fine. The risk is that the operative facts — the confirmed rent, the stated security deposit, the acknowledged defaults — are the whole point of the document, so an OCR miss on those specific figures matters more than usual. Verify the numbers the estoppel exists to confirm.
Scanned tax bills, CAM reconciliations, and vendor statements. Image-based financial documents with tables. Extraction works; the standing instruction is to check the figures that feed a bill or a model, not the boilerplate.
Appraisals and third-party reports (as PDFs). Long documents that mix native text, scanned exhibits, and image-based tables. The narrative reads well; the comparable-sales grids and reconciliation tables, often embedded as images, need a human eye.
The yellow-tier rule of thumb: let AI extract everything, then check the specific values where a wrong character is expensive. That verify-at-capture discipline is the difference between an abstraction project that holds up and one that quietly rots, which is a large part of why most lease abstraction projects fail.
Red: documents AI can’t be trusted to read alone
Red documents are ones where the information is either not stored as text or too degraded and unstructured for extraction to be dependable. AI can still help — it can transcribe a guess, summarize a drawing’s labels, or flag what it thinks it sees — but a person owns the answer.
Handwritten notes, margin annotations, and redlines. Handwriting recognition on unconstrained cursive is unreliable, and CRE handwriting shows up exactly where it matters most: a negotiated change scrawled into a lease margin, an initialed rider, a broker’s note on a rent roll. Treat any handwritten term as human-read until proven otherwise.
Signatures and signature blocks. AI can tell you a signature is present. It cannot verify whose it is or whether the signatory had authority — an interpretation-and-judgment task, not an extraction one. Execution and authority checks stay with people and counsel.
Poor faxes and multi-generation copies. A document that has been printed, faxed, scanned, and re-copied loses character fidelity at each hop. Below a legibility threshold, OCR output is a guess dressed as data. If you struggle to read it, assume the machine does too.
CAD drawings, site plans, floor plans, and surveys. These are not text documents. Whether a survey shows an encroachment, whether a site plan preserves required parking, whether a floor plan matches the rentable area on the lease — these are read by a person who understands drawings, sometimes a surveyor or architect. AI can pull labeled text off a plan; it cannot reliably interpret the geometry.
Property photographs. A model can describe a photo in general terms. It cannot assess roof condition, deferred maintenance, or code issues from an image with the reliability a diligence file requires. Photos inform people; they do not feed an extraction pipeline.
Complex nested tables and stamped-over documents. Multi-level headers, merged cells, and text stamped across other text defeat clean extraction even when the underlying scan is decent. When structure is the information, and the structure is tangled, keep a human in the loop.
Why the hard documents break
Knowing why a document lands in yellow or red helps you predict the ones this guide doesn’t name. Four mechanics account for almost all of it.
No text layer. A scan or photo is pixels. The machine cannot read it until OCR converts those pixels into characters, and OCR introduces its own errors — worse on low resolution, skew, noise, faint ink, and unusual fonts. Everything downstream inherits those errors.
Lost table structure. Reading the values in a table is easier than preserving which value belongs to which row and column. When a T-12 or a comps grid is an image, the model can recover “$412,800” and still attach it to the wrong month or line item.
Handwriting and degradation. Cursive, initials, and faded multi-copy documents fall outside what recognition handles dependably. The output looks confident and can be wrong, which is the dangerous combination.
Drawings and images are not language. A site plan encodes meaning in geometry, symbols, and scale, not sentences. Extracting text labels off it is not the same as understanding what it depicts, and only the latter answers a diligence question.
This is the same reality the broader document intelligence playbook is built around: match the tool to the document’s form, and put people where the form defeats the machine.
The rule that falls out of the tiers
The tiers are not trivia. They resolve into one operating rule a firm with no IT department can actually run.
Let AI extract from green and yellow documents, verify the yellow figures that carry money or dates at the moment of capture, and keep red documents human-led with AI as an assistant only. Then land the verified results in one structured place so the work is done once rather than repeated at every query.
The payoff of the tiering is focus. A small team cannot re-read every field of every document, and it does not have to. Green documents need a light spot-check. Yellow documents need a targeted check of the expensive fields. Red documents need a person from the start. Spending review attention that way — heavy where the form is risky, light where it is safe — is how a lean firm gets institutional-grade diligence without institutional headcount, the operating posture behind our manifesto on how small CRE shops out-operate larger competitors.
Keep confidential documents in the right account
Every document in this guide is confidential deal material, often under an NDA. Where you read it matters as much as how well it reads.
Use a business-tier or enterprise account from a major provider, whose terms state that your inputs are not used to train the model by default — confirm your plan’s current terms, since they change — and never paste a sensitive lease, PSA, or diligence file into a free consumer account. A single written rule covering which document types are cleared for which account is enough governance for a firm of this size. The most common mistake at a small firm is an analyst dropping a strict-NDA document into the wrong tool to save a few minutes; the tiering above is worthless if the confidentiality basics are skipped.
FAQ
Which CRE documents can AI read most reliably?
Digitally native documents whose text is machine-readable: modern leases and amendments generated in Word, LOIs and purchase agreements, rent rolls and operating models in Excel, and offering memoranda or tax bills issued as native PDFs. For these, a current assistant extracts stated terms at high accuracy, and a light spot-check clears the occasional error. The test is simple — if you can highlight and copy the text with your cursor, the machine can read it the same way.
Which documents can’t AI read reliably?
Anything that is not stored as clean text or is too degraded to trust: handwritten notes and margin annotations, signatures, poor faxes and multi-generation copies, CAD drawings, site plans, surveys, floor plans, property photographs, and complex nested or stamped-over tables. Some of these are not language at all (drawings and photos), and some are text the machine can only guess at (cursive, faded scans). Treat all of them as human-led, with AI assisting at most.
Can AI read scanned leases, or only digital ones?
It can read scanned leases, but with a verification pass. A scan has no text layer, so the tool runs optical character recognition to recover the words, and OCR accuracy depends heavily on scan quality — a crisp 300-dpi scan reads nearly as well as native text, while a faint or skewed copy drops and transposes characters. Because the characters most likely to be misread include rent figures and dates, confirm those specific values against the source before trusting them.
How accurate is AI at extracting lease terms?
On standard, digitally native leases, a business-tier assistant with a well-built, reusable prompt extracts terms at accuracy in the mid-to-high 90s. That is high enough to save real time and low enough that the remaining error rate still lands on fields where a mistake is expensive. The reliable pattern is to let the model extract everything, then have a person verify the high-cost fields — rent, escalations, critical dates, and options — against the source clause.
What’s the difference between AI reading a document and understanding it?
Reading is extraction: pulling stated facts off the page, like the base rent or expiration date. Understanding is interpretation: judging what those facts mean, like whether a clause is unusually favorable or whether a site plan shows an easement. Current assistants are strong at extraction and unreliable at interpretation. A model can pull every term from a lease accurately and still give an untrustworthy opinion on whether the lease is a good deal.
Can AI read handwritten notes on a lease or rent roll?
Not dependably. Handwriting recognition on unconstrained cursive and initials is unreliable, and in CRE the handwriting tends to sit exactly where it matters — a negotiated change written into a margin, an initialed rider, a broker’s note on a rent roll. The output can look confident and be wrong, which is the worst failure mode. Treat any handwritten term as human-read until a person confirms it.
Can AI interpret site plans, surveys, or floor plans?
No, not for diligence-grade questions. These documents encode meaning in geometry, symbols, and scale rather than sentences, so extracting text labels off them is not the same as understanding what they depict. Whether a survey shows an encroachment or a site plan preserves required parking is read by a person who understands drawings, sometimes a surveyor or architect. AI can transcribe labels; it cannot reliably read the geometry.
Is it safe to upload confidential CRE documents to an AI tool?
With safeguards, yes. Use a business-tier or enterprise account from a major provider, which states that inputs are not used to train the model by default — verify your plan’s current terms, since they change — and never use a free consumer account for confidential deal material. Set one written rule for which document types are cleared for which account, and withhold or anonymize anything under a strict NDA beyond what the task needs.
Do I need special software, or can a general assistant read these documents?
For most small-firm document work, a general business-tier assistant handles the green and yellow tiers well with a good prompt and a verification habit. Purpose-built proptech platforms add maintained extraction, source-linking, and portfolio querying, which is worth buying when volume is high or an audit trail is required. Neither replaces the human read on red-tier documents. Match the spend to your volume rather than the other way around.
Key takeaways
- “AI can read documents” is too coarse to act on; what matters is which of your specific documents it reads reliably, which it reads only with a check, and which it can’t be trusted with alone.
- The deciding variable is the document’s form — whether the words exist as real text the machine can access, or as pixels it has to guess at — not the document’s importance.
- Green (digitally native leases, LOIs, Excel rent rolls, digital OMs) extracts reliably with a light spot-check.
- Yellow (scanned leases, T-12s, estoppels, image-based financial tables) extracts well but needs a human check on the figures that carry money or dates.
- Red (handwriting, signatures, poor faxes, CAD drawings, site plans, photos, tangled tables) stays human-led, with AI assisting at most.
- Extraction is reliable; interpretation is not — even a clean read still needs a person to decide what the facts mean for the deal, in a business-tier account that keeps confidential documents safe.
Not sure which of your document workflows are safe to hand to AI and which still need a person? That depends on how your leases and diligence files actually arrive — native, scanned, or handwritten — which is exactly what a short working session sorts out. Book your free AI-readiness assessment →
Dirk Jan van Veen, PhD