Most lease abstraction projects at small commercial real estate firms do not fail because the AI reads the document wrong. They fail before a single lease is processed and after the demo ends — in the parts of the work no vendor sells you. A firm decides to “abstract the portfolio,” buys a tool or a service, watches a clean demo on a clean lease, and then hits the real stack: the scanned 1998 lease with six amendments, the hand-marked escalation clause, the estoppel that contradicts the abstract. The model was never the problem. The operating model was. The failures are predictable, and they cluster into five root causes that have almost nothing to do with which product you picked. Name them before you spend, and the project that stalls for most firms becomes the one that ships for yours.
What a lease abstraction project actually is
Lease abstraction is the work of reading a commercial lease and pulling its terms into a structured record: parties and premises, square footage, commencement and expiration dates, the base rent schedule, escalations, renewal and expansion options, the expense structure and base year, tenant-improvement allowances, security deposit, assignment and sublet rights, use and exclusive clauses, co-tenancy, and termination rights. A property manager needs those fields to bill correctly and never miss an option date. An acquisitions team needs them to underwrite. The abstract is the structured data every downstream decision runs on.
A “project” is the decision to do this at scale with software instead of by hand — across a portfolio, a new management assignment, or a diligence set. The reading layer that makes it possible is the same one covered in our guide to turning lease stacks into structured data. What follows is not about that reading layer. It is about the five things around it that decide whether the project succeeds, and the surveys back the concern: across industries, the majority of AI initiatives stall or fail to reach production, and the cause named again and again is organizational, not technical. Lease abstraction is a textbook case of that pattern in a small CRE firm.
Failure one: scoping by document count, not field standard
The most common way an abstraction project fails is at the moment it is scoped, and no one notices for weeks. A firm says “we have 400 leases to abstract” and treats the number of documents as the definition of the job. It is not. The document count tells you the volume; it tells you nothing about what “abstracted correctly” means, and that missing definition is where the project quietly breaks.
Consider one field: rent escalations. Across a real portfolio you will find fixed annual bumps, CPI-indexed increases with a floor and a cap, stepped schedules written into a rent table, and percentage rent tied to sales. “Extract the escalation” is four different extraction problems, and if the firm never wrote down which structures it expects, which format the output must take, and how ambiguous or conflicting clauses get resolved, there is no standard against which any output can be called right or wrong. The tool returns something. No one can say whether it passed.
Without that field standard, three things happen in sequence. Scope creeps, because every new lease structure is a surprise the project never planned for. Acceptance disputes multiply, because the firm and the vendor disagree about what a correct abstract looks like. And trust collapses, because staff catch obvious misses and conclude the whole output is unreliable. The fix costs a day and precedes any tooling decision: agree the field list, define the acceptable output format for each field, and write down how the hard clauses get handled. What that standard should contain, field by field, is the subject of our companion piece on what to extract from a lease and why. Skip it and you have not scoped a project — you have bought a volume of work with no finish line.
Failure two: measuring accuracy on the easy fields
The second failure is trusting a number that is technically true and practically misleading. Vendors report abstraction accuracy above 95%, and on the fields that number describes — parties, premises, commencement and expiration dates, base rent — it generally holds. Those fields are stated plainly and consistently across leases, and a competent reader, human or machine, gets them right most of the time.
The money is not in those fields. A wrong base rent is embarrassing and obvious; someone catches it the first month. The expensive errors hide in the clauses that are written differently in every lease: a CPI escalation with an unusual floor, a renewal option with a notice window and a fair-market-rent reset, an expense-recovery structure with a gross-up provision, a co-tenancy clause that lets a tenant go dark or cut rent if an anchor leaves. These are exactly the fields where extraction accuracy drops, because they are non-standard, and they are exactly the fields where a mistake changes what you bill, what you owe, or what you underwrite. Accuracy is highest where the stakes are lowest and lowest where the stakes are highest.
A project that measures itself on aggregate accuracy will report a healthy score and still be dangerous, because the aggregate is dominated by easy fields the model nails. The measurement that matters is field-weighted: how often is the abstract right on the specific clauses that carry financial consequence? Set the acceptance test on those clauses, sample them by hand, and judge the tool on the fields that would cost you money — not on a portfolio average that flatters every product in the market.
Failure three: no one owns the exception queue
The third failure is structural, and it is the one small firms walk into most often because it has no obvious owner. A working abstraction pipeline does not return finished data — it returns a first pass plus a set of flags: low-confidence extractions, clauses it could not classify, documents that did not match the expected structure. That exception queue is not a defect. It is the system doing its job, telling you which abstracts need a human before anyone relies on them.
Someone has to work that queue. In a large firm, an analyst does. In a 4-to-20-person shop with no analyst to spare, the queue lands on a broker or a property manager who already has a full desk, and one of two things happens. Either no one reviews it, and unverified abstracts flow into the rent roll and the underwriting model carrying errors the flags were trying to prevent — which is worse than manual abstraction, because it wears the authority of a system. Or the queue backs up, the review never happens, the promised time savings never arrive, and the project is declared a failure that was actually an unstaffed one.
The exception owner is not optional overhead you can trim to make the numbers work; it is the load-bearing role in the whole design. Before a project starts, name the person, size the queue honestly against their real capacity, and build the review time into the plan. A realistic view of that ongoing effort — the review load, not just the build — is part of what our breakdown of a lease-abstraction automation project’s scope, timeline, and budget exists to make visible before you commit.
Failure four: buying speed the team cannot check
The fourth failure is subtle because it looks like success at first. The tool is fast. It returns abstracts in minutes that used to take an hour each. The firm celebrates the speed and rolls it out — to a team that cannot yet tell a right abstract from a wrong one.
Reviewing an AI-generated lease abstract is a skill. It means reading the model’s output against the source clause and knowing when the two disagree in a way that matters: recognizing that the extracted renewal terms dropped the notice deadline, that the escalation was read as fixed when the lease indexes it, that the expense structure was labeled gross when the base year makes it something else. A person who can do that turns AI into a genuine accelerator. A person who cannot is simply approving output faster than they could ever verify it, and the speed becomes a liability — errors ship at machine pace instead of getting caught at human pace.
This is why fluency comes before tooling, not after. A team that can prompt ChatGPT, Claude, or Gemini to summarize a lease and can then read that summary critically against the document is ready to run an abstraction pipeline safely; a team that cannot is not, regardless of which product it buys. Building that judgment is a matter of a few focused sessions on the firm’s own leases — the entry point in the small-firm approach we lay out in the manifesto on how lean CRE shops out-operate larger competitors. The order is fixed: fluency first, then automate the work the fluent team already understands.
Failure five: treating a pipeline as a one-time build
The fifth failure is a mismatch of mental models. A firm treats abstraction as a project with a start and an end — abstract the portfolio, finish, move on — when the work that keeps the data correct is continuous. Leases are not static. New assignments bring new landlord paper. Amendments rewrite terms mid-term. The formats that the pipeline was tuned for on day one drift, and the underlying models change on the vendor’s schedule, not yours.
A pipeline that no one maintains decays quietly. The extraction that was accurate on last year’s lease templates degrades on this year’s, and nothing announces the drift — it surfaces the day a wrong option date means a renewal notice goes out late, or an amended escalation gets billed at the old rate. If the project was scoped as a one-time build with no maintenance owner and no re-check cadence, there is no one watching and no budget for the fix. The firm concludes the AI “stopped working,” when what actually happened is that a living system was treated as a finished one.
This is also the honest input to the build-versus-buy question. A one-time-feeling need often does not justify owning and maintaining a custom pipeline at all — a service or an off-the-shelf tool carries the maintenance for you, and the trade-offs among those routes are the whole subject of our three-way comparison of abstraction services, software, and custom automation. If you do decide to own accuracy-critical extraction that feeds your accounting, the maintenance obligation is exactly what our look at Trullion versus a custom lease-data extraction build weighs against the control a build gives you. Either way, plan for the tail, or the project ends the first time the paper changes.
The pattern underneath all five
Read the five failures together and the common thread is unmistakable: not one of them is a failure of the AI’s reading ability. The model can extract a co-tenancy clause. The pipeline can flag its own low-confidence output. The tools on the market are good enough that the reading layer is rarely the thing that breaks.
What breaks is the operating model around the reading layer — the definition of done, the acceptance test, the exception owner, the reviewer’s fluency, the maintenance plan. Every one of those is a governance and ownership decision the firm makes, or fails to make, independent of the vendor. That is why swapping tools almost never rescues a stalled abstraction project: the new tool inherits the same undefined standard, the same unstaffed queue, the same team that cannot check it. The question that actually predicts success is not “which abstraction product is best” but “is my firm set up so that any competent product can succeed here?” Get the operating model right and the tool choice becomes a detail. Get it wrong and no tool saves you.
How to de-risk before you spend
The de-risking sequence is short, cheap, and entirely upstream of buying anything.
Write the field standard first. List the fields you need, define the acceptable output for each, and write down how the hard clauses — escalations, options, expense recovery, co-tenancy — get handled. This is your definition of done and your acceptance test in one document.
Set the accuracy bar on the money clauses. Decide which fields carry financial consequence and require the tool to prove itself on those specifically, sampled by hand, not on a portfolio average.
Name the exception owner and size the queue. Identify the person who will review flagged abstracts, estimate the real weekly load against their actual capacity, and put it in the plan as work, not as an afterthought.
Build fluency before you automate. Make sure the people approving abstracts can read AI output critically against a lease. A few sessions on your own documents is enough to reach the bar, and it is the difference between an accelerator and a liability.
Plan for maintenance. Assume the paper will change and the models will update. Decide who re-checks accuracy and how often, or choose a service or tool that carries that burden for you.
Run one real, messy lease — not the clean demo lease — through whatever path you are considering, and check the output against the source on the clauses that would cost you money. That single test, against documents you have already worked, tells you more than any vendor comparison, because it exercises the exact operating-model questions the five failures come from.
Frequently asked questions
Why do most lease abstraction projects fail?
For operating-model reasons, not model accuracy. Projects fail because the firm scoped by document count instead of defining a field standard, measured accuracy on easy fields while errors hid in the money clauses, never staffed the exception queue that reviews flagged abstracts, rolled out to a team that could not yet check the output, or treated a living pipeline as a one-time build with no maintenance owner. The AI’s reading ability is rarely the cause. The governance and ownership around it usually is.
Is AI accurate enough for commercial lease abstraction?
On the standard fields — parties, premises, dates, base rent — accuracy is generally high, above 95% by common vendor reporting, and that holds up. It drops on the non-standard clauses: unusual escalations, renewal options with resets, expense-recovery structures, co-tenancy triggers. Those are exactly the fields with financial consequence, so a human should verify the money clauses against the source before anyone relies on the abstract. Treat the headline number as a claim about the easy fields.
What should we define before starting a lease abstraction project?
A field standard: the exact list of fields you need, the acceptable output format for each, and a written rule for how ambiguous or conflicting clauses get resolved. This becomes your definition of done and your acceptance test. Without it, you have bought a volume of work with no finish line, and scope creep and acceptance disputes are almost guaranteed.
Who should own lease abstraction quality control in a small firm?
A named person with the capacity to actually work the exception queue — the low-confidence and unclassified extractions the pipeline flags for review. In a 4-to-20-person firm with no analyst, this usually falls to a property manager or broker, so the review time has to be sized honestly and built into the plan. Unstaffed, the queue either backs up and stalls the project or gets skipped, letting unverified abstracts flow into the rent roll.
Should we build custom lease abstraction or buy a tool?
It depends on volume, standardization, and who maintains it between changes. A steady, high-volume, standardized portfolio can justify owning a pipeline; an episodic or one-time need usually does not, because a custom build carries a maintenance obligation a small firm cannot staff. An off-the-shelf tool or an outsourced service carries that burden for you. The trade-offs among service, software, and custom are worth working through before committing to a build.
Can we just use ChatGPT or Claude to abstract leases?
For low volume and ad hoc work, a thin workflow with saved prompts in ChatGPT, Claude, or Gemini can produce useful first-pass abstracts, with a person verifying the terms that matter. What you give up is a queryable data store, built-in confidence flags, and audit trails, which is why higher-volume, standardized operations move to a purpose-built tool or a service. The team still needs the fluency to read and correct the output either way.
How long does a lease abstraction project take?
The reading is the fast part; the operating-model work sets the real timeline. Defining the field standard, setting the acceptance test, staffing the exception queue, and building reviewer fluency take days to a few weeks depending on portfolio complexity, and skipping them is what turns a short project into a stalled one. A realistic scope-timeline-budget view plans for the review load and the maintenance tail, not just the initial pass.
What is the single biggest mistake firms make with lease abstraction AI?
Confusing a demo on a clean lease with production on a real stack. The demo lease is standard and legible; the portfolio is full of scanned documents, amendments, and non-standard clauses. A firm that buys on the demo and scopes on the document count has not tested the tool against the conditions that actually break abstraction, and discovers the gap only after the money is spent.
How do we know if our firm is ready to abstract leases with AI?
Ask whether you have a field standard, an accuracy bar on the money clauses, a named exception owner with capacity, a team that can read AI output critically, and a maintenance plan. If those are in place, most competent tools will succeed. If they are not, no tool will, and the readiness gap is the thing to close first.
Where to start
If any of the five failures sounds like a project you have already run or are about to start, the useful next step is not choosing a tool — it is checking whether your firm is set up so a tool can succeed. A free AI-readiness assessment does exactly that: a short working session that reviews your lease portfolio, your field requirements, your review capacity, and your team’s fluency, and returns an honest read on whether you are ready to abstract with AI, and if not, which of the five gaps to close first. Book a free AI-readiness assessment before you commit a budget to abstraction, so the project you fund is one that ships.
Dirk Jan van Veen, PhD