An AI can use a commercial real estate data source only when two things are true: the data arrives in a shape a machine can read without guessing, and every field traces back to a named, dated origin you could verify. Most of the CRE data a small firm touches fails one or both tests — a scanned rent roll, a broker’s free-text email, a PDF offering memorandum with numbers baked into an image. The vendor leaderboards rank sources by coverage and price, which is the wrong question for an AI-assisted shop. The right question is narrower and more useful: can my model read this cleanly, and can I trust where each number came from. This field guide sorts the CRE data landscape by that standard, because the constraint on AI usefulness is rarely how much data you have and almost always how usable it is — more than 80% of enterprise data is unstructured, and unstructured data is exactly where models invent numbers (AIMultiple, AI Data Quality 2026).
The one question that sorts every source
Before you evaluate any CRE data source for AI work, ask one thing: what shape does the data arrive in, and where did each number come from. Coverage, price, and asset-class focus matter for a human analyst. For a model, they are secondary to whether the data is structured enough to parse and traceable enough to trust.
Two properties decide it. Machine-readability is whether the data comes as fields a model can read deterministically — a JSON response, a spreadsheet, a clean export — versus prose or a picture it has to interpret. Provenance is whether each value carries a named, dated source you could open yourself. A source can be strong on one and weak on the other. A crisp API with no audit trail is fast but risky. A well-sourced PDF that only exists as a scan is trustworthy but unusable without a fragile extraction step.
The sources worth building a workflow on score well on both. Everything else is either a candidate for a careful extraction pass or a number you verify by hand before it reaches your model. Sort your whole data stack this way once and most of your tooling decisions make themselves.
The machine-readability ladder
CRE data arrives in four tiers, and where a source sits on this ladder predicts how much engineering — and how much hallucination risk — stands between you and a usable number.
Tier 1 — Structured feeds (API or clean export). Data delivered as fields: a JSON API response, a database export, a spreadsheet with consistent columns. A model reads these deterministically; there is nothing to interpret. This is the only tier you can automate end to end with confidence. Property-records APIs, rent-comp feeds, and foot-traffic exports live here.
Tier 2 — Templated documents. PDFs and files with a stable, predictable layout — a lender’s standard T-12 template, a platform-generated comp sheet. A model extracts these reliably because the structure repeats. Reliability drops the moment the template changes, so treat every extraction as needing a spot-check until you have seen the format hold.
Tier 3 — Scanned or free-form documents. A photographed rent roll, a broker’s bespoke offering memorandum, a lease PDF where numbers sit inside an image. Modern document-AI reads these far better than last-generation tools, but this is where errors enter: a transposed digit, a merged cell misread, a footnote dropped. Usable, but never without a verification step on the numbers that drive the deal.
Tier 4 — Unstructured prose. A broker’s email, a phone note, a paragraph of market color. There is no structure to parse, so a model treats the text as a suggestion, not a record. Fine for context; never a source of record for a figure you underwrite.
The practical rule: automate freely on Tier 1, automate with spot-checks on Tier 2, extract-then-verify on Tier 3, and never let Tier 4 supply a number that goes into a model. Knowing which tier a source belongs to tells you exactly how much human review it still needs.
The paid platforms, ranked by what an AI can do with them
The major CRE platforms differ less in coverage than in how cleanly their data leaves the platform. For a lean firm, exportability and structure matter more than the raw record count.
CoStar has the largest U.S. and global database — lease and sale comps, building attributes, tenant intelligence, ownership. Its depth is unmatched, but it is built for humans working inside its interface, and programmatic access through its API requires special arrangement (CoStar via PropAPIS). For an AI workflow, treat CoStar as a Tier 1 source only where you have a sanctioned export path; otherwise its data reaches your model as screenshots and copy-paste, which drops it to Tier 3.
Crexi functions as a deal marketplace showing what is actively marketed for sale or lease, rather than a comprehensive historical database. It is strong for live sourcing and accessible at lower cost than CoStar, which matters for a small shop. As an AI input, its structured listing and comp data is workable; just remember you are seeing on-market activity, not the full market.
Reonomy (owned by Altus Group) specializes in off-market ownership intelligence — it maps roughly 54M+ commercial properties and 30M+ owners, with entity, debt, and contact data delivered by web application and API (Reonomy). The API delivery makes it a genuine Tier 1 source for sourcing workflows, which is exactly why it shows up in off-market prospecting stacks.
HelloData is a strong example of a purpose-built structured feed for one asset class. It surveys more than 38.5M multifamily units daily from over 250,000 property websites and delivers rent, availability, fees, and concessions as weekly rent-comp updates through an API (HelloData). If you underwrite multifamily, this is the kind of source an AI can consume directly and repeatedly.
Placer.ai provides location intelligence and foot-traffic data — visit counts, trade-area demographics, consumer trends — as a structured feed that an AI can pull into a retail or mixed-use analysis (Placer.ai). It answers a question the comp platforms cannot: who actually goes to this location, and how has that changed.
The pattern across all of them: a platform is AI-usable to the degree it lets structured data out. Deep coverage locked behind a human-only interface is worth less to an automated workflow than a narrower feed with a clean API.
Document data: OMs, rent rolls, and T-12s
The documents at the center of every deal — offering memoranda, rent rolls, T-12 operating statements, leases — are where a lean firm spends most of its analysis time and where AI extraction pays off most. They are also where the machine-readability ladder matters most, because these arrive in every tier depending on who produced them.
A lender-generated T-12 on a standard template is often Tier 2: predictable structure, reliable extraction. The same firm’s rent roll, exported straight from property-management software, can be Tier 1. But an offering memorandum designed by a broker to look impressive is frequently Tier 3 — numbers set as graphics, inconsistent tables, footnotes that change the meaning of a figure two pages away.
The discipline is the same one that separates useful AI output from dangerous output everywhere else: extraction is not verification. A model can pull unit counts, in-place rents, expense lines, and lease dates out of these documents in seconds, which is a real gain over manual entry. It can also transpose a digit or misread a merged cell, and it will present the wrong number in the same confident tone as the right one. Let AI do the extraction; keep a human on the three or four figures the deal actually turns on — the rent roll total, the net operating income, the largest expense line, the lease-expiration schedule.
Free and public data a lean firm already has rights to
Before licensing a sixth paid feed, account for the public data you can already use — much of it structured, all of it traceable, and none of it billed monthly. The vendor listicles skip this because they sell subscriptions, but for a 4-20 person firm it is often the highest-provenance data in the building.
- County assessor and recorder records — parcel data, assessed values, recorded deeds, and mortgage filings. Availability and format vary by county, but where a jurisdiction offers a data portal or bulk export, this is authoritative Tier 1 or Tier 2 data with impeccable provenance: it is the public record.
- Census Bureau and Bureau of Labor Statistics — demographics, income, employment, and population trends as clean, downloadable, well-documented datasets. For trade-area and demand analysis, an AI can consume these directly, and the source is unimpeachable.
- Municipal zoning and permitting data — many cities publish zoning designations and building-permit activity through open-data portals, which surfaces supply-pipeline signals a paid feed may lag on.
- SEC filings — for deals involving public REITs or listed operators, EDGAR filings are structured, dated, and free.
None of this replaces a comps platform. It does mean a lean firm should map what it already has access to before assuming the answer is another subscription. Public data tends to score highest on provenance precisely because it is the record everything else is derived from.
Reconcile before you license
The most common data mistake a small firm makes is buying more data to fix a problem that is really about reconciliation. When your assessor record, your CoStar pull, and your own portfolio spreadsheet disagree about a property’s square footage, owner, or unit count, the issue is not missing data — it is that no source agrees on what the property is.
This is an entity-resolution problem, and it is the reason data-management platforms like Cherre exist: they ingest licensed sources, tax records, listings, and your own data, then resolve them to a single canonical property record (Cherre via NextAutomation). For a firm whose real problem is that five sources will not reconcile, that infrastructure is the answer — not a sixth feed that adds a sixth version of the truth.
For most 4-20 person firms, full data-warehousing infrastructure is more than the problem warrants. The lighter version of the same discipline: pick one source as the system of record for each field — the assessor for parcel and ownership, a comp platform for rents, your own file for actuals — and make your AI workflow reference that hierarchy rather than averaging across sources that were never meant to agree. An AI that knows which source wins for each field stops producing the contradictory numbers that erode trust in the whole workflow.
Fitting data sources into a deal process
A clean data stack is worth nothing sitting still; its value shows up inside a deal process that screens more opportunities without lowering the standard for what counts as a real number. The sourcing decision and the workflow decision are the same decision viewed from two ends.
The upstream question — when a platform like CoStar is enough and when a firm genuinely needs to assemble its own sources — is worked through in when CoStar is enough, and when you need your own data layer, and it pairs directly with the machine-readability lens here: the right stack is the one that feeds your model clean, traceable inputs at a cost your deal volume justifies. Downstream, the discipline for reading what an AI produces from that data — separating sourced signal from confident filler — is covered in decoding AI market reports: what’s signal, what’s filler. The full method for turning these inputs into screened and underwritten deals with a lean team sits in the deal-analysis playbook for small firms, and the broader case for how a disciplined 4-20 person shop out-operates far larger competitors is laid out in the small-firm AI manifesto.
The through-line: an AI is a multiplier on the quality of the data underneath it. Feed it Tier 1 sources with clean provenance and it screens deals faster than a team twice your size. Feed it scanned PDFs and unsourced prose and it manufactures confident errors at the same speed. The sourcing decision is the underwriting decision.
FAQ
What CRE data sources can an AI actually use?
The ones that arrive as structured, traceable data. That means API feeds and clean exports from platforms like Reonomy, HelloData, Crexi, and Placer.ai, structured public data such as county assessor records and Census datasets, and documents with predictable templates like a standard T-12. An AI can also extract from scanned rent rolls and offering memoranda, but those require a verification step on the key numbers. The dividing line is not the vendor’s reputation — it is whether the data comes in a shape a model can read without guessing and a source you could open yourself.
Is CoStar data machine-readable for AI?
Only where you have a sanctioned export or API path. CoStar has the deepest CRE database, but it is built for humans working inside its interface, and programmatic access requires special arrangement. When a firm gets CoStar data into a model by taking screenshots or copying and pasting, the data effectively drops to the scanned-document tier — usable, but requiring verification. If you have an approved export path, it becomes a clean structured source. Check your license terms before designing a workflow around it.
What is the difference between structured and unstructured CRE data?
Structured data arrives as fields a machine reads deterministically — a spreadsheet of rents, a JSON API response, a database export. Unstructured data is prose or images a model has to interpret — a broker’s email, a scanned lease, a paragraph of market color. Structured data is safe to automate because there is nothing to guess; unstructured data is where models invent numbers, because they fill gaps by prediction. More than 80% of enterprise data is unstructured, which is why data quality, not data quantity, is the real constraint on AI usefulness.
Can AI read rent rolls, T-12s, and offering memoranda?
Yes, and this is one of the highest-value uses of AI in CRE — but extraction is not verification. A model pulls unit counts, in-place rents, expense lines, and lease dates from these documents in seconds, far faster than manual entry. It can also transpose a digit or misread a merged cell and present the error with full confidence. The safe pattern is to let AI extract everything, then have a human confirm the three or four figures the deal turns on: the rent-roll total, the net operating income, the largest expense line, and the lease-expiration schedule.
What free CRE data sources can an AI use?
More than most firms realize. County assessor and recorder records provide parcel data, assessed values, and recorded deeds with impeccable provenance. Census Bureau and Bureau of Labor Statistics datasets give demographics, income, and employment as clean downloads. Municipal open-data portals publish zoning and building-permit activity, and SEC EDGAR filings cover public REITs and listed operators. Availability and format vary by jurisdiction, but where a data portal or export exists, this is often the highest-provenance data a lean firm has — because it is the public record everything else derives from.
Do I need CoStar if I am a small CRE firm?
Not necessarily. CoStar’s depth is genuinely valuable, but its cost and human-first interface may not fit a 4-20 person firm’s budget or AI workflow. Lower-cost structured sources — Crexi for on-market activity, Reonomy for off-market ownership, HelloData for multifamily rents, Placer.ai for foot traffic — combined with free public records often assemble a stack that feeds an AI workflow more cleanly than a single expensive subscription accessed by screenshot. The right answer depends on your asset classes and deal volume, not on which platform is largest.
Why does my AI give wrong numbers from CRE data?
Usually because it is working from unstructured or unsourced input. When you hand a model a scanned document, a free-text email, or a question it answers from training data rather than data you supplied, it fills gaps by prediction — and a predicted number looks identical to a real one. The fix is upstream: feed the model structured sources, keep a human on the figures that drive the deal, and never let it supply both a number and the narrative around it in one pass. Wrong numbers are almost always a data-shape problem, not a model problem.
What is entity resolution and why does it matter for CRE data?
Entity resolution is the process of deciding that records from different sources refer to the same property, owner, or tenant. It matters because your assessor record, your comp platform, and your own spreadsheet will often disagree about a property’s square footage, unit count, or owner — not because data is missing, but because none of them agree on what the property is. Data-management platforms like Cherre exist to resolve sources to one canonical record. For a small firm, the lighter version is to name one source as the system of record for each field and have your AI reference that hierarchy.
How many data sources does a small CRE firm actually need?
Fewer than the vendor leaderboards imply. A focused stack — one comp or marketplace source, one ownership-records source, the free public records for your markets, and clean exports from your own property-management software — covers most of what a 4-20 person firm underwrites. Adding sources past that point usually adds reconciliation problems faster than it adds insight. The goal is not maximum coverage; it is a small set of structured, traceable sources your AI workflow can consume without producing contradictory numbers.
Should a lean CRE firm build its own data layer or license one?
For most 4-20 person firms, neither extreme fits: full data-warehousing infrastructure is more than the problem warrants, and a single all-in-one platform rarely matches every asset class. The practical middle is to license a few structured feeds that fit your deal types, add the free public data you already have rights to, and impose a simple system-of-record hierarchy so your sources do not contradict each other. Reserve heavier infrastructure for the point where reconciliation, not sourcing, is your actual bottleneck.
Key takeaways
- An AI can use a CRE data source only when the data is machine-readable and every field traces to a named, dated origin — coverage and price are secondary to shape and provenance.
- Sort every source on a four-tier ladder: structured feeds (automate freely), templated documents (automate with spot-checks), scanned or free-form documents (extract then verify), and unstructured prose (never a source of record).
- Platforms are AI-usable to the degree they let structured data out; a narrow feed with a clean API beats deep coverage locked behind a human-only interface.
- Let AI extract from rent rolls, T-12s, and offering memoranda, but keep a human on the few figures the deal turns on — extraction is not verification.
- Account for free public data — county records, Census and BLS, zoning portals, SEC filings — before licensing another paid feed; it is often the highest-provenance data you have.
- When sources disagree, the problem is entity resolution, not missing data; name one system of record per field rather than buying a sixth source that adds a sixth version of the truth.
Want to know which of your current data sources an AI can safely automate against — and which are quietly feeding it errors? A short, free AI-readiness assessment maps your tools, data sources, and deal workflow, then shows where an AI-assisted process pays off and where it needs a human check. Book your free AI-readiness assessment → and we will size it to how your firm actually works.
Dirk Jan van Veen, PhD