An AI assistant will build you a comp set for any property you name, in any market, in any month. That is the problem. Comparable-sales analysis has an operating envelope — a band of market conditions where enough similar, recent transactions exist to bracket a value. Inside that envelope, AI is a genuine accelerator. Outside it, in a thin market or on a one-of-a-kind asset, the method itself stops producing a defensible number, and the tool that used to slow down and make you work no longer does. It returns a clean table with the same confidence it always has. This piece is about recognizing that boundary before you underwrite across it, and what a lean firm does the moment comps stop working.
Comps have an operating envelope, and AI hides its edges
The comparable-sales approach rests on one assumption: that a pool of genuinely similar properties has traded recently enough, and often enough, to reveal what the market pays. When that pool exists, the method works and any competent tool can speed it up. When it does not — a rural submarket with three sales a year, a specialty asset like a data center or a self-storage-to-industrial conversion, a submarket frozen mid-repricing — the approach has left its domain of validity. The number you can produce is no longer a market read; it is an extrapolation dressed as one.
For decades, the friction of doing comps by hand warned you when this happened. You would open the database, filter, and watch the screen empty out. Three sales, none within eighteen months, none the same product type. The scarcity was the signal. You felt the method straining because you were the one straining it. That warning is exactly what an AI assistant removes. Ask it for “recent comparable sales” in a thin market and it does not report scarcity — it produces a full-looking set, because a language model has no representation of “the data does not exist here.” It generates the most probable-looking answer to your prompt, and a confident table is more probable than an honest “there are not enough comps to value this.”
The comp-selection discipline itself does not change when AI enters the workflow; we lay out the full five-step version in our comp selection framework for an AI-assisted world. What changes at the thin-market boundary is that the discipline’s most important output — the empty screen, the honest “not enough” — is the one thing AI is structurally unable to give you unless you force it.
Why AI fails silently in thin markets
It helps to be precise about the failure, because “AI hallucinates” is too vague to underwrite around. The failure in thin markets is not random. It is structural, and it has three mechanical parts.
First, a model answers the prompt you gave it, not the one you should have asked. You asked for comparable sales. The honest answer is often “there are not enough to support a value.” But that answer does not look like a comp set, and the model is optimizing for a response that looks like what you requested. Absent explicit instruction to report scarcity, it fills the gap.
Second, the output carries no visible signal separating a sourced comp from a stretched or invented one. A transaction pulled from a real record and a transaction the model reached for to round out a thin set render identically — same columns, same decimals, same tone. The distinction that matters most in a thin market is the one the interface erases.
Third, confidence is decoupled from evidence. A model states a five-comp set with the same assurance whether it drew on fifty real transactions or three plus filler. In a deep market that decoupling is harmless because the underlying data is sound. In a thin one it is the whole risk: fluency masquerades as support. Telling real market signal from confident filler is the same skill you need reading any AI-drafted analysis, which we cover in reading AI market reports for what is signal and what is filler.
Put together, these mean the thin market is precisely where AI is most dangerous and feels least so. The deep market, where you barely need help, is where it is safest. That inversion is what makes the boundary worth learning to see.
Five signals that comps have stopped working
You do not need an appraisal license to recognize the boundary. You need a short list of signals to check before you trust any comp set, AI-assisted or not. Any one of these should slow you down; two or more means the method has likely left its envelope.
- Thin count. Fewer than roughly four to six genuinely comparable sales after honest screening. A set padded back up to a comfortable number is a set with filler in it.
- Stale transactions. Nothing within the last twelve months, or a market that has clearly repriced since the comps closed. Old comps in a moving market are data about a world that no longer exists.
- Forced geography. You had to reach two or three submarkets over, or into a different metro, to fill the set. Each mile of reach is an adjustment you now cannot support tightly.
- Asset singularity. The property has a feature that drives most of its value and that few or no comps share — special-purpose construction, an unusual tenancy, an entitlement or conversion play. There is no substitution set because there is nothing to substitute.
- Adjustments larger than the spread. When the net adjustments you are making exceed the price spread between the raw comps, the comps are no longer doing the work — your assumptions are. At that point you are valuing the property with judgment and using the comps as decoration.
Run these five against the deal, not against the AI’s output, because the AI’s output will look clean regardless. If the underlying market fails the test, a tidy table does not rescue it.
The AVM confidence-score trap
Automated valuation models and comp tools that emit a confidence score — a forecast standard deviation, a high/medium/low band, a percentage — seem to solve this. The vendor tells you when to trust the number. The trap is in how a busy underwriter reads a low score: as a haircut on precision, a wider error bar to note and move past. On a thin market or a unique asset, that reading is wrong.
A low-confidence AVM output usually means the model could not assemble a supportable comp set — the exact condition this article is about. It is not telling you the value is 620 dollars per square foot give or take a bit more than usual. It is telling you the comparable-sales method does not apply to this property right now. The correct response to a genuinely low confidence flag is not to widen the interval and proceed. It is to stop using that number as a value and switch approaches. Vendor documentation from providers like HouseCanary and CoreLogic is explicit that confidence degrades as comparable transactions grow scarce or heterogeneous; the score is doing its job. The failure is treating a “the method broke” signal as a “the method is a little fuzzy” signal. When the tool tells you it is not confident, believe it more than you believe the number next to it.
What to do when the comps run out
Recognizing the boundary is only useful if you have somewhere to go. When comps stop working, a lean firm has three moves, in order of preference.
Widen deliberately, and price the widening. Extend the geography or the time window on purpose, with your eyes open, and adjust harder for every difference you introduced. The key word is deliberately: you are choosing to add a distant comp and explicitly discounting its weight, not letting a tool quietly slip it in. Document each reach and each adjustment. A widened set can support a value if the widening is visible and defended; it cannot if it is hidden.
Change approaches. This is what appraisers do when sales comps run dry, and it is available to a non-appraiser too. Lean on the income approach — value the asset off its cash flow and a supportable cap rate — where the property produces income and rent evidence is stronger than sale evidence. For special-purpose properties, the cost approach (land plus depreciated replacement cost) sometimes carries more weight than any comp. The point is that comparable sales is one of three approaches to value, not the only one, and its failure is a cue to re-approach rather than to force it.
Value it as a range and say so. Sometimes the honest output is a wide range with the thinness stated as a limitation, not a point estimate with false precision. An investment committee or a lender can work with “we believe this sits between X and Y, and here is why the comps cannot narrow it further.” They cannot work with a confident single number that quietly rests on invented support and blows up in diligence. Stating the limitation is a mark of a disciplined shop, not a weak one. Where comp selection sits inside the broader underwriting sequence — and where each of these fallbacks fits — is laid out in our deal-analysis playbook for lean teams.
Where AI still earns its place past the boundary
Leaving the comp envelope does not mean putting the assistant away. It means moving it off the judgment and onto the mechanical work, which is where it belonged all along. Past the boundary, AI is still genuinely useful — as long as it is working on data and structure you handed it, never on transactions it sourced itself. The division of labor is the same one that keeps AI-assisted comps safe in normal markets; it just matters more when the market is thin.
| Task at the boundary | AI does this well | A human must own this |
|---|---|---|
| Building an income model | Structure the cash-flow model, run scenarios, check the math | Set the rent, vacancy, and cap-rate assumptions |
| Sourcing what few comps exist | Deduplicate and standardize a real export | Confirm each comp is real; refuse to invent the rest |
| Documenting the widening | Draft the “why we reached” narrative from your reasons | Decide how far to reach and how hard to discount |
| Writing the limitation memo | Draft clear language stating the comp scarcity | Judge that the method has failed and sign off |
| Pressure-testing the range | Argue both ends, surface what you missed | Own the final range and the recommendation |
Notice that in every row the AI works on inputs a person supplied and validated. That is the constant. The reason it survives being pointed at a thin market is that you never let it source the thing that is scarce. Knowing which underlying data sources are real and usable in the first place is its own skill, and we walk through it in our field guide to CRE data sources AI can actually use.
Make the boundary part of your process
The last step is to stop relying on whoever happens to run a deal to remember all of this. A firm that out-operates larger competitors does it by building judgment into a process rather than hoping for it case by case — the theme of our work on how small CRE firms punch above their weight with AI. For comps, that means a short standing rule: before any AI-assisted comp set becomes a value, someone runs the five signals; if two or more fire, the deal switches to a widen-or-re-approach path and the memo states the comp limitation. It costs a few minutes per deal and it catches the occasional underwriting where a clean table hides a market that could not support it.
That single discipline contains the risk without giving up the speed. In deep markets, where comps genuinely exist, AI compresses the work and the signals stay quiet. In thin ones, the signals fire, the method switches, and the assistant moves to the tasks it can do safely. Comp errors do not stay contained — a value sizes a loan and commits capital — so the firm that knows exactly where its comps stop working is the one that can move fast everywhere they do.
FAQ
When do AI comps stop working?
AI comps stop working the moment the comparable-sales method itself leaves its operating envelope: too few genuinely similar, recent transactions to bracket a value. That happens in thin markets — rural or low-liquidity submarkets, quiet periods, mid-repricing markets — and on unique or special-purpose assets that have no real substitution set. The tool will still produce a full-looking comp set in those conditions; that is the danger. The method has failed even though the output looks unchanged.
Why does AI produce comps that do not exist?
Because a language model answers the prompt it was given rather than reporting when the honest answer is “not enough data.” Asked for comparable sales, it generates the most probable-looking response, and a complete table is more probable than an admission of scarcity. It has no internal representation of “these transactions do not exist here,” so it fills the gap with plausible fill. The output shows no visible difference between a sourced comp and an invented one, which is why thin markets are where this failure does the most damage.
How do I know if I am in a thin market for comps?
Check five signals against the deal, not against the AI’s output: fewer than roughly four to six comparable sales after honest screening; nothing within the last twelve months, or a market that has repriced since; comps you had to reach two or three submarkets away to find; a property whose main value driver few or no comps share; and net adjustments larger than the price spread between the raw comps. Any one should slow you down. Two or more means the comp method has likely stopped working.
Should I trust an AVM confidence score?
Trust it more than the number beside it, and read a low score correctly. A low-confidence flag usually means the model could not assemble a supportable comp set — the same thin-market condition this article describes. It is not a slightly wider error bar to note and move past; it is a signal that comparable sales does not apply to this property right now. The right response is to switch approaches, not to widen the interval and proceed with the number anyway.
What should I do when comps run out?
Three moves, in order. Widen deliberately — extend geography or time on purpose, adjust harder for the added difference, and document every reach. Change approaches — lean on the income approach where the asset produces cash flow, or the cost approach for special-purpose property, since comparable sales is one of three approaches to value, not the only one. Or value it as a stated range with the comp scarcity named as a limitation, rather than a false-precision point estimate.
Can I still use AI once comps have failed?
Yes, on the mechanical work, never on the judgment. Past the boundary a general assistant can structure an income model, run scenarios, standardize the few real comps you have, draft the “why we widened” narrative, and write the limitation memo. In every case it works on inputs you supplied and validated. What it must never do is source the scarce thing itself — the transactions. The rule that keeps AI safe in normal markets holds harder in thin ones: the AI reasons over your data; it never invents it.
Is the income approach better than comps for unique assets?
Often, yes. When a property is special-purpose or sits in a market with too few sales, the income approach — valuing off cash flow and a supportable cap rate — usually rests on stronger evidence than a stretched comp set, provided the asset produces income and you have credible rent data. For non-income special-purpose property, the cost approach can carry more weight. The failure of comparable sales is a cue to re-approach, not to force comps that have stopped being informative.
How is this different from AI just making mistakes?
A general mistake is random and often catchable. This is structural and silent. In a thin market the model fails in a specific, predictable way — it fills a genuine data vacuum with confident fill because it cannot represent absence — and the output looks identical to a sound one. You cannot catch it by inspecting the table, because the table is clean. You catch it by testing the underlying market with the five signals before you trust any comp set. That is why the boundary has to be a process rule, not a spot check.
How do I build this into how my firm underwrites?
Add one standing step: before any AI-assisted comp set becomes a value, someone runs the five thin-market signals, and if two or more fire, the deal moves to a widen-or-re-approach path with the comp limitation stated in the memo. It takes a few minutes per deal and does not depend on whoever happens to run the file remembering to be careful. Building the judgment into the process rather than the person is how a lean firm gets institutional-grade discipline without an appraisal department.
Key takeaways
- Comparable sales has an operating envelope; in thin markets and on unique assets the method itself stops working, and AI removes the friction that used to warn you when you crossed the boundary.
- The failure is structural, not random: a model fills a real data vacuum with confident fill because it has no way to represent “the data does not exist,” and the output looks identical to a sound set.
- Test the market, not the AI’s table, with five signals — thin count, stale transactions, forced geography, asset singularity, and adjustments larger than the comp spread. Two or more means comps have stopped working.
- Read a low AVM confidence score as “the method does not apply,” not “the number is a little fuzzy,” and switch approaches rather than widening the interval and proceeding.
- When comps run out, widen deliberately, re-approach through income or cost, or state a range with the limitation named — and keep AI on the mechanical work while a human owns every judgment.
Not sure where AI is quietly crossing that boundary in your firm’s underwriting — where a clean comp table is hiding a market that cannot support it? A short assessment maps that against your deal volume, property types, and markets faster than any tool trial. Book your free AI-readiness assessment →
Dirk Jan van Veen, PhD