Home About Who We Are Team Services Startups Businesses Enterprise Case Studies Industries Commercial Real Estate Blog Guides Contact Connect with Us
All Commercial Real Estate guides
Real Estate 17 min read

The 10 Rules of a Fair AI Vendor Evaluation

The 10 Rules of a Fair AI Vendor Evaluation

A fair AI vendor evaluation is one where the best tool wins, not the best demo. The decision your firm makes is only as good as the process that produced it, and most small-firm evaluations are quietly rigged — by a slick presentation, by the vendor who called last, or by the tool a partner already likes. These ten rules fix the process, not the pitch. They give every option the same test on your own documents, force a score before the demo wears off, and keep the incumbent honest, so that when you sign, you know the winner earned it. This is written for a principal at a 4–20 person commercial real estate firm who makes the call in a partner meeting, without a procurement team or an IT department to run cover.

Most vendor-evaluation advice assumes a buyer you are not: an enterprise with a procurement office, a security officer, and an RFP process that runs for a quarter. A ten-person brokerage has none of that. The decision happens fast, over a couple of meetings, on the strength of two or three demos and a gut feeling — which is exactly the condition under which a good demo beats a good tool. The rules below scale a fair process down to a firm that decides on a Thursday. They pair with the single-vendor checklist of questions to ask any AI vendor; that checklist is the instrument you run on each tool, and these rules are how you run the whole comparison so the result holds up.

What “Fair” Means Here

“Fair” has two senses, and this article is about one of them. The first sense — algorithmic fairness, whether a model treats people equitably — matters, and it belongs in your data-terms review. The second sense is procedural: a fair fight between the tools you are comparing, where each gets the same test and the winner is the best fit rather than the best-connected salesperson. That second sense is what small firms get wrong, and it is what these rules protect.

A fair evaluation cuts both ways. It protects your firm from being sold, and it protects a genuinely good tool from being dismissed because it demoed second or because a partner had already decided. The goal is not a slower process. It is a process where the outcome would survive being questioned.

The 10 Rules

1. Define the Workflow and the Pass Bar Before Any Demo

Write down the one workflow you are trying to fix and the bar a tool must clear — before you watch a single presentation. Name the task (abstracting leases, ranking inbound deals, reconciling CAM, drafting market write-ups), how often it happens, and what “good enough to buy” looks like in a number you can check: reads nine of ten of our leases correctly, cuts the task from three hours to under one. Deciding the bar in advance is what stops a demo from redefining success on your behalf. If you have not yet settled which workflow deserves the attention at all, the discipline of mapping the workflow before shopping for a tool is the step before this one.

2. Give Every Vendor the Identical Test Set

Fairness starts with the same exam for everyone. Assemble a fixed set of your own materials — ten real documents, a live scenario, the messiest files you own — and make every vendor run that exact set. The two worst files earn their place: a scanned T-12, a lease with three amendments, the broker blast with forty attachments. Tools read tidy sample deals well and stumble on the ugly reality that fills a CRE inbox, so a demo on the vendor’s clean example tells you nothing about your work. When each tool faces the identical set, the comparison is real; when each shows you its own best case, you are grading three different exams and calling it a bake-off.

3. Weight the Criteria Before the Demos, Not After

Decide what matters and how much before you are charmed. List the criteria — accuracy on your documents, data terms, how it connects to Outlook and Excel and your CRM, twelve-month cost, support — and assign each a weight that reflects your firm’s priorities. Do this on paper first. Weights set after the demos are not criteria; they are a rationalization of the tool you already liked. A simple scorecard, each criterion rated one to five and multiplied by its weight, turns a fuzzy impression into a number you can defend in a partner meeting and revisit in a year.

4. Score Each Vendor Immediately and in Writing

Fill in the scorecard the moment a demo or trial ends, while the detail is fresh, and have each partner score independently before you compare. Memory decays fast and unevenly: the vendor who presented last, or most confidently, gets a halo that has nothing to do with the tool. Independent scores captured on the spot beat that drift. Where two partners diverge sharply on the same criterion, the gap is a flag to investigate, not an average to paper over. The written record is also what lets you reopen the decision later without relitigating it from memory.

5. Judge the Demo You Drive, Not the One They Scripted

A scripted demo is a sales asset built to hide the seams. Send your scenarios in advance, and in the meeting ask to drive — your document, your edge case, your question typed live. Watch what the tool does when a figure is missing or a scan is unreadable: does it flag the gap, or invent a clean-looking number to fill it? The second behavior is the expensive one, because it looks finished. Score only what runs live in front of you. A vendor confident in the product will hand you the keyboard; one who insists on the rehearsed path is protecting something.

6. Run a Time-Boxed Trial on Real Data

A demo is a pitch in a controlled room; a trial is a stress test in your world. For any tool that will touch real work, run a short, time-boxed proof on live data — one or two weeks, a defined slice of actual documents or deals, a clear success bar from Rule 1. Firms that pilot before they commit catch the integration snags and accuracy gaps that a demo hides, and they walk away from bad fits before the money is spent. Keep the trial boxed so it does not sprawl into a free consulting engagement. If the tool cannot clear your bar on a small, real sample, it will not clear it at scale.

7. Price Every Option on the Same Twelve-Month Basis

Compare full costs, not stickers. The monthly fee is rarely the real number: setup, integration, training your team, and the hours the tool consumes before it saves any usually make up most of the first year. Put every option on one twelve-month, all-in basis that includes your own time. A subscription’s true cost is the fee plus onboarding plus the staff hours it eats; a custom build generally runs a one-time project in the range of roughly $25,000 to $150,000 depending on scope, plus a modest monthly run. Only a common denominator makes the comparison honest — and it often reveals that the cheapest sticker is not the cheapest tool.

8. Put the Incumbent in the Bracket

The fairest evaluations include the option nobody is selling: keep doing it the way you do now. “Do nothing,” “hire a part-time analyst,” and “write a prompt over a tool we already pay for” are legitimate contestants, and a fair bracket scores them on the same card as the vendors. Firms skip this because no salesperson champions the status quo, so it wins or loses by default instead of on merit. Score it honestly and it will sometimes beat the shiny option — a well-built prompt over ChatGPT or Claude, run against a spreadsheet you already own, can clear the bar for a low-volume task at no new cost. Deciding whether a real tool is even warranted is the heart of the buy-versus-build question.

9. Separate the Tool From the Salesperson

Grade the product and the relationship on separate lines. A likable rep and a responsive team are worth something — support is the difference between a working tool and a dead login for a firm with no IT department — but they are a distinct criterion from whether the tool reads your leases correctly. Blur the two and a great salesperson carries a mediocre product across the line. Check references yourself with pointed questions about accuracy, hidden costs, and what broke after signing, and score the answers on their own row. If you are weighing a custom-build partner, the same discipline surfaces the tells that separate a real partner from a vendor who disappears at handoff, which the guide to reading an AI project proposal works through in detail.

10. Write the Decision Down

Close the evaluation with a short written verdict: what won, what it beat, on which criteria, and what would change your mind. A page is enough. It converts a meeting into a record you can defend to a skeptical partner, revisit when the tool underdelivers, and reuse the next time you evaluate anything. It also exposes a weak decision to yourself — if you cannot write down why the winner won, the process was not fair, it was fast. Buying capability deliberately, and being able to say why, is the habit that lets a small firm out-operate much larger competitors rather than accumulate a drawer of unused logins.

Same Rules, Three Kinds of Vendor

The ten rules apply whether you are weighing an off-the-shelf subscription, a custom build, or standing pat. What a strong answer looks like differs by path, so the same test produces fair — not identical — expectations.

Rule Subscription Custom build Status quo / DIY
Identical test set Runs on your worst files in a trial Tuned to your documents during the build A prompt over a general tool, tried on the same files
Pass bar Clears the accuracy bar out of the box Clears it after tuning, verified at handoff Clears it well enough for a low-volume task
Twelve-month cost Fee plus onboarding plus your hours One-time project plus a modest monthly run Near zero new cost, plus your time
Support Vendor tier with response times A named retainer and a model-update plan You maintain it
When it wins A standard product fits your volume The workflow is specific and crosses systems Volume is low or the task is rare

Buy when a standard tool fits your asset classes and volume, because a subscription starts faster and cheaper. Build only when your workflow is specific enough that no product fits, or when it crosses systems no vendor connects — and confirm your firm is ready for that commitment using the readiness framework for custom automation. Keep the status quo when the honest scorecard says the problem is not yet worth paying to solve. A fair process that returns “none of these, not yet” has done its job.

Frequently Asked Questions

What does a fair AI vendor evaluation actually mean?

It means a process where the best-fitting tool wins rather than the best demo or the best-connected salesperson. Every option gets the same test on your own documents, the criteria and their weights are set before the demos, each tool is scored immediately and independently, and the incumbent or “do nothing” option competes on the same card. Fairness here is procedural — a fair fight between the tools — not the separate question of whether a model treats people equitably, which belongs in your data-terms review. A fair process is one whose outcome would survive being questioned by a skeptical partner.

How do I compare AI vendors apples-to-apples?

Give every vendor the identical test set — the same ten real documents or the same live scenario, including your two ugliest files — and score each on the same weighted criteria immediately after its demo or trial. When each tool shows you its own hand-picked sample instead, you are grading three different exams. The same-exam discipline, plus weights fixed in advance and scores captured while fresh, is what makes the comparison genuinely like-for-like rather than a contest of presentation skill.

How do I keep a good demo from biasing my decision?

Set your pass bar and criteria weights before you watch any demo, drive the demo yourself with your own documents rather than following the vendor’s script, and score each tool in writing the moment the meeting ends. A polished, confidently delivered demo earns a halo that has nothing to do with how the tool handles your worst lease. Deciding what “good” means in advance, testing on live edge cases, and locking in a score before the next vendor presents are the three defenses that stop the pitch from buying the decision.

Should a small firm use a weighted scorecard for a two-vendor decision?

Yes, and it takes about fifteen minutes to build. List your criteria — accuracy on your documents, data terms, integrations, twelve-month cost, support — weight each by importance, and rate every option one to five per criterion, multiplied by its weight. Even for two vendors, the scorecard converts a fuzzy impression into a number you can defend and revisit later. It also surfaces disagreement between partners on specific criteria, which is more useful than a vague overall preference that hides where the real doubt lives.

How long should an AI proof-of-concept take?

One to two weeks on a defined slice of real data is usually enough for a small firm. Long enough to expose integration snags and accuracy gaps a demo hides, short enough that it does not turn into an unpaid consulting project. Fix the success bar before it starts — the same bar from your criteria — and a clear slice of actual documents or deals to run against. If a tool cannot clear that bar on a small, real sample, it will not clear it at scale, and you have saved the cost of finding out the expensive way.

How many AI vendors should a small firm evaluate at once?

Two or three is the practical range. Fewer than two and you have no comparison; more than three and the evaluation stalls under its own weight, since each one needs the same test set, demo, and scoring. Shortlist to a manageable few using the single-vendor checklist as a first filter, then run the full fair process on the finalists. Include the status-quo option as one of the contestants so you are always comparing against the honest baseline of doing nothing new.

How do I evaluate the incumbent or “do nothing” option fairly?

Score it on the same card as the vendors. Estimate what the current approach costs in hours and errors, what a prompt over a tool you already pay for could do against the same test files, and whether either clears your pass bar. Because no salesperson argues for the status quo, firms let it win or lose by default rather than on merit — which is unfair in both directions. Sometimes the honest score says a well-built prompt over ChatGPT or Claude already clears the bar for a low-volume task; sometimes it confirms the manual cost is real and a purchase is justified.

Is it fair to test AI vendors on my messiest documents?

It is the only fair test. Your firm does not run on clean sample deals; it runs on scanned T-12s, amended leases, and cluttered inboxes, so those are the conditions that decide whether a tool actually helps. A vendor confident in the product will run your worst files; one who insists on demoing a tidy example is hiding the accuracy that matters to you. Testing on the ugly reality is not stacking the deck against the vendor — it is refusing to let the vendor stack it in their own favor.

How do I run a fair evaluation without an IT department or committee?

Scale the process, not the rigor. You do not need a procurement office to fix the workflow and pass bar in advance, hand every vendor the same ten files, weight your criteria on one page, and score each tool the day you see it. Two partners scoring independently is enough of a panel. The point of these rules is to replace the enterprise apparatus a small firm lacks with a lightweight discipline it can actually run — a couple of meetings, a shared scorecard, and a one-page written verdict.

What is the difference between a fair process and a slow one?

Speed and fairness are not opposites. A fair evaluation for a small firm can run in two or three weeks: a week to set the bar and assemble the test set, a few days of driven demos, a short trial on real data, and a written verdict. What makes it fair is structure, not duration — the same test, weights set in advance, scores captured while fresh, the incumbent in the bracket. A slow process without those elements is just a delayed gut call; a fast one with them is a defensible decision.

Where to Start

Every one of these rules assumes you already know which workflow to point the evaluation at — and most firms start collecting demos before they have decided which problem is worth solving first. That inventory is the real starting line: which of your workflows a tool should fix, in what order, and whether the honest answer is a subscription, a build, or a prompt you already own. A free AI-readiness assessment produces exactly that — a working session that maps your firm’s workflows, flags the ones where a tool pays for itself, and hands you a ranked plan with real cost ranges before you sit across from any salesperson. Book a free AI-readiness assessment if you want that map first. If the fair answer for your firm right now is “none of these yet,” the assessment will tell you so — and you will still leave with the plan.

Last Updated: Aug 15, 2026

AW

Arthur Wandzel

SFAI Labs helps companies build AI-powered products that work. We focus on practical solutions, not hype.

Make your firm fluent in AI — then automate what works

  • Hands-on training applied to LOIs, lease summaries, and market write-ups
  • Automation across documents, deals, communications, and back office
  • Built for 4–20-person firms with no IT department

Related articles