Skip to content
       

Blog

How to Run a Property Management Software Evaluation: Test Your Portfolio, Not the Demo

How to Run a Property Management Software Evaluation: Test Your Portfolio, Not the Demo

The short answer

Most property software evaluations fail on method rather than on shortlist. The failure is structural: a demo is a controlled performance in which the vendor chooses the data, the entity structure and the scenario, and every platform looks capable under those conditions.

The fix is to invert who supplies the conditions. Bring your own worst month, your own messiest entity, your own most irregular lease, and score every vendor against the same set of them.

A demo tells you what a platform can do. Only your own data tells you what it will do for you.

This blog is about how to run the buying process. If you are still deciding whether to evaluate at all, the four triggers covers that question first, and our complete buyer's guide to evaluating property management software covers which criteria matter.

Why do software evaluations go wrong?

Because the process is usually designed by the people selling, not the people buying.

Enterprise procurement practice is unusually direct on this point. Analysis of vendor evaluation identifies letting demos drive the decision as the single biggest scoring mistake, on the reasoning that demos are effective at hiding architectural debt while written answers expose it. The recommended sequence inverts the usual one: issue the same written requirements to every vendor simultaneously, score against a weighted rubric, and reserve demos for verifying claims rather than forming them.

The risk of a demo-led selection is that a successful pilot gets mistaken for a successful deployment. Gartner predicted that at least 30% of generative AI projects would be abandoned after proof of concept by the end of 2025, citing poor data quality, inadequate risk controls, escalating costs or unclear business value. That research is specific to generative AI rather than property software, but it illustrates a broader evaluation lesson: a proof of concept that succeeds under controlled conditions does not by itself establish that a production deployment will.

Property has its own version of this, and it is specific. The demo dataset has one entity, clean history, standard residential leases and no open reconciliation. Your portfolio has none of those things. The evaluation therefore tests a business you do not run.

The Adversarial Evaluation Method

Five stages, each designed so the vendor cannot choose the conditions.

Stage

What you do

What it exposes

1. Assemble the hard cases

Document your worst entity, weirdest lease, hardest close

Where your complexity actually lives

2. Written before visual

Same questions to every vendor in writing, before any demo

Architectural gaps a demo can route around

3. Demo on your data

Vendors demonstrate against your cases, not their dataset

Whether capability is native or assembled

4. Reference the failures

Ask references what went wrong, not whether they are happy

The gap between scope and reality

5. Test the exit

Establish what leaving costs before entering

Whether the decision is reversible

The word adversarial is deliberate, and it is not hostile. A good vendor should welcome this, because it shortens their sales cycle and reduces the chance of an unhappy implementation. A vendor who resists it is telling you something useful.

Stage one: assemble the hard cases before you talk to anyone

This is the stage that determines everything downstream, and it happens before any vendor is contacted.

Write down, from your own operation:

  • Your most complex entity.
    The one with a joint venture partner, an unusual ownership split, or intercompany charges that take a day to reconcile.

  • Your strangest lease.
    The percentage rent arrangement, the side letter with a bespoke escalation, the ground lease with a formula agreed decades ago. Every portfolio has three or four of these and they are usually the valuable ones, which we covered as exception debt in the four consolidation debts.

  • Your hardest close.
    Whichever month end took longest last year, and specifically what made it long.

  • Your ugliest data.
    The property with duplicated vendor records, inconsistent unit naming or ten years of unmaintained history.

  • Your most awkward report.
    The one an owner, lender or auditor asked for that took a week to assemble.

  • The single most useful question to ask internally is where the spreadsheets are.
    Every workbook maintained alongside the official system is a requirement the current platform does not meet, and it is a more honest requirements document than anything a committee will produce in a meeting.

    Five to eight cases is enough. This becomes the spine of the evaluation, and every vendor answers the same set.

Stage two: written answers before any demo

Reverse the conventional order, because the conventional order favours the seller.

Send the same document to every shortlisted vendor at the same time, with a written deadline. Ask them to describe, in writing, how their platform handles each of your hard cases. Not whether it can. How.

Written responses are harder to route around than a live demonstration, where a presenter can navigate away from an awkward area without anyone noticing. They also create a record you can compare side by side and refer back to during implementation.

Three rules for reading the responses, drawn from enterprise procurement practice and worth applying strictly:

  • Treat "available in a future release" as not available, unless the date is contractually committed.

  • Treat "available via a third-party integration" as not available, unless the integration is pre-built and maintained by the vendor.

  • Require evidence for every compliance claim, rather than an assertion.

The second rule matters more in property software than in most categories, because the phrase "we integrate with that" can mean anything from a maintained native connector to a middleware project you will pay for.

Stage three: make the demo run on your data

Now hold the demos, and change what is being demonstrated.

Vendor demonstrations use curated datasets by design, and the industry is candid about why: sandbox environments let a prospect see the product under realistic but controlled conditions without production risk. That is reasonable from the seller's side. It is not sufficient from yours.

Provide a sample of your own data, anonymised if necessary, and ask each vendor to configure your hard cases in it. What you are testing is not whether the screen looks right. It is:

  • How long configuration takes.
    A capability requiring three weeks of professional services is a different product from one configured in an afternoon.

  • Who does the configuring.
    If only the vendor's implementation team can do it, every future change is a paid change request.

  • Whether the answer is native or assembled.
    Ask which module, add-on or partner product delivers each capability, and what each costs.

  • What breaks.
    Feed in the duplicated records and the inconsistent naming deliberately. Watching how a platform handles bad data tells you more than watching it handle clean data, and your data is not clean.

Red flags here are consistent across software categories: generic demos on sanitised test data rather than a proof of concept using your actual data and systems, and implementation timelines quoted without any data discovery or integration scoping. Both hide the same thing, which is the gap between the demo and your migration.

Stage four: reference the failures, not the successes

Vendors supply reference customers who are happy. That is expected, and it makes the standard reference call almost worthless.

Ask different questions. Enterprise evaluation practice suggests four that consistently produce useful answers:

  1. What did the integration actually require compared with what was scoped in the proposal?

  2. How long did production deployment take relative to the original timeline?

  3. What did change management support look like in practice?

  4. How does the system perform today versus what was projected at purchase?

References who answer all four with specifics are themselves the strongest signal. References who answer in generalities are telling you something too.

Two additions specific to property. Ask for a reference at your scale and in your asset mix, since a platform excellent at single-asset-class residential may be untested in your mixed portfolio. And ask a reference who has been live for at least two years, because the first year of any platform is flattering and the interesting problems surface in year two.

Stage five: price the exit before you enter

The stage almost nobody runs, and the one that most affects your position later.

Before signing, establish in writing: what data you can extract, in what format, at what cost, and how long you retain access after termination. Negotiate data portability rights, audit access and exit provisions before signing rather than after implementation, because your leverage is never higher than it is now.

This matters in property specifically because the structures hardest to extract are the ones you most need. Historical CAM and recovery calculations. Trust and deposit ledgers. Recurring charge schedules with future-dated escalations. A vendor who can export a tenant list has not answered the question.

Ask one question and listen carefully to the answer: if we leave in five years, what exactly do we take with us? The specificity of the response tells you how the vendor thinks about the relationship.

What to score, and what not to

Weight the scorecard before you see any responses, so the weighting is not retrofitted to the vendor you have started to like.

Weight higher

Weight lower

Performance against your hard cases

Feature count

Native versus assembled capability

Interface aesthetics

Configuration ownership after go-live

Demo polish

Reference specificity at your scale

Analyst rankings

Data extraction and exit terms

Roadmap promises

Implementation scope realism

Discount depth

The right column is not worthless. It is simply where evaluations tend to over-index, because those things are visible early and easy to compare, while the left column requires work.

One governance point sits above the scorecard. If no executive can settle a disagreement between departments about what a term means, the evaluation will stall regardless of which platform wins. Occupancy, turn time and recoverable expense all mean different things to leasing, accounting and asset management, and the definitions register that resolves them belongs in evaluation rather than in implementation.

Where evaluations typically go wrong

Symptom

What went wrong

Fix

Every vendor looked good

Demo-led evaluation

Written responses on your hard cases first

The shortlist changed after one impressive demo

Weighting set after seeing vendors

Weight the scorecard before responses arrive

Implementation cost doubled after signing

No data discovery in scoping

Require scoping against your actual data

A key capability turned out to need a third-party product

"We integrate with that" accepted at face value

Ask which module delivers it and what it costs

Finance raised objections in month four

Wrong people in the room

Include whoever owns the close from stage one

Nobody can agree what a report should show

Definitions never settled

Build the definitions register during evaluation

Switching later proved impossible

Exit terms never negotiated

Price the exit before signing

Frequently asked questions

Q1. How should you run a property management software evaluation?
Start by documenting your hardest cases: most complex entity, most unusual lease, longest close, worst data, most awkward report. Send the same written questions to every vendor before any demo, then require demonstrations against your data rather than theirs.

Q2. What questions should you ask property management software vendors?
Beyond capability questions, ask how each hard case would be configured and by whom, which module or partner product delivers each capability, what implementation scope was based on, and what data extraction looks like at exit.

Q3. Should you run a proof of concept?
Yes, but only with defined success criteria set in advance and a clear path to production. Gartner predicted that at least 30% of generative AI projects would be abandoned after proof of concept by the end of 2025, citing poor data quality, inadequate risk controls, escalating costs or unclear business value. The transferable lesson is to define success before the test begins rather than after it.

Q4. What are the red flags in a software evaluation?
Demos on sanitised data rather than yours, implementation timelines quoted without data discovery, capability delivered through unnamed third parties, reluctance to provide references at your scale, and vague answers on data extraction.

Q5. Who should be involved in choosing property management software?
Whoever owns the month-end close, whoever owns lease administration, whoever handles the hardest reports, and an executive who can settle definitional disputes between departments. Site operations should be represented but rarely surfaces the constraints that break implementations.

Q6. How long should an evaluation take?
Longer than a demo cycle and shorter than a year. The variable is not vendor count but how long it takes to assemble your hard cases honestly, which is internal work no vendor can accelerate.

Q7. What should you negotiate before signing?
Data portability rights, audit access, exit provisions and implementation scope tied to your actual data. Leverage is highest before signature and drops sharply afterwards.

The real job of a software evaluation

An evaluation is usually run as a search for the best product. That framing quietly hands control to the vendors, because it invites them to demonstrate strength rather than to be tested against difficulty.

The better framing is that an evaluation is a test you design. You supply the conditions, you supply the data, you supply the awkward cases, and you score every vendor against the same ones. What you are buying is not the demo. It is how the platform behaves in your hardest month, with your messiest entity, in front of the report you dread producing.

The companies making disciplined software decisions are not necessarily the ones watching the most demos. They are the ones doing the internal work of documenting their complexity before anyone shows them a screen, because that document becomes the evaluation.

RIOO is built on NetSuite. Bring us your hardest cases, and we'll show you how the platform handles them. See how the platform is structured