Skip to content
       

Blog

Vendor Scorecard for Property Management: How to Measure Contractor Performance

Vendor Scorecard for Property Management: How to Measure Contractor Performance

A vendor scorecard measures contractor performance using operational data rather than opinion. For property teams, that usually means five metrics pulled from work order history: first-time fix rate, response and completion against target, callback rate, cost variance against quote, and documentation completeness. Scored consistently within trade cohorts and run on a fixed cycle, it lets you compare maintenance vendors fairly, spot decline early, and hold a difficult conversation from evidence rather than from memory.

Ask a property manager which of their contractors is best and you will get an answer immediately. Ask how they know, and the answer changes shape. He always picks up the phone. She sorts things out. They have been with us for years.

None of that is measurement. It is familiarity, and familiarity is a poor proxy for performance. The contractor who answers the phone quickly may also be the one returning to the same unit three times a month. The one nobody likes talking to may have the best completion record in the portfolio.

A vendor scorecard replaces impression with evidence. Most of the data needed already sits in your work order history.

Key takeaways

  • Most vendor reviews measure responsiveness and rapport, neither of which predicts outcome.

  • Five metrics do most of the work: first-time fix, response against target, callback rate, cost variance, and documentation completeness.

  • Definitions matter more than the metrics. An unclear definition of first-time fix produces a number nobody can act on.

  • Comparing an elevator company against a plumber requires normalising by trade and job type, not scoring everyone on one scale.

  • The scorecard's real value is that it converts a difficult conversation into a documented one.

In this guide

  • What is a vendor scorecard?

  • Who should use a vendor scorecard?

  • Why most vendor reviews are worthless

  • The five vendor performance metrics that matter most

  • Vendor scorecard example

  • Get the definitions right before you measure anything

  • The comparison problem: trades, complexity, and property mix

  • Why total spend is the wrong lens on cost

  • Building the scorecard: weighting, bands, and cadence

  • What to do with the result

  • The evidence you need before terminating a contractor

  • Common vendor evaluation mistakes

  • Frequently asked questions

What is a vendor scorecard?

Short answer: A vendor scorecard is a structured, repeated assessment of a contractor's performance against defined measures, calculated from operational data rather than from opinion. In property management it typically scores response and completion against agreed targets, quality of work as measured by repeat visits, cost behavior against quote, and administrative reliability. It is run on a fixed cycle so performance can be compared over time and across the vendor base.

The scorecard is not a satisfaction survey. Satisfaction tells you how a contractor made your team feel. The scorecard tells you what the contractor delivered.

Who should use a vendor scorecard?

Any property team that engages more than a handful of contractors and cannot say, with evidence, which of them performs best. That usually means:

  • Residential and multifamily portfolio managers

  • Commercial and office property teams

  • HOA and community association managers

  • Facility management teams running multi-site contracts

  • Build-to-rent and single-family rental operators

  • Student housing and institutional operators

The threshold is not portfolio size. It is whether the same work is being done by more than one vendor, because that is the point at which comparison becomes possible and useful.

Why most vendor reviews are worthless

Three failure patterns account for nearly all bad vendor performance evaluations.

They measure the wrong thing. Responsiveness gets scored because it is visible. A contractor who answers immediately and then takes eleven days to complete scores well on the only thing anyone recorded.

They are retrospective and unstructured. The review happens at renewal, from memory, six or twelve months after the events being judged. Recent incidents dominate. A good year with a bad December scores as a bad year.

They have no denominator. "We had three complaints about them" means nothing without knowing they completed 400 jobs. The maintenance vendor with one complaint out of twelve jobs is the problem, and nobody notices.

The fix is not more scrutiny. It is fewer, better-defined measures applied consistently.

The five vendor performance metrics that matter most

Metric

How to calculate

What it tells you

First-time fix rate

Jobs resolved on the first visit ÷ total jobs

Whether the contractor arrives prepared and competent

Response and completion against target

Jobs meeting the agreed window ÷ total jobs, split by priority

Whether the service level is real or aspirational

Callback rate

Return visits for the same fault within 30 days ÷ total jobs

Work quality, independent of how the first visit felt

Cost variance against quote

(Invoiced ÷ quoted) − 1, averaged across jobs

Estimating discipline and scope-creep behavior

Documentation completeness at close

Jobs closed with required evidence ÷ total jobs

Whether you can defend the charge or the claim later

First-time fix rate is the percentage of jobs resolved on the first visit, with no return trip for parts, further diagnosis, or rework. Field service benchmarking puts the industry average at around 80%, with performance near 90% considered strong and anything under 70% treated as a warning sign. At 60%, four out of every ten jobs need a second visit, which means double the labour, double the travel, and a second disruption to the occupant.

Callback rate is the percentage of maintenance jobs requiring another visit for the same fault within a defined period, usually 30 days. Many property teams set a callback rate below 5% as a performance target, though acceptable thresholds vary by trade and by contract. It is harder to game than first-time fix, because a repeat visit for the same fault is unambiguous. A vendor with a high first-time fix rate and a high callback rate is closing jobs that were not finished.

Cost variance against quote is the average difference between quoted and invoiced amounts across completed jobs, expressed as a percentage. A maintenance vendor whose invoices consistently land 15% above quote is not more expensive per job than the competition. They are less predictable, which is worse, because every budget built on their quotes is wrong.

Two further metrics are worth adding once the first five are stable: appointments kept, which matters enormously in residential where an occupant took time off work, and mean time to complete by priority, which distinguishes a contractor who is fast at everything from one who is fast only at the easy jobs.

Vendor scorecard example

Here is what a completed quarterly scorecard looks like for three plumbing contractors working the same residential portfolio. All three are scored within the same trade cohort, against the same targets, over the same period.

Metric

Target

Vendor A

Vendor B

Vendor C

Jobs completed

Min 20

84

61

19

First-time fix rate

85%

88%

91%

74%

Response within target, P1

95%

96%

82%

100%

Completion within target, P2

90%

91%

88%

79%

Callback rate, 30 days

Under 5%

3.6%

9.8%

5.3%

Cost variance against quote

Under 5%

2.1%

4.4%

18.7%

Documentation complete at close

95%

97%

71%

88%

Weighted score

 

91

74

Insufficient volume

Band

 

Performing

Improvement required

Not scored

Three things this example shows that a ranked list would hide.

Vendor B looks good on the metric everyone watches. The highest first-time fix rate in the cohort, and the worst callback rate by a wide margin. Those two together mean jobs are being closed that were not finished. First-time fix on its own would have made B the top performer.

Vendor A is the strongest and it is not close. Highest volume, consistent across every measure, no single number spectacular. That profile is what reliability looks like on a scorecard.

Vendor C is not scored at all. Nineteen jobs is below the minimum volume threshold, so the numbers are recorded but not published as a result. The 18.7% cost variance is worth a conversation now, but it is not yet a pattern.

The action list from this quarter is short: share A's score, open an improvement plan with B on callback rate and documentation, and ask C about quoting before their volume grows.

Get the definitions right before you measure anything

This is the step that determines whether the scorecard is useful or decorative, and it is the one most teams skip.

The pattern is consistent enough to predict. A team agrees to start tracking first-time fix, runs it for a quarter, and then discovers at the review that the contractor counts a job as fixed when they leave the site, while the property team counts it as fixed when the tenant stops reporting the fault. Both were measuring in good faith. Neither number means anything, and the quarter is wasted.

Before the first scorecard runs, agree in writing:

  • What counts as a first-time fix. Does a job requiring a part order count as a failure? Most definitions say yes, because parts availability is the contractor's responsibility. Does a job where access was refused count? Usually no, because that is outside their control.

  • When the clock starts. At tenant report, at dispatch, or at contractor acknowledgement? These differ by hours and change the result substantially.

  • What counts as a callback. Same fault, same location, within a set window. Thirty days is common. A different fault in the same unit is not a callback.

  • What "complete" means. Work finished, or work finished with documentation submitted? If it is the latter, say so, because it changes behavior.

  • Which jobs are excluded. Emergencies, jobs where access failed, jobs cancelled by the property team. Exclusions must be defined in advance or they become an argument later.

Write these into the contract or the service agreement. A metric the contractor disputes at review is a metric you cannot act on.

For a view of the intake side that generates all of this data, see RIOO's guide to managing maintenance requests.

The comparison problem: trades, complexity, and property mix

A single leaderboard ranking every maintenance vendor in the portfolio is worse than no scorecard, because it produces confident conclusions from incomparable data.

Three sources of distortion:

Trade. An elevator contractor works on a statutory schedule with long lead times on parts. A plumber attending a leak works in hours. Cleaning, security, and other soft services have different rhythms again. Scoring all of them against a four-hour response target penalises some for the nature of their work.

Job complexity. A contractor handling only routine tasks will out-score one handling the difficult ones. If the difficult jobs are routed to your best contractor, the scorecard will show them as your worst.

Property mix. A contractor working an older building with legacy systems will show a higher callback rate than one working new stock, for reasons that have nothing to do with competence.

Distortion

Fix

Trade differences

Score within trade cohorts, never across them

Job complexity

Set targets by job category, not one target per contractor

Property mix

Compare a contractor to its own trailing performance as well as to peers

Low job volume

Set a minimum job count before a score is published

Emergency work

Report emergency and planned work as separate lines

The last one matters more than it looks. A contractor who takes every emergency callout at 2am and performs adequately is worth more than one who declines them and performs excellently on scheduled work. The scorecard should be able to show that.

Why total spend is the wrong lens on cost

Almost every portfolio ranks vendors by annual spend. It is the easiest number to produce and the least informative.

Total spend reflects how much work you gave a contractor, not what that work cost you. The right unit is cost per job type, compared across the vendors doing the same category of work.

Question

What total spend tells you

What cost per job type tells you

Is this contractor expensive?

Nothing, without job volume

Directly comparable to peers

Where is the saving?

Only the largest vendor

The specific job categories with rate variance

Are quotes reliable?

Nothing

Variance between quoted and invoiced

Should we consolidate?

Suggests consolidating to the cheapest

Shows which categories are worth consolidating

Add the hidden cost that never appears on an invoice. A callback consumes labour hours, travel, and scheduling capacity with no additional revenue on a fixed-price job. A contractor with a 12% callback rate is charging you a premium that does not appear anywhere in their pricing.

Building the scorecard: weighting, bands, and cadence

Weighting. Not every metric matters equally in maintenance vendor management, and the weighting should reflect what you actually need from that trade. A workable starting point for reactive maintenance:

Metric

Suggested weight

First-time fix rate

25%

Response and completion against target

25%

Callback rate

20%

Cost variance against quote

20%

Documentation completeness

10%

Adjust by category. For statutory and compliance work, documentation completeness should carry far more weight, because the evidence is the deliverable.

Bands. Convert scores into three or four bands rather than a continuous ranking. Something like: performing, watch, improvement required, and at risk. Bands make the result actionable. A ranked list of eleven contractors mostly tells you who is eleventh.

Cadence. Quarterly is the practical rhythm for most portfolios. Monthly produces noise on small job volumes. Annual is too late to change anything. A common pattern in contracted repairs work is to report quarterly against annual targets, which balances the two: measure often, judge on the trend. 

Portfolio-level dashboards and reports are where the scorecard becomes something the team looks at rather than a document produced once and filed.

What to do with the result

The scorecard is only worth building if the contractor performance review changes something.

Performing. Share the score. Contractors who see their numbers usually work to protect them, and the ones who ask how the score is calculated are the ones worth keeping.

Watch. One metric below target, trend unclear. Raise it at the next review meeting, note it, and look again next quarter.

Improvement required. Specific, written, time-bound. Name the metric, the current level, the target, and the review date. Most contractors respond to this, particularly when they can see they are being measured on the same basis as everyone else.

At risk. The conversation about whether the relationship continues. This is where the scorecard earns its existence, because the discussion is about a documented pattern rather than about a memorable bad job.

Reviews should be a conversation, not a verdict. A contractor scoring badly on response time because your team is dispatching them without unit access details has a legitimate defence, and you want to hear it. A meaningful share of the poor scores any portfolio surfaces at first run are caused by something on the property team's side.

The evidence you need before terminating a contractor

If you may need to end a relationship, the scorecard is the record that makes it defensible.

  • A consistent measurement basis applied to all vendors in the same cohort, not a case assembled after the decision

  • Multiple periods showing a pattern rather than one bad quarter

  • Documented notification that performance was below target, with dates

  • A recorded improvement opportunity with specified targets and a review date

  • The underlying job records supporting each figure

  • Contract terms on service levels, notice, and termination, confirmed with legal advice before acting

Termination terms and any statutory or contractual obligations vary by jurisdiction and by the agreement in place, so treat this as an operational checklist and take advice on the specific contract.

Common vendor evaluation mistakes

Mistake

Why it distorts

Scoring across trades on one scale

Penalises trades with inherently longer cycles

Publishing scores on tiny job volumes

Two bad jobs out of five is not a trend

Counting access failures against the contractor

Rewards contractors who avoid hard-to-access properties

Measuring response but not completion

A contractor who arrives fast and finishes slowly scores well

Changing definitions mid-year

Destroys the trend, which is the most useful part

Never sharing the score

The scorecard changes nothing if the contractor cannot see it

Weighting cost above everything

Produces a vendor base that is cheap per visit and expensive per year

Frequently asked questions

1. What is a vendor scorecard?
A structured, repeated assessment of a contractor's performance against defined measures, calculated from operational data rather than opinion. In property management it typically covers response and completion against target, first-time fix, callback rate, cost variance against quote, and documentation completeness.

2. How do you measure contractor performance?
By defining a small set of metrics, agreeing exactly how each is calculated, applying them consistently within trade cohorts, and running the assessment on a fixed cycle so trends are visible. The underlying data usually already exists in work order history.

3. What KPIs should you track for maintenance contractors?
First-time fix rate, response and completion against agreed targets by priority, callback rate within a defined window, cost variance against quote, and documentation completeness at close. Appointments kept is a valuable addition in residential portfolios.

5. What should a vendor scorecard template include?
Vendor name and trade, the review period, job volume, each metric with its target and actual result, a weighted total, a performance band, and a notes field for context. Keep it to one page per vendor. A template nobody can read at a review meeting does not get used at review meetings.

6. What is a good first-time fix rate?
Field service benchmarking places the industry average at around 80%. Performance near 90% is considered strong and below 70% is generally treated as a warning sign, though the appropriate target varies by trade and job type.

7. What is a callback rate and what is an acceptable level?
The proportion of jobs requiring a return visit for the same fault within a defined window, commonly 30 days. Many property teams set a target below 5%, though thresholds vary by trade and contract. It is a harder metric to game than first-time fix, because a repeat visit for the same fault is unambiguous.

8. What is the difference between vendor evaluation and vendor management?
Vendor management covers the full relationship: sourcing, onboarding, compliance documentation, contracting, and payment. Vendor evaluation is the measurement part, assessing how a contractor performed against agreed standards once work is underway. A scorecard is an evaluation tool, not a management system.

9. How do you compare contractors working in different trades?
You do not compare them directly. Score within trade cohorts, set targets by job category rather than one target per contractor, and compare each contractor against its own trailing performance as well as against peers in the same trade.

10. How often should you review vendor performance?
Quarterly suits most portfolios. Monthly generates noise on low job volumes, and annual reviews arrive too late to change the year being reviewed. Measure frequently and judge on the trend.

11. What evidence do you need before terminating a contractor?
A consistent measurement basis applied across the cohort, multiple periods showing a pattern, documented notification that performance was below target, a recorded improvement opportunity with targets and a review date, and the underlying work order records. Confirm the contractual and legal position before acting.

12. How do you start if you have never scored vendors before?
Pick two metrics, define them precisely in writing, run them for one quarter, and share the result with the vendors. Two well-defined metrics beat eight loosely defined ones, and the definitions are the part that takes time to get right.

Maintenance contractors are the largest controllable operating cost in most portfolios, and the one most often managed by relationship rather than by evidence. That is not because property teams are careless. It is because the data sits in work order records that nobody has been asked to read as a performance signal.

The scorecard changes what the annual conversation is about. Not whether the contractor is any good, which is unanswerable, but whether the numbers moved, and if not, what happens next.