A vendor scorecard measures contractor performance using operational data rather than opinion. For property teams, that usually means five metrics pulled from work order history: first-time fix rate, response and completion against target, callback rate, cost variance against quote, and documentation completeness. Scored consistently within trade cohorts and run on a fixed cycle, it lets you compare maintenance vendors fairly, spot decline early, and hold a difficult conversation from evidence rather than from memory.
Ask a property manager which of their contractors is best and you will get an answer immediately. Ask how they know, and the answer changes shape. He always picks up the phone. She sorts things out. They have been with us for years.
None of that is measurement. It is familiarity, and familiarity is a poor proxy for performance. The contractor who answers the phone quickly may also be the one returning to the same unit three times a month. The one nobody likes talking to may have the best completion record in the portfolio.
A vendor scorecard replaces impression with evidence. Most of the data needed already sits in your work order history.
Key takeaways
-
Most vendor reviews measure responsiveness and rapport, neither of which predicts outcome.
-
Five metrics do most of the work: first-time fix, response against target, callback rate, cost variance, and documentation completeness.
-
Definitions matter more than the metrics. An unclear definition of first-time fix produces a number nobody can act on.
-
Comparing an elevator company against a plumber requires normalising by trade and job type, not scoring everyone on one scale.
-
The scorecard's real value is that it converts a difficult conversation into a documented one.
In this guide
-
What is a vendor scorecard?
-
Who should use a vendor scorecard?
-
Why most vendor reviews are worthless
-
The five vendor performance metrics that matter most
-
Vendor scorecard example
-
Get the definitions right before you measure anything
-
The comparison problem: trades, complexity, and property mix
-
Why total spend is the wrong lens on cost
-
Building the scorecard: weighting, bands, and cadence
-
What to do with the result
-
The evidence you need before terminating a contractor
-
Common vendor evaluation mistakes
-
Frequently asked questions
What is a vendor scorecard?
Short answer: A vendor scorecard is a structured, repeated assessment of a contractor's performance against defined measures, calculated from operational data rather than from opinion. In property management it typically scores response and completion against agreed targets, quality of work as measured by repeat visits, cost behavior against quote, and administrative reliability. It is run on a fixed cycle so performance can be compared over time and across the vendor base.
The scorecard is not a satisfaction survey. Satisfaction tells you how a contractor made your team feel. The scorecard tells you what the contractor delivered.
Who should use a vendor scorecard?
Any property team that engages more than a handful of contractors and cannot say, with evidence, which of them performs best. That usually means:
-
Residential and multifamily portfolio managers
-
Commercial and office property teams
-
HOA and community association managers
-
Facility management teams running multi-site contracts
-
Build-to-rent and single-family rental operators
-
Student housing and institutional operators
The threshold is not portfolio size. It is whether the same work is being done by more than one vendor, because that is the point at which comparison becomes possible and useful.
Why most vendor reviews are worthless
Three failure patterns account for nearly all bad vendor performance evaluations.
They measure the wrong thing. Responsiveness gets scored because it is visible. A contractor who answers immediately and then takes eleven days to complete scores well on the only thing anyone recorded.
They are retrospective and unstructured. The review happens at renewal, from memory, six or twelve months after the events being judged. Recent incidents dominate. A good year with a bad December scores as a bad year.
They have no denominator. "We had three complaints about them" means nothing without knowing they completed 400 jobs. The maintenance vendor with one complaint out of twelve jobs is the problem, and nobody notices.
The fix is not more scrutiny. It is fewer, better-defined measures applied consistently.
The five vendor performance metrics that matter most
|
Metric |
How to calculate |
What it tells you |
|---|---|---|
|
First-time fix rate |
Jobs resolved on the first visit ÷ total jobs |
Whether the contractor arrives prepared and competent |
|
Response and completion against target |
Jobs meeting the agreed window ÷ total jobs, split by priority |
Whether the service level is real or aspirational |
|
Callback rate |
Return visits for the same fault within 30 days ÷ total jobs |
Work quality, independent of how the first visit felt |
|
Cost variance against quote |
(Invoiced ÷ quoted) − 1, averaged across jobs |
Estimating discipline and scope-creep behavior |
|
Documentation completeness at close |
Jobs closed with required evidence ÷ total jobs |
Whether you can defend the charge or the claim later |
First-time fix rate is the percentage of jobs resolved on the first visit, with no return trip for parts, further diagnosis, or rework. Field service benchmarking puts the industry average at around 80%, with performance near 90% considered strong and anything under 70% treated as a warning sign. At 60%, four out of every ten jobs need a second visit, which means double the labour, double the travel, and a second disruption to the occupant.
Callback rate is the percentage of maintenance jobs requiring another visit for the same fault within a defined period, usually 30 days. Many property teams set a callback rate below 5% as a performance target, though acceptable thresholds vary by trade and by contract. It is harder to game than first-time fix, because a repeat visit for the same fault is unambiguous. A vendor with a high first-time fix rate and a high callback rate is closing jobs that were not finished.
Cost variance against quote is the average difference between quoted and invoiced amounts across completed jobs, expressed as a percentage. A maintenance vendor whose invoices consistently land 15% above quote is not more expensive per job than the competition. They are less predictable, which is worse, because every budget built on their quotes is wrong.
Two further metrics are worth adding once the first five are stable: appointments kept, which matters enormously in residential where an occupant took time off work, and mean time to complete by priority, which distinguishes a contractor who is fast at everything from one who is fast only at the easy jobs.
Vendor scorecard example
Here is what a completed quarterly scorecard looks like for three plumbing contractors working the same residential portfolio. All three are scored within the same trade cohort, against the same targets, over the same period.
|
Metric |
Target |
Vendor A |
Vendor B |
Vendor C |
|---|---|---|---|---|
|
Jobs completed |
Min 20 |
84 |
61 |
19 |
|
First-time fix rate |
85% |
88% |
91% |
74% |
|
Response within target, P1 |
95% |
96% |
82% |
100% |
|
Completion within target, P2 |
90% |
91% |
88% |
79% |
|
Callback rate, 30 days |
Under 5% |
3.6% |
9.8% |
5.3% |
|
Cost variance against quote |
Under 5% |
2.1% |
4.4% |
18.7% |
|
Documentation complete at close |
95% |
97% |
71% |
88% |
|
Weighted score |
91 |
74 |
Insufficient volume |
|
|
Band |
Performing |
Improvement required |
Not scored |
Three things this example shows that a ranked list would hide.
Vendor B looks good on the metric everyone watches. The highest first-time fix rate in the cohort, and the worst callback rate by a wide margin. Those two together mean jobs are being closed that were not finished. First-time fix on its own would have made B the top performer.
Vendor A is the strongest and it is not close. Highest volume, consistent across every measure, no single number spectacular. That profile is what reliability looks like on a scorecard.
Vendor C is not scored at all. Nineteen jobs is below the minimum volume threshold, so the numbers are recorded but not published as a result. The 18.7% cost variance is worth a conversation now, but it is not yet a pattern.
The action list from this quarter is short: share A's score, open an improvement plan with B on callback rate and documentation, and ask C about quoting before their volume grows.
Get the definitions right before you measure anything
This is the step that determines whether the scorecard is useful or decorative, and it is the one most teams skip.
The pattern is consistent enough to predict. A team agrees to start tracking first-time fix, runs it for a quarter, and then discovers at the review that the contractor counts a job as fixed when they leave the site, while the property team counts it as fixed when the tenant stops reporting the fault. Both were measuring in good faith. Neither number means anything, and the quarter is wasted.
Before the first scorecard runs, agree in writing:
-
What counts as a first-time fix. Does a job requiring a part order count as a failure? Most definitions say yes, because parts availability is the contractor's responsibility. Does a job where access was refused count? Usually no, because that is outside their control.
-
When the clock starts. At tenant report, at dispatch, or at contractor acknowledgement? These differ by hours and change the result substantially.
-
What counts as a callback. Same fault, same location, within a set window. Thirty days is common. A different fault in the same unit is not a callback.
-
What "complete" means. Work finished, or work finished with documentation submitted? If it is the latter, say so, because it changes behavior.
-
Which jobs are excluded. Emergencies, jobs where access failed, jobs cancelled by the property team. Exclusions must be defined in advance or they become an argument later.
Write these into the contract or the service agreement. A metric the contractor disputes at review is a metric you cannot act on.
For a view of the intake side that generates all of this data, see RIOO's guide to managing maintenance requests.
The comparison problem: trades, complexity, and property mix
A single leaderboard ranking every maintenance vendor in the portfolio is worse than no scorecard, because it produces confident conclusions from incomparable data.
Three sources of distortion:
Trade. An elevator contractor works on a statutory schedule with long lead times on parts. A plumber attending a leak works in hours. Cleaning, security, and other soft services have different rhythms again. Scoring all of them against a four-hour response target penalises some for the nature of their work.
Job complexity. A contractor handling only routine tasks will out-score one handling the difficult ones. If the difficult jobs are routed to your best contractor, the scorecard will show them as your worst.
Property mix. A contractor working an older building with legacy systems will show a higher callback rate than one working new stock, for reasons that have nothing to do with competence.
|
Distortion |
Fix |
|---|---|
|
Trade differences |
Score within trade cohorts, never across them |
|
Job complexity |
Set targets by job category, not one target per contractor |
|
Property mix |
Compare a contractor to its own trailing performance as well as to peers |
|
Low job volume |
Set a minimum job count before a score is published |
|
Emergency work |
Report emergency and planned work as separate lines |
The last one matters more than it looks. A contractor who takes every emergency callout at 2am and performs adequately is worth more than one who declines them and performs excellently on scheduled work. The scorecard should be able to show that.
Why total spend is the wrong lens on cost
Almost every portfolio ranks vendors by annual spend. It is the easiest number to produce and the least informative.
Total spend reflects how much work you gave a contractor, not what that work cost you. The right unit is cost per job type, compared across the vendors doing the same category of work.
|
Question |
What total spend tells you |
What cost per job type tells you |
|---|---|---|
|
Is this contractor expensive? |
Nothing, without job volume |
Directly comparable to peers |
|
Where is the saving? |
Only the largest vendor |
The specific job categories with rate variance |
|
Are quotes reliable? |
Nothing |
Variance between quoted and invoiced |
|
Should we consolidate? |
Suggests consolidating to the cheapest |
Shows which categories are worth consolidating |
Add the hidden cost that never appears on an invoice. A callback consumes labour hours, travel, and scheduling capacity with no additional revenue on a fixed-price job. A contractor with a 12% callback rate is charging you a premium that does not appear anywhere in their pricing.
Building the scorecard: weighting, bands, and cadence
Weighting. Not every metric matters equally in maintenance vendor management, and the weighting should reflect what you actually need from that trade. A workable starting point for reactive maintenance:
|
Metric |
Suggested weight |
|---|---|
|
First-time fix rate |
25% |
|
Response and completion against target |
25% |
|
Callback rate |
20% |
|
Cost variance against quote |
20% |
|
Documentation completeness |
10% |
Adjust by category. For statutory and compliance work, documentation completeness should carry far more weight, because the evidence is the deliverable.
Bands. Convert scores into three or four bands rather than a continuous ranking. Something like: performing, watch, improvement required, and at risk. Bands make the result actionable. A ranked list of eleven contractors mostly tells you who is eleventh.
Cadence. Quarterly is the practical rhythm for most portfolios. Monthly produces noise on small job volumes. Annual is too late to change anything. A common pattern in contracted repairs work is to report quarterly against annual targets, which balances the two: measure often, judge on the trend.
Portfolio-level dashboards and reports are where the scorecard becomes something the team looks at rather than a document produced once and filed.
What to do with the result
The scorecard is only worth building if the contractor performance review changes something.
Performing. Share the score. Contractors who see their numbers usually work to protect them, and the ones who ask how the score is calculated are the ones worth keeping.
Watch. One metric below target, trend unclear. Raise it at the next review meeting, note it, and look again next quarter.
Improvement required. Specific, written, time-bound. Name the metric, the current level, the target, and the review date. Most contractors respond to this, particularly when they can see they are being measured on the same basis as everyone else.
At risk. The conversation about whether the relationship continues. This is where the scorecard earns its existence, because the discussion is about a documented pattern rather than about a memorable bad job.
Reviews should be a conversation, not a verdict. A contractor scoring badly on response time because your team is dispatching them without unit access details has a legitimate defence, and you want to hear it. A meaningful share of the poor scores any portfolio surfaces at first run are caused by something on the property team's side.
The evidence you need before terminating a contractor
If you may need to end a relationship, the scorecard is the record that makes it defensible.
-
A consistent measurement basis applied to all vendors in the same cohort, not a case assembled after the decision
-
Multiple periods showing a pattern rather than one bad quarter
-
Documented notification that performance was below target, with dates
-
A recorded improvement opportunity with specified targets and a review date
-
The underlying job records supporting each figure
-
Contract terms on service levels, notice, and termination, confirmed with legal advice before acting
Termination terms and any statutory or contractual obligations vary by jurisdiction and by the agreement in place, so treat this as an operational checklist and take advice on the specific contract.
Common vendor evaluation mistakes
|
Mistake |
Why it distorts |
|---|---|
|
Scoring across trades on one scale |
Penalises trades with inherently longer cycles |
|
Publishing scores on tiny job volumes |
Two bad jobs out of five is not a trend |
|
Counting access failures against the contractor |
Rewards contractors who avoid hard-to-access properties |
|
Measuring response but not completion |
A contractor who arrives fast and finishes slowly scores well |
|
Changing definitions mid-year |
Destroys the trend, which is the most useful part |
|
Never sharing the score |
The scorecard changes nothing if the contractor cannot see it |
|
Weighting cost above everything |
Produces a vendor base that is cheap per visit and expensive per year |
Frequently asked questions
1. What is a vendor scorecard?
A structured, repeated assessment of a contractor's performance against defined measures, calculated from operational data rather than opinion. In property management it typically covers response and completion against target, first-time fix, callback rate, cost variance against quote, and documentation completeness.
2. How do you measure contractor performance?
By defining a small set of metrics, agreeing exactly how each is calculated, applying them consistently within trade cohorts, and running the assessment on a fixed cycle so trends are visible. The underlying data usually already exists in work order history.
3. What KPIs should you track for maintenance contractors?
First-time fix rate, response and completion against agreed targets by priority, callback rate within a defined window, cost variance against quote, and documentation completeness at close. Appointments kept is a valuable addition in residential portfolios.
5. What should a vendor scorecard template include?
Vendor name and trade, the review period, job volume, each metric with its target and actual result, a weighted total, a performance band, and a notes field for context. Keep it to one page per vendor. A template nobody can read at a review meeting does not get used at review meetings.
6. What is a good first-time fix rate?
Field service benchmarking places the industry average at around 80%. Performance near 90% is considered strong and below 70% is generally treated as a warning sign, though the appropriate target varies by trade and job type.
7. What is a callback rate and what is an acceptable level?
The proportion of jobs requiring a return visit for the same fault within a defined window, commonly 30 days. Many property teams set a target below 5%, though thresholds vary by trade and contract. It is a harder metric to game than first-time fix, because a repeat visit for the same fault is unambiguous.
8. What is the difference between vendor evaluation and vendor management?
Vendor management covers the full relationship: sourcing, onboarding, compliance documentation, contracting, and payment. Vendor evaluation is the measurement part, assessing how a contractor performed against agreed standards once work is underway. A scorecard is an evaluation tool, not a management system.
9. How do you compare contractors working in different trades?
You do not compare them directly. Score within trade cohorts, set targets by job category rather than one target per contractor, and compare each contractor against its own trailing performance as well as against peers in the same trade.
10. How often should you review vendor performance?
Quarterly suits most portfolios. Monthly generates noise on low job volumes, and annual reviews arrive too late to change the year being reviewed. Measure frequently and judge on the trend.
11. What evidence do you need before terminating a contractor?
A consistent measurement basis applied across the cohort, multiple periods showing a pattern, documented notification that performance was below target, a recorded improvement opportunity with targets and a review date, and the underlying work order records. Confirm the contractual and legal position before acting.
12. How do you start if you have never scored vendors before?
Pick two metrics, define them precisely in writing, run them for one quarter, and share the result with the vendors. Two well-defined metrics beat eight loosely defined ones, and the definitions are the part that takes time to get right.
Maintenance contractors are the largest controllable operating cost in most portfolios, and the one most often managed by relationship rather than by evidence. That is not because property teams are careless. It is because the data sits in work order records that nobody has been asked to read as a performance signal.
The scorecard changes what the annual conversation is about. Not whether the contractor is any good, which is unanswerable, but whether the numbers moved, and if not, what happens next.