Weighted Procurement Scoring for Service Robots
At a glance: By the time proposals are on the table, every shortlisted vendor has a plausible story. The difference between a good fleet and an expensive mistake is rarely the hardware, it is the decision method. A weighted scoring matrix converts opinion into arithmetic: each criterion is scored on a fixed scale, multiplied by a weight agreed in advance, and totalled. The result is a ranked shortlist you can defend to a board, to an auditor and to the losing bidder.
Why Unweighted Evaluation Fails in Robot Procurement
Most buying committees evaluate robot vendors on a single axis: price, or price plus a gut feeling about the demo. That works when the product is a commodity. A service robot fleet is not a commodity. It is a multi-year operating commitment whose true cost is dominated by uptime, service response, integration effort and the replacement cycle, none of which are visible in the purchase price.
The failure mode is predictable. The vendor with the slickest demonstration and the lowest headline price wins, and eighteen months later the buyer is absorbing downtime, paying for integration the proposal implied was free, and discovering the battery pack is a proprietary part with a six-week lead time. An unweighted decision has no place to record those risks before they become costs.
A weighted matrix does not remove judgement. It forces judgement into the open, where each assumption has a number and each number can be challenged. The output is not a proof, but it is auditable, and auditability is what protects the decision when the fleet underperforms and someone asks how the vendor was chosen.
The Seven Criteria That Carry the Decision
Boil the evaluation down to seven criteria. More than that and the committee loses the thread; fewer and you omit a deal-breaker. The table below gives the standard set, the evidence to request for each, and a starting weight. Adjust the weights to your building type, but fix them before proposals are opened.
| Criterion | What it measures | Evidence to request | Weight |
|---|---|---|---|
| Total cost of ownership | 5-year cost including service, parts, power | Itemised TCO model, not a price sheet | 25% |
| Technical fit | Coverage, navigation, payload against your scenes | Measured coverage rate on comparable floors | 20% |
| Service and uptime | Response time, spares availability, first-time fix | SLA with credits, named service region | 20% |
| Integration capability | API, BMS/CMMS hooks, data export | Documented API, sample export file | 10% |
| Vendor stability | Financial health, install base, references | Two reference sites you can visit | 10% |
| Safety and compliance | Certification against the standards you must meet | Test certificates, not marketing claims | 10% |
| Contract terms | Data ownership, exit, escalators, remedies | Draft contract, not a proposal annex | 5% |
Notice where the weight sits. TCO, technical fit and service together carry 65%. Price alone is a component of TCO, not a criterion in its own right, which is precisely the re-framing that stops the cheapest bid winning by default. The commercial depth of that cost model is set out in the five-year cost of ownership method, and the arithmetic behind the technical-fit scores is covered in the throughput planning guide.
Scoring Each Criterion Without Fooling Yourself
Use a fixed 1-to-5 scale for every criterion and define what each point means in writing. The scale is the guard against grade inflation, where every criterion drifts to 4 or 5 and the matrix stops discriminating.
| Score | Meaning | Use when |
|---|---|---|
| 5 | Exceeds requirement, documented | Evidence proves it beats your spec |
| 4 | Meets requirement fully | Evidence proves the spec is met |
| 3 | Meets with caveats | Yes, but with a stated condition |
| 2 | Partially meets | Gap you must absorb or pay to close |
| 1 | Does not meet | Unresolved or unproven |
Score on evidence, not on the response document. A vendor who claims a 98% uptime figure but supplies no measurement method scores a 3, not a 5. A vendor who supplies the last twelve months of fleet telemetry scores a 5 even if the number is 96%, because it is verifiable. The rule is simple: no evidence, no top score. Apply the procurement-adjacent due diligence in the vendor evaluation framework as the evidence trail that feeds these scores.
Normalising and Applying the Weights
Once every criterion is scored 1-5, multiply by the weight and total. The arithmetic is deliberately simple so a committee member can reproduce it by hand.
| Criterion | Weight | Vendor A score | A weighted | Vendor B score | B weighted |
|---|---|---|---|---|---|
| Total cost of ownership | 25% | 5 | 1.25 | 3 | 0.75 |
| Technical fit | 20% | 4 | 0.80 | 5 | 1.00 |
| Service and uptime | 20% | 4 | 0.80 | 3 | 0.60 |
| Integration capability | 10% | 3 | 0.30 | 5 | 0.50 |
| Vendor stability | 10% | 5 | 0.50 | 3 | 0.30 |
| Safety and compliance | 10% | 5 | 0.50 | 4 | 0.40 |
| Contract terms | 5% | 4 | 0.20 | 2 | 0.10 |
| Total | 100% | 4.35 | 3.65 |
Vendor B wins the demo on technical fit and integration, yet loses the decision on cost and service. That inversion is exactly what the matrix exists to surface. Had the committee scored on demo impressions alone, B would have won, and the 5-year cost gap would have arrived as a surprise.
Handling Disagreement and Grade Inflation
Two safeguards keep the scores honest. First, have each evaluator score independently before the committee meets, then reconcile. The reconciliation discussion is where hidden assumptions surface: one reviewer scores service a 3 because a 4-hour response is a caveat, another scores it a 4 because it is contractual. Neither is wrong until the group defines the standard.
Second, run a sensitivity check. Raise and lower each weight by five points and see whether the ranking changes. If a two-point shift in the TCO weight flips the winner, the decision is genuinely close and the committee should look for another differentiator, such as a pilot. If the ranking holds across the range, the decision is robust and you can stop debating weights. The go/no-go gate that a pilot adds to this arithmetic is set out in the site survey and readiness method.
Tie-Breakers and the Award Decision
A matrix can produce a tie or a near-tie. Resolve it in a fixed order, written down before scoring, so the tie-break is not a backdoor for the preferred bidder.
- Highest TCO score wins the tie, because cost certainty compounds over the whole contract.
- If still level, highest service score wins, because uptime is the operational risk you cannot recover.
- If still level, the vendor who will accept a paid pilot with defined success gates wins.
- If still level, defer to the reference site visit result, not the demo.
Record the final matrix, the evidence behind each score and the tie-break used. That record is what makes the award defensible in a procurement audit and what lets you hold the winning vendor to the claims that earned them the score. For a live fleet, the same scoring discipline reappears at renewal, where the incumbent is scored against challengers using identical criteria, a process that keeps pricing honest without the disruption of a full re-tender.
Frequently Asked Questions
How many vendors should be scored? Three to five. Two gives no spread and becomes a binary argument; more than five and the evidence-gathering cost exceeds the value of the extra comparison.
Should price be its own criterion? No. Price is an input to TCO. Scored separately it double-counts and drifts the decision back toward the cheapest bid, which is the outcome the matrix exists to prevent.
Who should score? At least three people with different exposure: operations, who live with uptime; finance, who own the TCO; and IT or facilities, who own integration. Single-evaluator scoring defeats the purpose. The scoring discipline here also underpins the lease versus buy decision, where the same weighted logic decides the financing route rather than the vendor.
