Structuring a Service Robot Pilot That De-Risks the Fleet
At a glance: A pilot is the cheapest insurance in robot procurement and the most frequently botched. Run as a free demo, it produces a polite report and no decision. Run as a bounded experiment with a written hypothesis, a defined duration, measured KPIs and thresholds agreed before the robot arrives, it produces the one thing procurement cannot otherwise generate: evidence from your own floor about whether the fleet will work.
Why a Pilot Must Be an Experiment, Not a Demo
A demo answers the wrong question. It shows that a robot can work in ideal conditions, on the vendor's routes, for the length of a sales visit. The question you actually need answered is whether this robot, on your floors, with your clutter, your Wi-Fi and your staff, delivers the coverage and reliability the business case assumes. Only a structured pilot answers that.
The difference is discipline. An experiment has a hypothesis, a duration, a measured result and a decision threshold fixed in advance. A demo has none of those, which is why a demo almost never produces a no. If the success criteria are set after the robot has been running, they will be set to match whatever the robot happened to do, and the pilot becomes a rubber stamp for a decision already made.
Write the hypothesis before the unit ships. A workable one reads: "A single AOMAN C2 Pro on the second-floor office loop will hold coverage above 92% of the addressed area across a four-week trial at a navigation success rate above 97%, with fewer than two human interventions per shift." Every clause is measurable, and every clause can fail.
Scoping the Pilot: One Scene, One Fleet Size
Pilots fail from scope creep more often than from hardware. A pilot that tries to cover the whole building in the first week produces data you cannot interpret, because when something goes wrong you cannot tell whether the fault is the robot, the route or the workload. Narrow the scope deliberately.
| Scope element | Recommended pilot setting | Why |
|---|---|---|
| Area | One floor or one defined zone | Isolates variables; makes route mapping tractable |
| Fleet size | One or two units | Tests the product; fleet coordination is a later stage |
| Shift | One shift, ideally the harder one | Tests under realistic, not best-case, occupancy |
| Task set | The single highest-frequency task | Clean signal on the work being replaced |
| Duration | Four to six weeks | Long enough to hit failures and a full maintenance cycle |
| Human role | One named operator per shift | Produces the intervention data you need |
Four to six weeks is not arbitrary. A one-week trial will not encounter the battery degradation curve, the filter-loading cycle or the software update that breaks a route. Those events are exactly what the pilot must expose, and they appear on the second or third week, not the first. The readiness prerequisites that make a pilot interpretable in the first place are covered in the site survey and readiness assessment.
The KPIs That Decide the Pilot
Measure a small set of KPIs, and fix the target for each before the unit arrives. Five is enough. More and the daily report becomes unreadable and the evaluation loses focus.
| KPI | Definition | Sample target | Data source |
|---|---|---|---|
| Coverage rate | Addressed area actually cleaned per shift | > 92% | Fleet dashboard, logged daily |
| Navigation success | Routes completed without human rescue | > 97% | Intervention log |
| Interventions per shift | Times a person had to unstick or assist | < 2 | Operator log |
| Availability | Shifts the robot started on time and finished | > 95% | Shift log |
| Actual cycle time | Minutes to clean the pilot route | Within 15% of plan | Time-stamped route data |
These five KPIs map directly onto the numbers a fleet business case depends on. Coverage and cycle time feed the throughput arithmetic in the deployment throughput guide. Interventions and availability feed the uptime and cost assumptions. If the pilot hits its targets, the business case is validated with your own data; if it misses, you have learned it for the price of one unit instead of a fleet. The contractual framing that lets you turn pilot results into firm obligations is set out in the uptime and SLA contract guide.
Paid, Not Free: Why the Commercial Frame Matters
A free pilot is treated as a favour by both sides, and favours do not generate rigorous data. Pay a defined pilot fee, however nominal, and the arrangement changes: the vendor commits resources to make it succeed, and you carry an obligation to run it properly and produce a verdict.
Structure the fee so it credits against the fleet order if you proceed. That removes the incentive to view the pilot as a cost and makes it, in commercial terms, a deposit on the decision. Spell out in the pilot agreement what the vendor provides: the unit, the install, the onboarding, the dashboard access and a named support contact with a defined response time. A pilot without a support commitment measures the robot with the vendor watching over it, which is not the fleet you will actually run.
The Go/No-Go Gates
Do not leave the decision to the end of the pilot as a single judgement. Set three gates, and let each one end the pilot early if it fails. Early termination is a feature, not a failure: it saves the remaining weeks and the remaining cost.
- Gate 1, week 1, install and baseline: the robot maps the route, completes a full clean cycle, and produces interpretable dashboard data. Fail here means the scoping was wrong and the pilot should be re-scoped, not continued.
- Gate 2, week 3, reliability: interventions per shift are trending below target and navigation success is holding. Fail here means the environment or the product is not ready, and a fleet rollout would multiply the problem.
- Gate 3, week 6, economics: measured coverage and cycle time validate the business case range, and the projected fleet cost per cleanable square metre beats the manual baseline. Fail here is the honest finding that this robot does not pay back in this building, which is worth knowing now.
The baseline that Gate 3 compares against is the manual productivity measurement set out in the productivity baseline method. Running the pilot without that baseline leaves Gate 3 with nothing to compare to, and an economic gate that cannot be evaluated is not a gate at all.
What to Do with the Pilot Result
A passed pilot is the strongest negotiating position a buyer ever holds. You are no longer evaluating a proposal, you are accepting a proven configuration, and the terms can be tied to the measured results: the SLA can reference the intervention rate you observed, the coverage commitment can reference the coverage you measured, and the support response time can reference the support you actually received.
A failed pilot is equally valuable. It stops a fleet purchase that would have underperformed for years, and the intervention log tells the vendor precisely which environmental factor defeated the unit, which is often fixable. Either outcome of a well-run pilot is a win, which is why the discipline of fixing the hypothesis, the KPIs and the gates before the robot arrives is the cheapest part of the whole procurement.
Frequently Asked Questions
How much should a pilot cost? Structure it as a defined fee that credits against the fleet order, typically a fraction of one unit's price. The point is commitment, not revenue.
Can one pilot cover both cleaning and delivery robots? No. Run them separately. A combined pilot doubles the variables and halves the clarity, and the KPIs for the two use cases barely overlap.
What if the vendor refuses a paid pilot with defined gates? Treat it as a warning. A vendor confident in the product accepts measurement; a refusal to define success criteria before the trial is a signal about the support you will receive after the sale. Where several vendors pass their pilots, the comparison between them is decided by the weighted procurement scoring matrix.
