Service Robot Pilot Program: A 30-60-90 Day Deployment Roadmap for Facility Managers
At a glance: Most pilots fail on design, not hardware: a demo pilot shows the robot in the easiest 20% of the facility, while an experimental pilot with a 14-day baseline answers whether it works in yours. This roadmap gives you the 30-60-90 structure — baseline, stabilization, evaluation — and the thresholds to sign before a single robot arrives.
Consider an illustrative pilot run on two floors of the same facility. The vendor’s materials suggest a significant reduction in overnight cleaning labor. One zone is hard-surface corridor — the robot performs as expected, approaching the target. The other is carpeted floors with gurneys, spills and patient traffic — the robot’s improvement barely registers, dragging the average below the go/no-go threshold. Procurement concludes the technology underperformed and shelves the initiative. What actually failed was the pilot design: one robot, two very different operating conditions, one averaged number.
This is the pattern across service robot pilots: teams design them as demonstration projects — “let’s see if the robot can do the job” — rather than as structured experiments that isolate variables and quantify outcomes. The vendor configures the robot for the easiest part of the facility, staffs the demo themselves, and handles every edge case off-camera. An experimental pilot asks the harder question: under what conditions does the robot succeed, under what conditions does it fail, and what is the net operational impact across your actual facility? This guide is the 30-60-90 framework that makes a pilot a procurement-grade decision tool.

Why Most Pilots Fail: Demonstration vs. Experiment
A demonstration pilot asks: can the robot do the job? The answer is almost always yes — because the robot is configured for a showcase environment and the vendor team handles the edge cases. An experimental pilot asks: under what conditions does it succeed, and what is the net impact when those conditions are averaged across our facility? The design differences are concrete:
- Zone selection: the vendor wants the easiest zone; you need zones that represent your operational variance — floor surface, traffic, obstacles, lighting, Wi-Fi
- Measurement ownership: the vendor’s dashboard reports robot uptime and task counts; your team must independently measure labor hours, quality and delivery times
- Success thresholds: declare the go/no-go criteria before any robot data arrives, so mixed results cannot be rationalized after the fact
A 12-point vendor scorecard applied to a poorly designed pilot produces precise-looking data that is operationally meaningless. Design the pilot first; evaluate vendors within that design.
Phase 1: Days 1–30 — Baseline Collection and Pilot Design
The first month of a service robot pilot should not involve robots at all. This is the most commonly violated rule in pilot design, and the most expensive to ignore.
Weeks 1–2: Establish Your Baseline
Before any robot arrives, collect 14 days of baseline data on the metrics the robot is supposed to improve — measured manually, with stopwatch timing, staff logs and existing checklists, not vendor dashboards.
For cleaning robots (C1 / C2 Pro):
- Square footage cleaned per shift — actual, measured, not spec-sheet
- Labor hours allocated to the pilot zone, including supervision, supply replenishment and equipment maintenance
- Post-cleaning inspection pass rate: swab tests for hygiene-sensitive areas, visual checklist elsewhere
- Complaint and re-clean request frequency
For delivery robots (D1):
- Average request-to-fulfillment time by hour, to capture peak vs. off-peak variance
- Staff hours allocated to the task, including walking time to and from supply rooms
- Delivery error rate — wrong item, wrong room, wrong quantity
- Time diverted from primary duties to manage deliveries
For guidance robots (G1):
- Visitor wait time from arrival to service
- Queue length during peak hours
- Desk staff hours allocated to repeated directional questions
- Visitor satisfaction scores from a pre-pilot survey
Weeks 3–4: Design the Pilot Protocol
- Select pilot zones by operational variance, not convenience. Two or three zones that differ on the variable most likely to affect robot performance — floor surface, traffic density, obstacle frequency, lighting or Wi-Fi strength.
- Define success thresholds before seeing any robot data. Write: “we will proceed to full deployment if the metric improves by at least X% in Y of pilot zones and degrades below baseline in none.” Sign and date it.
- Assign measurement to your team. If you cannot measure the operational outcome without the vendor’s software, you cannot evaluate the pilot.
The ROI guide provides the financial frame for converting these operational metrics into dollars — but the raw data must come from your facility.
Phase 2: Days 31–60 — Active Pilot Execution
Weeks 5–6: Deployment and Stabilization
Deploy robots to the selected zones. The first two weeks produce unreliable data: mapping is still being refined, staff are still learning the interaction protocols, edge cases are still being resolved. Document everything — but keep weeks 5–6 out of the go/no-go decision. Key activities:
- Map the pilot zones completely, including elevator lobbies, fire doors and dock transitions
- Train staff on what to do when the robot stops and how to report issues without bypassing the robot entirely
- Run change management in parallel — staff resistance during the first week or two is a leading source of false-negative results
- Log every stoppage with timestamp, location, root cause and resolution time
Weeks 5–8: Making the Vendor Answer for the Pilot
An experimental pilot is still a joint project, and the vendor’s contribution should be in writing before it starts:
- Mapping and calibration support: the vendor maps all pilot zones and owns the stabilization weeks — fixing their system, not yours
- Integration checkpoints: door, elevator and network integration is tested on a fixed schedule, with written status per checkpoint
- Failure-response commitment: a named engineer for the pilot window and a written response target per failure category
- Data access: export rights for all pilot data — tasks, failures, per-robot hours — so the evaluation is not limited to their dashboard
- Replacement unit policy: a spare unit on site for the evaluation window, or a written swap commitment within 48 hours, so a hardware fault does not invalidate the dataset
All of these are standard to request and easy to get a yes on. The vendors who resist written commitments during a pilot are the ones who save their best behavior for after the signature.
Weeks 7–8: Data Collection Under Normal Operations
For 14 consecutive days, collect the same metrics as baseline — same measurement methods, same times, same people. The only variable that should have changed is the presence of robots. Do not switch measurement methods mid-collection; if baseline used stopwatch timing and the pilot uses the vendor dashboard, the comparison is invalid. Watch what the maintenance and TCO guide calls the hidden labor: who empties collection bins, refills tanks and clears obstructions — and count that time.
Phase 3: Days 61–90 — Analysis and Decision
Weeks 9–10: Data Analysis
Compare pilot metrics against baseline per zone, never as one average. A robot that improves Zone A by a third and degrades Zone B slightly is not “a modest average improvement” — it is a robot that works under Zone A conditions and fails under Zone B conditions. The deployment decision should reflect that: deploy to Zone A, redesign Zone B, or pick a different robot for Zone B. The question the pilot must answer is whether the robot’s performance envelope covers your facility’s operational variance — not whether it works somewhere.
Weeks 11–12: Procurement Decision and Contract Negotiation
If the pilot passed your pre-defined thresholds, you now have leverage: actual throughput, actual downtime rate, actual labor displacement in your environment. Use it to negotiate:
- Performance guarantees tied to pilot-verified metrics, not spec-sheet values
- Uptime SLAs with penalty clauses calibrated to what the pilot actually showed
- Right-sized fleet: the pilot tells you how many units you need, which is often fewer than the initial proposal
If the pilot failed, the data still has value. An honest “no” is worth more than a false “yes” — the false yes leads to a deployment that underperforms for years before anyone admits it. Document the failure mode, share it across departments, and revisit in 12–18 months.
The Pilot Design Checklist
Before signing a pilot agreement with any vendor, confirm all of the following:
- Baseline data collected for at least 14 days on all target metrics
- Pilot zones selected to represent operational variance, not vendor convenience
- Success thresholds defined and signed before robot deployment
- Measurement performed by your team using methods identical to baseline
- Weeks 5–6 designated as stabilization; data excluded from evaluation
- Weeks 7–8 designated as evaluation; data compared to baseline per zone, not averaged
- Vendor dashboard data cross-validated against independent measurement
- Staff trained on interaction protocols and failure reporting before deployment
- Change management plan active from day one of robot deployment
- Go/no-go decision scheduled for day 90 with pre-defined criteria
A pilot that follows this framework costs approximately the same as a demonstration pilot — the difference is not in budget but in discipline. And the discipline pays for itself the first time it prevents a fleet-scale deployment that would have underperformed because the pilot was built to succeed rather than to inform. Tell us your facility and the metric you care about most — AOMAN FUTURE will structure a pilot with pre-defined thresholds, your measurement team, and a day-90 decision gate built in.
