Service Robot Incident Response and Escalation Protocol
At a glance: Every fleet will fail. A scrubber will stall mid-aisle, a delivery unit will mis-dock, a sensor will drop a reading at the worst moment. What decides whether that failure costs ten minutes or a lost client is not the robot, it is the protocol agreed long before the incident. This guide sets out severity tiers, the escalation ladder with clock times, the metrics that expose a weak response, and the contract clauses that keep the supplier accountable for the recovery clock.
Why an Incident Protocol Must Exist Before the Fleet Arrives
The moment a robot stops in a live operation, the response is being written whether or not a document exists. If no protocol exists, the response is improvised: whoever notices tries to fix it, the supplier is called at an unknown number, and the delay is discovered only when a stakeholder asks why the lobby was not cleaned. Improvisation is expensive, and its cost lands entirely on the operator.
A written protocol converts that improvisation into a procedure. Every incident has a severity, every severity has a first responder and a supplier contact with a clock, and every recovery has a measured duration. The economics are direct: the difference between a 20-minute recovery and a 6-hour recovery on the same fault is the difference between a routine event and a service failure the operator has to explain upward. The method for measuring that recovery clock is the same discipline used in the KPI benchmark set, applied to failure rather than throughput.
Build the protocol during the deployment phase, not afterward. The site survey and readiness assessment that precede installation, covered in the facility readiness guide, already define who will own the robot on site. That owner is the first responder. Naming them now costs nothing; naming them during an incident costs the response time.
Severity Tiers: Classifying Before You React
The first decision in any incident is not "what do we do" but "how serious is this". Tiers must be defined in advance so the response is automatic. A workable four-tier scheme follows; adjust the boundaries to the operation, but keep four tiers so that routine faults never trigger the emergency path.
| Tier | Definition | Operational impact | First response target |
|---|---|---|---|
| S4, Advisory | Non-blocking warning; robot still productive | None today; trend worth watching | Log, review at weekly maintenance |
| S3, Degraded | One subsystem impaired; output reduced but continuing | Coverage or throughput below plan | Remote diagnosis within 4 working hours |
| S2, Stopped | Robot halted and cannot resume unaided | Scene uncovered; manual cover required | On-site or remote fix within 8 hours |
| S1, Safety or Security | Contact risk, blocked egress, door or lift jam, data event | Immediate hazard or compliance exposure | Immediate isolation; supplier paged within 1 hour |
Two rules make the tiers work. First, the tier is set by the worst plausible outcome, not the most likely one: a stopped robot near a fire exit is S1 even if the fault itself looks trivial. Second, the tier can only be raised by the first responder, never lowered without the operations manager's sign-off. Lowering a tier to avoid escalation is the most common way a protocol quietly stops working.
The Escalation Ladder: Named People, Clock Times
An escalation ladder without clock times is a phone tree. Each rung needs three things: a named role, a contact method, and the elapsed time since incident detection at which that rung activates. The elapsed time is the whole point, because it removes the judgement call that delays escalation in practice.
| Elapsed from detection | Rung | Action |
|---|---|---|
| 0 min | On-site operator | Confirm tier, make safe, photograph, log the incident ID |
| 15 min | Internal technical lead | Attempt documented remote recovery steps; confirm coverage gap |
| 1 hour | Supplier support (S1 and S2) | Ticket opened with incident ID, logs attached, response clock started |
| 4 hours | Supplier account manager | Written status and estimated recovery time demanded |
| 24 hours | Contract escalation clause | Service-credit or SLA-breach notice triggered in writing |
| 72 hours | Executive review | Continuity decision: substitute units, partial redeployment, or suspension |
Attach the robot's incident ID, the fault code, the time of detection and a photograph to every rung hand-off. A supplier who receives a ticket with logs and a photograph diagnoses faster than one who receives a phone description, and the difference shows up directly in the recovery clock. The commercial teeth for the 24-hour rung come from the service-level terms described in the uptime SLA guide; an escalation ladder without a credit clause behind it is advisory.
The Metrics That Expose a Weak Response
A protocol that is never measured decays. Track four figures per month and review the trend, not the individual incident. Each is cheap to record if the incident log is disciplined.
- Mean time to detect (MTTD): how long between a fault occurring and someone noticing. High MTTD means monitoring gaps, and it is the single most fixable number on this list. Remote monitoring that raises an alert on fault codes cuts it from hours to minutes.
- Mean time to respond (MTTR-ack): how long to the first supplier acknowledgement. This is a supplier metric; hold them to it in the contract and compare it against the promised first-response window.
- Mean time to recover (MTTR): how long until the robot is productive again. Field data across service-robot fleets commonly shows a split of roughly 60 to 70 percent remote-resolvable faults and 30 to 40 percent requiring a site visit, so MTTR is best tracked separately for each class.
- Repeat-incident rate: the share of incidents sharing a root cause with a previous incident. A rising repeat rate means fixes are temporary and the maintenance schedule is not absorbing the lesson. The interval arithmetic for that schedule lives in the preventive maintenance guide.
Publish the four numbers monthly alongside uptime. A supplier whose MTTR-ack is drifting upward is a supplier whose support is being deprioritised, and the trend shows it weeks before the client complaints do.
Root-Cause Discipline: Closing the Loop
Recovery is not resolution. An incident is only closed when the root cause has been identified, classified, and fed back into either the maintenance schedule, the operating procedure, or the supplier's defect record. Adopt a simple three-way classification and record it against the incident ID:
| Classification | Meaning | Corrective owner |
|---|---|---|
| Product defect | Fault traced to hardware or firmware under warranty | Supplier, with defect report and firmware or part replacement |
| Environmental | Fault traced to floor, layout, network or traffic pattern | Operator, via the survey and readiness checklist |
| Operational | Fault traced to procedure, consumable or handling error | Operator, via training and the induction plan |
The classification matters because it routes the fix to the party who can actually prevent recurrence. Misclassify an environmental cause as a product defect and the supplier replaces a part that was never broken; the fault returns within the month. Operational causes are the most under-recorded and the most common in the first ninety days, when staff habits are still forming; the phased induction described in the staff induction plan is the upstream fix for most of them.
What to Write Into the Supply Contract
The protocol is only as strong as the clauses behind it. Before signature, confirm the contract contains, at minimum: a defined first-response window per severity tier, a service-credit or remedy for breach of that window, an obligation to supply incident logs and fault codes on request, a spare-parts lead time for the failure modes you consider critical, and the right to escalate to a named account manager rather than a general support queue. The parts-planning consequences of those lead times are covered in the spares and consumables planning guide.
Where the fleet is leased or financed rather than bought, the same clauses must survive the financing structure. A lease that bundles support into a monthly fee still needs a written response window, or the operator is paying for availability without any mechanism to enforce it. The separation of hardware, service and financing is set out in the service contract versus warranty economics guide.
Frequently Asked Questions
How many severity tiers should a small fleet use? Four. Three collapses safety and stoppage together, which is exactly the distinction that must stay sharp. Five invites debate during an incident, which is the wrong time to argue.
Who should be the first responder on site? The named operator who owns the robot day to day, not the IT or facilities manager. Proximity and familiarity resolve more S4 and S3 incidents than any remote tool.
Is a supplier's own incident app enough? No. The supplier's tool logs their side of the ticket; you need your own incident log to measure MTTD and MTTR and to prove an SLA breach. Run both, and reconcile them monthly.
How quickly should a stopped robot be physically covered? Within the same shift. Manual cover is the bridge that keeps the scene compliant while the response clock runs, and it must be scheduled in advance, not improvised during the incident.
