Robot fault-recovery acceptance: reset is not restart
A cleared alarm does not prove that the robot, tooling, part and surrounding equipment are ready to continue. Test six event classes through a state-based recovery contract.

The short answer: permit an industrial robot task to continue only after the team has classified the cause, re-established the physical state, kept reset separate from start authorization, and validated the recovery point. Alarm acknowledged, safety reset, drives enabled and program executing are four different facts. One catch-all “Reset” command can otherwise restart motion with an unknown part, an invalid sequence step or people still inside the affected area.
A cleared fault is not a recovered cell
China's official standards portal lists GB/T 42983.2-2023, Industrial robots—Operation and maintenance—Part 2: Fault diagnosis, as a current recommended national standard. Fault diagnosis is therefore a distinct operations-and-maintenance layer, not merely an HMI message. The public scope of ISO 10218-2:2025 covers integration, commissioning, operation and maintenance of industrial robot applications and cells, including hazards in intended use and reasonably foreseeable misuse. A recovery decision must consequently include the end effector, part, fixture, conveyor, people and task controller—not only the robot arm.
ISO 14118:2017 addresses prevention of unexpected start-up from electrical, hydraulic and pneumatic supplies, stored energy such as gravity or springs, and external influences. ISO 13850:2015 sets design principles for emergency-stop functions. Their public scopes make the boundary clear: recovery is not just a software program counter. It must account for energy, access, suspended mechanisms, residual pressure and every machine that might move when permission returns.
Six event classes need distinct recovery contracts
| Event | What must be known first | Acceptable exit evidence | Common mistake |
|---|---|---|---|
| Normal pause/hold | Task step, part and device states remain coherent | Hold clears and the existing start permission remains valid | Treating every pause as a fault reset |
| Protective stop | Trigger cause, contact/limit state, people and obstruction | Checks, reset and recovery point follow the specific robot and application documents | Repeatedly clearing the alarm without finding the cause |
| Emergency stop | Who triggered it, whether the hazard is gone, device and stored-energy states | Device released and reset; a separate action then permits start | Release automatically resumes the task |
| Process/task fault | Part ownership, operation result and real tool/fixture feedback | Rework, reject, intervention or bounded retry updates task state | Blind resume from the interrupted line |
| Power/comms loss | Whether frames, program state, outputs, part ownership and peripherals remain trustworthy | Read back reality; re-home, unload or rebuild the baseline where needed | Trusting pre-interruption memory |
| Software component error | Failed component, output disposition and whether state is reconstructable | Known inactive state, then verified configuration and dependencies before activation | Process restart immediately enables motion |
Separate five actions: acknowledge, reset, ready, start and resume
The ABB AC500-S emergency-stop function-block page cites a central IEC 60204-1 principle: resetting the command must not restart the machinery; it only permits restarting. Device implementations still differ. The Universal Robots stop-recovery documentation, for example, distinguishes a powered, brake-released “Running” robot mode from a program that is actually executing. It also documents different sequences and program restart points for protective stop, robot emergency stop, system emergency stop and safety fault. This vendor example is not a universal procedure; it shows why a PLC or HMI should not compress every layer into one generic robot-normal bit.
- Acknowledge: record that a person or supervisor has seen the event; do not change the physical state.
- Reset: after the cause is removed, leave a specific safety or fault state without issuing a motion command.
- Ready: re-prove robot mode, axes/drives, tool, part, fixture, peripherals and protected-area conditions for this task.
- Start: use the independent action and permission defined by the risk assessment and control design to initiate automatic operation.
- Resume: continue only from a validated program point. If reality no longer matches the saved task step, route to unload, retract, re-home or supervised recovery.
A seven-step FAT/SAT recovery test
- Inventory reachable events. Combine alarm history, risk assessment, manuals and process FMEA. Include protective and emergency stops, contradictory sensors, failed grasp, tool fault, communication loss and power loss.
- Define the state vector. For each event, capture operating mode, program/task step, joint or TCP region, part ownership, tool/fixture feedback, peripheral state, safety state and software version. An alarm code alone cannot reconstruct a cell.
- Specify the stop outcome. State which actions stop, which energy remains, whether the part is retained, how outputs change and when recovery must hand over to an energy-isolation maintenance procedure.
- Separate authority. Name who may acknowledge, reset, authorize start and perform local-only steps. Give each HMI action a distinct label and feedback so that one click cannot silently cross multiple gates.
- Select a recovery target. Choose resume-in-place, return to a known pose, complete or abort the current process, unload and restart, or supervised intervention. Define entry conditions, attempt limit, timeout and terminal failure for each path.
- Recover software by state. ROS 2 Managed Nodes separates unconfigured, inactive, active, error-processing and finalized states under supervisory transitions. It is a useful ordinary-software pattern, not a safety function; hardware, safety control and application risk retain their own authority.
- Trigger and retest. Use FAT to prove logic, messages, permissions and records, then SAT to create approved real-device conditions. Verify the stop result, operator action, recovery path, part outcome, retry bound and event log; sample frequent recovery paths again in production.
When is automatic recovery worth evaluating?
- the risk assessment, application design and specific equipment manuals permit that reset or recovery;
- the trigger is observable and its disappearance cannot conceal a person or unknown obstruction;
- the actual robot, tool, part and peripheral states can all be re-established;
- the recovery has one target, bounded path, attempt limit, timeout and human handoff;
- safety reset, equipment readiness and task start remain separate auditable states;
- FAT and SAT have triggered the event and verified both logs and the resulting part.
Safety boundary: this is a state-and-acceptance method, not a safety program, automatic-reset authorization or conformity decision for a specific cell. Robot brands, stop categories, safety functions, tools and local requirements differ. Actual reset, restart and access conditions must come from released standards, the project risk assessment, equipment manuals and a validated safety design.
Frequently asked questions
Why not resume as soon as the alarm disappears?
A cleared diagnostic state does not prove robot position, part retention, fixture and peripheral feedback, protected-area status or a valid program recovery point. Re-establish those facts separately.
Can safety reset and start share one button?
Do not combine them merely to reduce operator steps. Reset should permit a later restart, not initiate motion by itself. The project risk assessment, applicable standards and safety design must define devices, locations, authority and logic.
Can a protective stop recover automatically?
There is no cross-brand, cross-application answer. Check the specific robot documentation, cause, risk assessment and observable site state. If the trigger, people or obstruction state is uncertain, automatic continuation is not justified.
If forced inputs passed in FAT, why repeat recovery tests in SAT?
Simulated FAT inputs mainly prove program logic. SAT adds real sensors, wiring, tools, fixtures, networks, operator actions and physical recovery paths. Safety-function validation should still follow its separately approved plan.
Sources
These primary sources support the material facts and engineering boundaries discussed above.
- 国家标准全文公开系统 — GB/T 42983.2-2023《工业机器人 运行维护 第2部分:故障诊断》
- ISO 10218-2:2025 — Industrial robot applications and robot cells
- ISO 14118:2017 — Prevention of unexpected start-up
- ISO 13850:2015 — Emergency stop function
- ABB AC500-S Safety User Manual — SF_EmergencyStop
- Universal Robots — Stop recovery
- ROS 2 Design — Managed nodes
Evaluating robot control, bimanual manipulation or a mobile platform?
Talk to Matrix Dimension →