Robot end-to-end latency: build a sensor-to-actuator timing budget
Inference time and control rate do not equal robot reaction time. Trace one physical event, budget every stage, validate clocks and accept tail timing, missed deadlines and stale-data behavior under representative load.

The short answer: do not accept robot timing from model inference milliseconds or average loop rate. Define the start and end of one physical event from sensor capture to command application, set its maximum permissible data age, allocate the budget across acquisition, transport, scheduling, compute, control and hardware, then test tail latency, deadline misses, loss and timeout behaviour under representative load.
Separate four timing quantities before measuring
ROS 2 REP-2014 distinguishes latency, system reaction time, software system reaction time, message latency and execution latency. That distinction changes acceptance. Publisher-to-subscriber latency covers one software interval; it does not include camera exposure or scanner acquisition, and it may end well before a drive applies a command or motion becomes observable.
| Quantity | Practical start and end | What it answers | What it cannot prove alone |
|---|---|---|---|
| Period | Planned or actual time between loop iterations or samples | Whether a component runs at its stated rate | Data freshness or timely command application |
| Execution time | Entry to exit of one callback, inference or control calculation | Compute consumed by that stage | Queueing, transport or hardware wait |
| Message age | Source timestamp to receipt or use | How old the consumed data is | True age when clocks are not aligned |
| System reaction time | Physical stimulus or sample to observable actuator response | The result the task actually experiences | Which stage caused a long tail without stage traces |
What belongs in a sensor-to-actuator budget?
Draw a timestamped event chain for every control-relevant path instead of adding nominal module rates. Exposure, scan or encoder sampling marks data creation. Driver delivery, ROS 2 publish/take, callback start, inference or planning completion, controller update, bus transmission, drive application and the first feedback change are separate events. Fixed event semantics let teams decide whether they are improving data age, computation or physical response.
| Stage | Evidence to record | Common blind spot | First checks when over budget |
|---|---|---|---|
| Capture and driver | Exposure/scan/sample time, sequence, driver delivery | Treating driver callback time as capture time | Sensor buffering, batching, drops and timestamp source |
| Transport and queueing | Publish, receive, take, callback start and queue depth | Testing only idle ping or one topic | Payload, QoS, network load and executor backlog |
| Perception/inference/planning | Input revision, start/end and state timestamp used | Fast inference over stale input | Pre-processing, synchronisation waits, GPU queue and batching |
| Control update | read–update–write period, execution and overruns | Average rate hiding isolated deadline misses | Scheduling, memory, locks, hardware I/O and priority inversion |
| Hardware and physical response | Command sequence, drive receive/apply and first feedback change | Equating software publish with motion | Bus cycle, drive buffer, inner loop and mechanism dynamics |
Each tool covers a segment; acceptance joins the evidence
ROS 2 topic statistics can report message age and period with mean, minimum, maximum, standard deviation and sample count. That helps expose transport and arrival irregularity, but its interpretation still depends on timestamp semantics and clock integrity. The official ROS 2 tracing tutorial shows how to capture execution traces and analyse callback duration. Tracing can locate software queueing and execution; it does not automatically see a sensor's internal buffer or when a motor physically responds.
The ros2_control Controller Manager documentation defines its real-time loop as read, update and write, and exposes diagnostics for periodicity, controller and hardware execution time, and overruns. Those diagnostics belong in an acceptance record, but their defaults are not universal robot requirements. Derive project limits from motion, contact, speed, stopping distance, data freshness and the required degraded response.
Clock synchronisation passing does not mean latency passes
Before subtracting timestamps across computers or device clocks, prove their relationship. LinuxPTP ptp4l synchronises PTP clocks, while phc2sys is commonly used to align a system clock with a PTP hardware clock. Keep synchronisation state, offset, source, time scale and faults with the timing result. Clock error corrupts cross-device latency; a small offset, however, does not remove application queueing, scheduling, buffering or compute tails.
A seven-step end-to-end acceptance workflow
- Derive deadlines from the task. Define input freshness, reaction endpoint and timeout action separately for avoidance, vision guidance, force control, teleoperation or learned policy control. Do not copy a universal millisecond target.
- Draw the event chain. Give capture, driver delivery, publish, take, callback, compute, control update, hardware apply and feedback change explicit event IDs.
- Inventory clocks. Record whether each timestamp comes from the sensor, host, hardware clock or controller. Verify offset and clock-jump handling.
- Measure the segments first. Use topic statistics, tracing, controller diagnostics and device logs, then connect the path with a shared sequence or trigger.
- Run representative load. Include real payload sizes, cameras or point clouds, inference, recording, network traffic and expected background services. Keep idle results as a baseline, not the release result.
- Report distributions and violations. Preserve sample count, percentiles, observed maximum, misses, loss/reordering, clock faults and test duration. A mean cannot replace tail evidence.
- Inject stale and late data. Verify whether the system rejects, holds, slows, stops in a controlled way or asks for intervention when state, command, cycle or synchronisation expires. Connect that behaviour to the robot recovery-state contract.
Boundary: this is a measurement and acceptance framework, not a universal latency requirement for any robot, controller, network or AI model. Software timing evidence is not safety-function validation. Safety stops, collaborative operation and functional safety require the responsible parties to apply the relevant standards, risk assessment and validated safety chain.
Frequently asked questions
Why can a robot with low average latency still fail acceptance?
A mean hides isolated queueing, scheduling, buffering and missed cycles. Control tasks also need tail latency, observed maximum, data age, violation count and the action taken after a deadline is missed.
Does a 1 kHz control rate prove sub-millisecond reaction time?
No. It states the nominal period of one loop. Capture, phase wait, transport, compute, bus transfer and actuator response can span multiple periods and must be traced on the same event.
Can ROS 2 topic statistics measure total sensor-to-motor latency?
Not alone. They can report message age and period, but end-to-end acceptance also needs capture, internal execution, hardware application and physical feedback evidence, with validated clock relationships.
If PTP synchronisation passes, is load testing still necessary?
Yes. Synchronisation makes cross-clock comparisons credible; it does not remove queues, scheduler delays, congestion, GPU waits, drive buffers or deadline misses. Accept the two evidence sets separately.
Sources
These primary sources support the material facts and engineering boundaries discussed above.
Evaluating robot control, bimanual manipulation or a mobile platform?
Talk to Matrix Dimension →