Part III: Running a trained policy · §36 of 43

Change one thing at a time, and always on short --duration runs first.

36.1 How often the policy looks: n_action_steps

ACT predicts 100 actions at once. With the default n_action_steps=100, the robot plays all 100 (3.3 s at 30 fps) without looking, then asks the network again. If the object moves or the grasp slips, the policy notices only at the next chunk.

Lower values make it re-plan more often:

bash run_policy.sh --policy.n_action_steps=30     # re-plan every 1 s
  • Pros: reacts faster to what the cameras see.
  • Cons: more network calls (still cheap for ACT), and more chunk boundaries, i.e. more small hitches.
  • n_action_steps can’t exceed chunk_size (100).

36.2 Smoother motion: temporal ensembling (ACT)

Instead of playing one chunk and then switching to the next, ACT can run the network every tick and average the overlapping predictions for the current tick. This removes the chunk-boundary hitches and usually gives the smoothest motion:

bash run_policy.sh --policy.n_action_steps=1 --policy.temporal_ensemble_coeff=0.01
  • n_action_steps must be 1 with ensembling (LeRobot refuses otherwise).
  • 0.01 is the value from the ACT paper. Larger values weight the oldest predictions more.
  • The network now runs 30 times per second. ⚠ verify on the RTX 5050 that the cadence summary still shows ~30 Hz (the infer line).

36.3 Smoother commands: —interpolation_multiplier

bash run_policy.sh --interpolation_multiplier=2

Sends 2 commands per policy action (60 Hz to the motors) by interpolating linearly between consecutive actions. The policy and any recording stay at 30 Hz. Helps with “stair-step” motion at kp 240. It doesn’t fix chunk-boundary jumps (use §36.2 for that). Each motor command also reads the current position when max_relative_target is set, so going above 2–3 can saturate the CAN loop. Check the cadence summary.

36.4 Step limit: max_relative_target

The guard against a single huge jump. For each command and joint: target = current + clamp(target − current, ±max_relative_target).

ValueEffect
unsetNo guard at all. Don’t.
5Very cautious, but the low-stiffness wrist joints (kp 24–31) stall a few degrees short of their target and the log fills with “clamped” warnings (§23.3).
10Recommended start: limits a wild jump to 10° per tick, rarely clamps a normal policy motion.
20+Little protection.

The value is per arm. A per-joint form also exists ({joint_1: 10.0, …, gripper: 10.0}, all 8 names required). ⚠ verify the command-line syntax before relying on it.

36.5 Stiffness: position_kp and position_kd

LeRobot’s follower is ~3.4× stiffer than the ROS defaults (§20). Softer gains make contacts gentler but the arm sags more and lags its targets:

bash run_policy.sh \
  --robot.left_arm_config.position_kp='[120,120,120,120,24,31,25,25]' \
  --robot.right_arm_config.position_kp='[120,120,120,120,24,31,25,25]'

The policy was trained on observations recorded at kp 240. With softer gains, the joint angles it sees differ from training (more sag), so it may behave differently. Treat it as an experiment, not a default.

36.6 Loop rate: —fps

Keep --fps equal to the dataset’s fps (30). What you can control is whether the loop actually reaches it. Every run ends with a cadence summary, for example:

Cadence summary — whole run …
  effective cadence: 29.84 Hz policy …
  cycles over the 33.3 ms work budget: 12/1197 (1.0%) — work mean 18.7 ms, worst 48.1 ms
  loop-body steps (share of measured work):
    observe      mean   3.13 ms …
    infer        mean   5.18 ms …
    send         mean   0.40 ms …
  pacing headroom: 7.4 ms slept per tick on average …
  • effective cadence well below 30 Hz → the motion is slower than in training.
  • High observe share → cameras (fewer or smaller cameras, no Rerun).
  • High infer share → the network (no ensembling, smaller model, or RTC for VLAs).
  • pacing headroom near 0 → no margin; expect jerks.

36.7 Faster inference: —use_torch_compile

--use_torch_compile=true compiles the network on the first calls. The robot waits during the warm-up (a few seconds to minutes) and then inference is faster. Not needed for ACT. Worth trying for Diffusion or SmolVLA if infer dominates the cadence summary.