Part III: Running a trained policy · §38 of 43
38.1 The idea
A policy fails in situations its demonstrations didn’t cover. DAgger collects exactly those: the policy runs, and when it’s about to fail you pause it, take over with the Minis, show the correction, and hand back. Every correction is saved as an episode, tagged intervention=True. Retraining on the original demos plus the corrections fixes the weak spots. LeRobot’s own guide for this (docs/source/hil_data_collection.mdx) uses exactly bi_openarm_follower + bi_openarm_mini.
Needs the keyboard (or a foot pedal)
Pausing and taking over are done with keys read by
pynput. On zeus’s current Wayland session they probably don’t work (§37.3). Switch to an X11 session first, or use a USB foot pedal (--strategy.input_device=pedal). Without either, you can’t take over, so don’t start a DAgger run.
38.2 Before you start
python ~/Documents/openarm_lerobot/check_mini_calibration.pymust pass. A wrong Mini calibration makes the takeover snap (§23.2).- Practise the key sequence (§38.4) once with
--strategy.num_episodes=1and a short task.
38.3 The command
lerobot-rollout \
--strategy.type=dagger \
--strategy.num_episodes=20 \
--policy.path=$HOME/Documents/lerobot/outputs/train/act_openarm_pick_cube/checkpoints/last/pretrained_model \
--robot.type=bi_openarm_follower \
--robot.id=my_bimanual_follower \
--robot.left_arm_config.port=can0 --robot.left_arm_config.side=left --robot.left_arm_config.max_relative_target=10.0 \
--robot.right_arm_config.port=can1 --robot.right_arm_config.side=right --robot.right_arm_config.max_relative_target=10.0 \
--robot.cameras='{top: {type: opencv, index_or_path: /dev/video0, width: 640, height: 480, fps: 30}}' \
--teleop.type=bi_openarm_mini \
--teleop.left_arm_config.port=/dev/serial/by-id/usb-1a86_USB_Single_Serial_5876043720-if00 --teleop.left_arm_config.side=left \
--teleop.right_arm_config.port=/dev/serial/by-id/usb-1a86_USB_Single_Serial_5876043592-if00 --teleop.right_arm_config.side=right \
--teleop.id=my_mini \
--dataset.repo_id=<hf_user>/rollout_hil_pick_cube \
--dataset.single_task="Pick up the red cube and place it in the bowl" \
--dataset.no_stamp=true \
--dataset.push_to_hub=false \
--device=cudaAt start the Minis ask “ENTER to use existing calibration”: press Enter. --strategy.num_episodes is the number of corrections to collect (falls back to --dataset.num_episodes when unset).
38.4 Controls
| Key (default) | Effect |
|---|---|
| Space | Pause / resume the policy. On pause the followers hold, and LeRobot powers the Minis and drives them to the followers’ pose over 2 s. Hands off the Minis while they move; take hold of them once they stop. |
| Tab | Start / stop a correction. On start the Minis go limp and the followers follow them: you demonstrate. On stop the Minis are powered again to hold. |
| Space (again, while paused) | Resume the policy; the Minis go limp. |
| Enter | Push the dataset to the Hub (only with push_to_hub=true). |
| Esc | End the session (followers return to the start pose, then torque off). |
Typical cycle: policy runs → it’s about to fail → Space (pause, wait for the Minis to stop moving, grab them) → Tab (correct the motion) → Tab (end correction) → Space (policy continues).
Other options: --strategy.record_autonomous=true also records the autonomous stretches (continuous, size-based episodes). --strategy.input_device=pedal with --strategy.pedal.device_path=/dev/input/by-id/… uses a USB foot pedal. --strategy.smooth_handover=false disables the powered Mini move. Don’t: the follower would then jump to wherever the Mini is when the correction starts.
38.5 Retrain on demos + corrections
Training on two datasets at once isn’t supported in v0.6.2 (MultiLeRobotDataset isn't supported for now), so merge them first, then train as in §24.3 with a new --output_dir / --job_name:
lerobot-edit-dataset \
--new_repo_id <hf_user>/openarm_pick_cube_plus_hil \
--operation.type merge \
--operation.repo_ids "['<hf_user>/openarm_pick_cube', '<hf_user>/rollout_hil_pick_cube']"
lerobot-train --dataset.repo_id=<hf_user>/openarm_pick_cube_plus_hil … (as §24.3, new --output_dir/--job_name)Use the real names of both datasets (ls ~/.cache/huggingface/lerobot/<hf_user>/). Without --dataset.no_stamp=true, both carry a date-time suffix. Repeat rollout → corrections → merge → train until the policy stops needing corrections.
38.6 Other recording strategies
| Strategy | What it records | When useful |
|---|---|---|
highlight | Keeps the last --strategy.ring_buffer_seconds (default 10 s) in memory. Press s to save that buffer and keep recording, s again to close the episode. h pushes to the Hub. | Long autonomous runs where you only want to keep the interesting moments (e.g. failures). Needs the keyboard. |
sentry | Everything, continuously, in size-based episodes, uploading every --strategy.upload_every_n_episodes (5). | Unattended long runs. Not for this setup yet: the robot must never run unattended. |
Both need --dataset.repo_id=<hf_user>/rollout_….