Part III: Running a trained policy · §37 of 43
37.1 What it does
--strategy.type=episodic runs the policy as a series of episodes with a reset phase between them, and records everything as a LeRobot dataset, like lerobot-record but with the policy instead of you. Use it to measure a success rate, compare checkpoints, or keep good rollouts as extra training data.
For each episode:
- Policy state is cleared; the policy runs for up to
--dataset.episode_time_sseconds (→ ends it early). - Episode saved.
- Reset phase of
--dataset.reset_time_sseconds: the arms go back to the start pose (in 1 s, faster than the 3 s at the end of a run) and hold there, torque on, while you reset the objects. With the Minis attached you teleoperate the reset instead (§38). - Next episode, until
--dataset.num_episodesare recorded or Esc.
37.2 The command
lerobot-rollout \
--strategy.type=episodic \
--policy.path=$HOME/Documents/lerobot/outputs/train/act_openarm_pick_cube/checkpoints/last/pretrained_model \
--robot.type=bi_openarm_follower \
--robot.id=my_bimanual_follower \
--robot.left_arm_config.port=can0 --robot.left_arm_config.side=left --robot.left_arm_config.max_relative_target=10.0 \
--robot.right_arm_config.port=can1 --robot.right_arm_config.side=right --robot.right_arm_config.max_relative_target=10.0 \
--robot.cameras='{top: {type: opencv, index_or_path: /dev/video0, width: 640, height: 480, fps: 30}}' \
--dataset.repo_id=<hf_user>/rollout_act_pick_cube_last \
--dataset.single_task="Pick up the red cube and place it in the bowl" \
--dataset.num_episodes=10 \
--dataset.episode_time_s=40 \
--dataset.reset_time_s=20 \
--dataset.fps=30 \
--dataset.push_to_hub=false \
--device=cudaRules for the dataset flags, all enforced by LeRobot v0.6.2:
| Rule | Why / error if broken |
|---|---|
The dataset name must start with rollout_ | Dataset names for rollout must start with 'rollout_'. (V2.2 of this guide used eval_… and hil_…, which are rejected.) |
A date-time tag is appended to the name, e.g. rollout_act_pick_cube_last_20261010_143000 | Each session gets a new dataset. Add --dataset.no_stamp=true to keep the exact name (a second run with the same name then needs --resume=true, or fails because the dataset exists ⚠ verify the exact error). |
Always set --dataset.push_to_hub=false | The default is true: at the end it tries to upload to the Hub and fails without hf auth login. |
Set num_episodes, episode_time_s, reset_time_s | Defaults are 50 episodes of 60 s with 60 s resets. |
--dataset.single_task | Becomes the task text (no separate --task needed). |
--dataset.* flags with --strategy.type=base | Rejected: base strategy does not record data. |
The recording goes to ~/.cache/huggingface/lerobot/<hf_user>/rollout_…/. View it with lerobot-dataset-viz --repo-id <the stamped name> --episode-index 0.
37.3 Keyboard controls and the Wayland problem on zeus
| Key | Effect |
|---|---|
| → | End the current episode (or reset phase) now |
| ← | Discard the current episode and record it again |
| Esc | Stop the session (return to start pose, torque off) |
These keys are read by pynput, which needs an X11 session. zeus currently logs in to a Wayland session (echo $XDG_SESSION_TYPE → wayland), where the listener most likely receives nothing ⚠ verify. Without the keys the session still works: episodes and resets simply run for their full time, and Ctrl+C once ends the session (an episode in progress at that moment may be lost). To get the keys: log out, click the gear icon on the login screen, choose “Ubuntu on Xorg”, log in again.
37.4 Scoring
LeRobot doesn’t know whether the task succeeded. You decide, per episode. Keep a tally in the day’s log:
| Checkpoint | Episode | Result | Notes |
|---|---|---|---|
last (100k) | 0 | ✅ | |
last (100k) | 1 | ❌ | missed the grasp by ~2 cm |
| … |
With 10 episodes per checkpoint you can compare checkpoints. Point --policy.path at …/checkpoints/040000/pretrained_model, …/060000/…, etc. The last checkpoint isn’t always the best (overfitting).