Part III: Running a trained policy · §37 of 43

37.1 What it does

--strategy.type=episodic runs the policy as a series of episodes with a reset phase between them, and records everything as a LeRobot dataset, like lerobot-record but with the policy instead of you. Use it to measure a success rate, compare checkpoints, or keep good rollouts as extra training data.

For each episode:

  1. Policy state is cleared; the policy runs for up to --dataset.episode_time_s seconds (→ ends it early).
  2. Episode saved.
  3. Reset phase of --dataset.reset_time_s seconds: the arms go back to the start pose (in 1 s, faster than the 3 s at the end of a run) and hold there, torque on, while you reset the objects. With the Minis attached you teleoperate the reset instead (§38).
  4. Next episode, until --dataset.num_episodes are recorded or Esc.

37.2 The command

lerobot-rollout \
  --strategy.type=episodic \
  --policy.path=$HOME/Documents/lerobot/outputs/train/act_openarm_pick_cube/checkpoints/last/pretrained_model \
  --robot.type=bi_openarm_follower \
  --robot.id=my_bimanual_follower \
  --robot.left_arm_config.port=can0  --robot.left_arm_config.side=left  --robot.left_arm_config.max_relative_target=10.0 \
  --robot.right_arm_config.port=can1 --robot.right_arm_config.side=right --robot.right_arm_config.max_relative_target=10.0 \
  --robot.cameras='{top: {type: opencv, index_or_path: /dev/video0, width: 640, height: 480, fps: 30}}' \
  --dataset.repo_id=<hf_user>/rollout_act_pick_cube_last \
  --dataset.single_task="Pick up the red cube and place it in the bowl" \
  --dataset.num_episodes=10 \
  --dataset.episode_time_s=40 \
  --dataset.reset_time_s=20 \
  --dataset.fps=30 \
  --dataset.push_to_hub=false \
  --device=cuda

Rules for the dataset flags, all enforced by LeRobot v0.6.2:

RuleWhy / error if broken
The dataset name must start with rollout_Dataset names for rollout must start with 'rollout_'. (V2.2 of this guide used eval_… and hil_…, which are rejected.)
A date-time tag is appended to the name, e.g. rollout_act_pick_cube_last_20261010_143000Each session gets a new dataset. Add --dataset.no_stamp=true to keep the exact name (a second run with the same name then needs --resume=true, or fails because the dataset exists ⚠ verify the exact error).
Always set --dataset.push_to_hub=falseThe default is true: at the end it tries to upload to the Hub and fails without hf auth login.
Set num_episodes, episode_time_s, reset_time_sDefaults are 50 episodes of 60 s with 60 s resets.
--dataset.single_taskBecomes the task text (no separate --task needed).
--dataset.* flags with --strategy.type=baseRejected: base strategy does not record data.

The recording goes to ~/.cache/huggingface/lerobot/<hf_user>/rollout_…/. View it with lerobot-dataset-viz --repo-id <the stamped name> --episode-index 0.

37.3 Keyboard controls and the Wayland problem on zeus

KeyEffect
→End the current episode (or reset phase) now
←Discard the current episode and record it again
EscStop the session (return to start pose, torque off)

These keys are read by pynput, which needs an X11 session. zeus currently logs in to a Wayland session (echo $XDG_SESSION_TYPE → wayland), where the listener most likely receives nothing ⚠ verify. Without the keys the session still works: episodes and resets simply run for their full time, and Ctrl+C once ends the session (an episode in progress at that moment may be lost). To get the keys: log out, click the gear icon on the login screen, choose “Ubuntu on Xorg”, log in again.

37.4 Scoring

LeRobot doesn’t know whether the task succeeded. You decide, per episode. Keep a tally in the day’s log:

CheckpointEpisodeResultNotes
last (100k)0✅
last (100k)1❌missed the grasp by ~2 cm
…

With 10 episodes per checkpoint you can compare checkpoints. Point --policy.path at …/checkpoints/040000/pretrained_model, …/060000/…, etc. The last checkpoint isn’t always the best (overfitting).