Part III: Running a trained policy · §41 of 43
Start-up errors (the policy never takes control). Errors raised after Robot connected (camera or feature mismatches) end the program without the normal shutdown, so torque may stay on ⚠ verify. If the arms are still stiff afterwards: support them, then openarm-can-cli -i can0 disable and openarm-can-cli -i can1 disable (the arms go limp).
| Message / symptom | Cause / fix |
|---|---|
lerobot-rollout: command not found | conda activate lerobot. |
--policy.path is required for rollout | Flag missing, or the line with it was cut off by a broken \ (§32.4). |
| Hub “repository not found” / “repo id must be in the form …” for a local model | The path doesn’t exist, so LeRobot treated it as a Hub name. Check it with ls <path>/config.json; use an absolute path ending in pretrained_model. |
Visual feature mismatch between policy and robot hardware. Policy expects: {…} Robot provides: {…} | Camera names differ from training. Use the recording’s exact --robot.cameras, or --rename_map. |
| Error about missing / unexpected features, or a size mismatch (e.g. 16 vs 48) | use_velocity_and_torque or the robot type differs from training, or a bimanual policy on a single arm. |
observation.state order … is a permutation of the action dispatch order | Joint order of the robot doesn’t match the checkpoint. Shouldn’t happen with the standard follower; check the policy was trained on a bi_openarm_follower dataset. |
| LeRobot asks to put the follower “hanging straight down” | --robot.id missing or different from my_bimanual_follower. Ctrl+C; don’t press Enter (§21). |
Failed to connect to CAN bus: … Network is down | PCAN adapter replugged. Re-run setup_openarm_lerobot.sh. |
| Motors “No response” | Motor power off, or another program (ROS, teleop) owns the bus. §27. |
CUDA out of memory | Model too big for 8 GB: smaller policy, --device=cpu (slow), or async inference on a bigger GPU (§40). |
Dataset names for rollout must start with 'rollout_' | Recording strategies need --dataset.repo_id=<hf_user>/rollout_<name>. |
base strategy does not record data: drop the --dataset.* flags … | Remove the --dataset.* flags, or use episodic. |
episodic strategy requires --dataset.repo_id to be set (same for dagger, highlight, sentry) | Add --dataset.repo_id=<hf_user>/rollout_<name>. |
dagger strategy requires --teleop.type to be set | Add the --teleop.* flags of §38.3. |
--interactive=true supports --strategy.type=base or sentry | Interactive mode only works with those two strategies. |
RTC inference is not supported by policy type 'act' | Drop --inference.type=rtc. |
`n_action_steps` must be 1 when using temporal ensembling | Add --policy.n_action_steps=1. |
During the run:
| Symptom | Cause / fix |
|---|---|
| Arms snap at the start | Not in the training start pose (§34 step 5), or max_relative_target not set. |
| Policy moves but “does nothing sensible” | Cameras moved, lighting changed, cameras swapped, wrong task text (VLAs), or simply too few / inconsistent demos. Replay a training episode with lerobot-replay (§24.2) to check the setup itself. |
| Policy works at the start, then drifts or freezes | Situation not covered by the demos: collect DAgger corrections (§38) or more demos. Also try a lower n_action_steps (§36.1). |
| A small jerk every ~3 s | ACT chunk boundaries. Temporal ensembling (§36.2). |
| Jerky motion all the time | Loop slower than --fps (cadence summary), Rerun on, or a slow VLA without RTC. Try --interpolation_multiplier=2. |
| One joint creeps or stalls, constant “had to be clamped” warnings for it | max_relative_target too small for that low-gain joint; raise it a little. |
| Gripper doesn’t close fully | Gripper targets are clipped to −65…0°. The policy learned the Mini’s gripper range; check the Mini gripper calibration used during recording (§23.2). |
| The arms go limp at the end | Expected: torque off after the return to the start pose. Support them. |
| Nothing printed in interactive mode | Expected: routine logs are muted during an interactive session; only errors and cadence summaries appear. |
| The computer talks | --play_sounds (default true) reads events aloud. --play_sounds=false to silence. |
| → / ← / Esc / Space / Tab do nothing | Keyboard listener needs X11; zeus is on Wayland (§37.3). |
| DAgger: Minis move on their own when pausing | Intended: the smooth handover drives the Minis to the follower pose. Hands off until they stop. |
| DAgger: follower jumps when a correction starts | Mini calibration wrong or smooth_handover disabled. Run check_mini_calibration.py; recalibrate (§23.2). |
Training / dataset names:
| Symptom | Cause / fix |
|---|---|
lerobot-train can’t find <hf_user>/openarm_pick_cube | lerobot-record appended a date-time tag to the name (§24.1). ls ~/.cache/huggingface/lerobot/<hf_user>/ shows the real name. Use --dataset.no_stamp=true when recording to avoid this. |
Dataset names starting with 'eval_' are reserved for policy evaluation | lerobot-record refuses eval_… names. Evaluate with lerobot-rollout --strategy.type=episodic (§37). |