Part III: Running a trained policy · §30 of 43
30.1 The idea in one paragraph
A policy is a neural network that has learned, from your demonstrations, “when the robot looks like this, move like that”. Running it (“rollout”, “inference” or “deployment”; all three words mean the same here) means a program repeats a short loop about 30 times per second: read the arms’ joint angles and the camera images, give them to the network, take the joint targets it answers with, and send those targets to the motors. Nobody holds the Minis. The network is the operator. In LeRobot this program is the command lerobot-rollout.
It is exactly the teleoperation loop of §23 with the Minis replaced by the network:
| Teleoperation (§23) | Policy rollout (Part III) | |
|---|---|---|
| Who decides the next joint targets | You, through the Minis | The trained network |
| Where the targets come from | teleop.get_action() (Mini angles) | policy.select_action(observation) |
| What the follower does with them | Same: clipped to the side limits and max_relative_target, then sent as MIT commands at kp 240 | Same |
| What it needs to see | Nothing (you look at the scene yourself) | Joint angles and camera images, exactly as during recording |
30.2 Words you will meet
| Word | Meaning on zeus |
|---|---|
| Policy | The trained network plus its settings. For a first policy this is ACT ([[guide/part-2/24-recording-training-and-running-policies#24.4 Which policy to train |
| Checkpoint | A snapshot of the policy saved during training, e.g. after 20 000 and 40 000 training steps. Each one is a folder. The newest is reachable through the last link. |
pretrained_model folder | The folder inside a checkpoint that holds everything needed to run the policy. This is what --policy.path points to ([[guide/part-3/31-what-you-need-before-the-first-run#31.2 Find and check the trained model |
| Observation | What the policy sees at one instant: 16 joint angles in degrees (7 joints + gripper, per arm), named left_joint_1.pos … right_gripper.pos, plus one image per camera. |
| Action | What the policy answers: 16 joint targets in degrees, same names and order. They are absolute positions (“go to 37°”), not changes (“move 2°”). |
| Action chunk | ACT doesn’t predict one action, it predicts the next 100 at once (chunk_size). The loop then plays them back one per tick. |
n_action_steps | How many actions of each chunk are actually played before the network is asked again. Default 100 for ACT, i.e. the network looks at the scene once every 100 / 30 ≈ 3.3 s and moves “blind” in between. See [[guide/part-3/36-tuning-how-the-policy-moves#36.1 How often the policy looks: n_action_steps |
| fps / tick | One pass of the loop is a tick. --fps=30 means 30 ticks per second, one action per tick. It must match the fps of the recorded dataset. |
| Normalisation stats | The mean and spread of every joint and camera channel in the training data. The network works on normalised numbers, so the stats are saved with the checkpoint and applied automatically. You never touch them, but they’re why a policy trained on one dataset can’t simply be reused on a differently set-up robot. |
| Task | A sentence describing what to do, e.g. “Pick up the red cube and place it in the bowl”. ACT and Diffusion ignore it (they only know the one task they were trained on). Language models (SmolVLA, Pi0) use it, so it must match the training text. |
| Start pose / initial position | The joint angles the arms have when lerobot-rollout connects. LeRobot records them, and when the run ends it moves the arms back there over 3 s before switching torque off. |
| Strategy | What the run does around the policy: base (just run it), episodic (run repeated evaluation episodes and record them), dagger (let you take over with the Minis and record corrections), highlight, sentry. |
| Inference engine | How the policy is called: sync (inline, the default, right for ACT) or rtc (for slow VLA models). |
| Device | Where the network runs: cuda = the RTX 5050 GPU (use this), cpu = processor (slow). |
| Episode | One attempt at the task, from a reset scene to the end. Datasets are made of episodes. |
30.3 The loop, drawn
The rollout loop. Every tick takes one action from the queue; when the queue runs out, the network looks at the scene again and refills it.
The same diagram as text, as written in the guide
lerobot-rollout ├─ 1. load the policy (pretrained_model/ → GPU) no hardware touched yet ├─ 2. connect the robot: open can0/can1, cameras, TORQUE ON ├─ 3. remember the current joint angles = "initial position" │ ├─ 4. loop, 30 times per second, until --duration or Ctrl+C: │ observe ← 16 joint angles (deg) + camera images │ policy → if its action queue is empty: run the network once, │ get a chunk of 100 future actions, queue n_action_steps of them │ → pop the next action (16 joint targets, deg) │ safety → clip each target to the `side` joint limits (§20) │ → clip each step to ±max_relative_target from where the joint is now │ send → MIT command per motor (kp 240 on the big joints) │ ├─ 5. move back to the initial position in a straight line, over 3 s └─ 6. TORQUE OFF (the arms go limp) and disconnect
30.4 Why the setup has to match the training
The network only knows the world it saw in the demonstrations. It has no idea what a “cube” is. It learned which pixels and joint angles came before which motions. So anything that makes today’s observation look different from the recorded ones makes the policy behave worse, sometimes wildly:
| Must be identical to the recording | Why | What happens if not |
|---|---|---|
Robot type (bi_openarm_follower), cabling (can0 = left, can1 = right), side | Joint i of the observation must be the same physical joint | Wrong arm or mirrored joints receive the commands |
Camera names (top, …) | The policy looks up each image by name | Start-up error (Visual feature mismatch) |
| Camera resolution and fps | The network’s input size is fixed | Start-up error or garbage |
| Camera position and angle | Pixels must mean the same thing | Policy “does nothing sensible” (most common cause of a policy that worked yesterday failing today) |
| Which physical camera has which name | top must still be the top camera | Silently wrong behaviour |
Loop --fps | Each action was learned as “1/30 s later” | Motions too fast or too slow |
use_velocity_and_torque (default off) | Changes the size of the observation | Start-up error |
| Task text (VLAs only) | Language-conditioned policies read it | Wrong or no behaviour |
| Lighting, background, objects | Same pixels | Degrades gradually; train with some variety to make it robust |