Part III: Running a trained policy · §30 of 43

30.1 The idea in one paragraph

A policy is a neural network that has learned, from your demonstrations, “when the robot looks like this, move like that”. Running it (“rollout”, “inference” or “deployment”; all three words mean the same here) means a program repeats a short loop about 30 times per second: read the arms’ joint angles and the camera images, give them to the network, take the joint targets it answers with, and send those targets to the motors. Nobody holds the Minis. The network is the operator. In LeRobot this program is the command lerobot-rollout.

It is exactly the teleoperation loop of §23 with the Minis replaced by the network:

Teleoperation (§23)Policy rollout (Part III)
Who decides the next joint targetsYou, through the MinisThe trained network
Where the targets come fromteleop.get_action() (Mini angles)policy.select_action(observation)
What the follower does with themSame: clipped to the side limits and max_relative_target, then sent as MIT commands at kp 240Same
What it needs to seeNothing (you look at the scene yourself)Joint angles and camera images, exactly as during recording

30.2 Words you will meet

WordMeaning on zeus
PolicyThe trained network plus its settings. For a first policy this is ACT ([[guide/part-2/24-recording-training-and-running-policies#24.4 Which policy to train
CheckpointA snapshot of the policy saved during training, e.g. after 20 000 and 40 000 training steps. Each one is a folder. The newest is reachable through the last link.
pretrained_model folderThe folder inside a checkpoint that holds everything needed to run the policy. This is what --policy.path points to ([[guide/part-3/31-what-you-need-before-the-first-run#31.2 Find and check the trained model
ObservationWhat the policy sees at one instant: 16 joint angles in degrees (7 joints + gripper, per arm), named left_joint_1.pos … right_gripper.pos, plus one image per camera.
ActionWhat the policy answers: 16 joint targets in degrees, same names and order. They are absolute positions (“go to 37°”), not changes (“move 2°”).
Action chunkACT doesn’t predict one action, it predicts the next 100 at once (chunk_size). The loop then plays them back one per tick.
n_action_stepsHow many actions of each chunk are actually played before the network is asked again. Default 100 for ACT, i.e. the network looks at the scene once every 100 / 30 ≈ 3.3 s and moves “blind” in between. See [[guide/part-3/36-tuning-how-the-policy-moves#36.1 How often the policy looks: n_action_steps
fps / tickOne pass of the loop is a tick. --fps=30 means 30 ticks per second, one action per tick. It must match the fps of the recorded dataset.
Normalisation statsThe mean and spread of every joint and camera channel in the training data. The network works on normalised numbers, so the stats are saved with the checkpoint and applied automatically. You never touch them, but they’re why a policy trained on one dataset can’t simply be reused on a differently set-up robot.
TaskA sentence describing what to do, e.g. “Pick up the red cube and place it in the bowl”. ACT and Diffusion ignore it (they only know the one task they were trained on). Language models (SmolVLA, Pi0) use it, so it must match the training text.
Start pose / initial positionThe joint angles the arms have when lerobot-rollout connects. LeRobot records them, and when the run ends it moves the arms back there over 3 s before switching torque off.
StrategyWhat the run does around the policy: base (just run it), episodic (run repeated evaluation episodes and record them), dagger (let you take over with the Minis and record corrections), highlight, sentry.
Inference engineHow the policy is called: sync (inline, the default, right for ACT) or rtc (for slow VLA models).
DeviceWhere the network runs: cuda = the RTX 5050 GPU (use this), cpu = processor (slow).
EpisodeOne attempt at the task, from a reset scene to the end. Datasets are made of episodes.

30.3 The loop, drawn

The rollout loop. Every tick takes one action from the queue; when the queue runs out, the network looks at the scene again and refills it.

observe16 joint angles + camera imagespolicypop the next actionsafetyclip to limits and max_relative_targetsendMIT, kp 2401 tick= 1/30 s on zeusaction queueone action played per tick, oldest firstACT network (GPU)runs once to refill the queue with a new chunkSlowed down here, with a 10-action chunk.On zeus: 30 ticks/s, chunks of 100 actions,n_action_steps 100: the network looks every ~3.3 s.

30.4 Why the setup has to match the training

The network only knows the world it saw in the demonstrations. It has no idea what a “cube” is. It learned which pixels and joint angles came before which motions. So anything that makes today’s observation look different from the recorded ones makes the policy behave worse, sometimes wildly:

Must be identical to the recordingWhyWhat happens if not
Robot type (bi_openarm_follower), cabling (can0 = left, can1 = right), sideJoint i of the observation must be the same physical jointWrong arm or mirrored joints receive the commands
Camera names (top, …)The policy looks up each image by nameStart-up error (Visual feature mismatch)
Camera resolution and fpsThe network’s input size is fixedStart-up error or garbage
Camera position and anglePixels must mean the same thingPolicy “does nothing sensible” (most common cause of a policy that worked yesterday failing today)
Which physical camera has which nametop must still be the top cameraSilently wrong behaviour
Loop --fpsEach action was learned as “1/30 s later”Motions too fast or too slow
use_velocity_and_torque (default off)Changes the size of the observationStart-up error
Task text (VLAs only)Language-conditioned policies read itWrong or no behaviour
Lighting, background, objectsSame pixelsDegrades gradually; train with some variety to make it robust