Part III: Running a trained policy · §40 of 43
For flow-matching VLAs (SmolVLA, Pi0, Pi0.5) one inference takes longer than one control tick. Real-Time Chunking (RTC) computes the next chunk in the background while the current one plays, and blends them so the motion doesn’t stall:
lerobot-rollout --strategy.type=base --inference.type=rtc \
--inference.rtc.execution_horizon=10 --inference.rtc.max_guidance_weight=10.0 \
--policy.path=<hf_user>/smolvla_openarm_pick_cube \
… same --robot.* flags as §32.1 … \
--task="Pick up the red cube and place it in the bowl" --device=cudaLeRobot refuses RTC for policies that don’t support it (ACT, Diffusion: RTC inference is not supported by policy type …); use the default sync for those.
If the model doesn’t fit on the laptop GPU, LeRobot’s async inference runs the policy on a GPU server and only the robot client on zeus (pip install -e ".[async]" on both machines):
# on the GPU server
python -m lerobot.async_inference.policy_server --host=0.0.0.0 --port=8080
# on zeus
python -m lerobot.async_inference.robot_client --server_address=<server_ip>:8080 \
--robot.type=bi_openarm_follower … --policy_type=pi05 --pretrained_name_or_path=<hf_user>/pi05_openarm \
--actions_per_chunk=50 --chunk_size_threshold=0.5 --task="…"⚠ The async client’s official robot list doesn’t include the OpenArm yet (SUPPORTED_ROBOTS in async_inference/constants.py). The check is commented out, so it should run, but it’s untested here. The async client is a separate program from lerobot-rollout: no strategies, no return-to-start-pose. See docs/source/async.mdx for the full flag list.