Part III: Running a trained policy · §40 of 43

For flow-matching VLAs (SmolVLA, Pi0, Pi0.5) one inference takes longer than one control tick. Real-Time Chunking (RTC) computes the next chunk in the background while the current one plays, and blends them so the motion doesn’t stall:

lerobot-rollout --strategy.type=base --inference.type=rtc \
  --inference.rtc.execution_horizon=10 --inference.rtc.max_guidance_weight=10.0 \
  --policy.path=<hf_user>/smolvla_openarm_pick_cube \
  … same --robot.* flags as §32.1 … \
  --task="Pick up the red cube and place it in the bowl" --device=cuda

LeRobot refuses RTC for policies that don’t support it (ACT, Diffusion: RTC inference is not supported by policy type …); use the default sync for those.

If the model doesn’t fit on the laptop GPU, LeRobot’s async inference runs the policy on a GPU server and only the robot client on zeus (pip install -e ".[async]" on both machines):

# on the GPU server
python -m lerobot.async_inference.policy_server --host=0.0.0.0 --port=8080
# on zeus
python -m lerobot.async_inference.robot_client --server_address=<server_ip>:8080 \
  --robot.type=bi_openarm_follower … --policy_type=pi05 --pretrained_name_or_path=<hf_user>/pi05_openarm \
  --actions_per_chunk=50 --chunk_size_threshold=0.5 --task="…"

⚠ The async client’s official robot list doesn’t include the OpenArm yet (SUPPORTED_ROBOTS in async_inference/constants.py). The check is commented out, so it should run, but it’s untested here. The async client is a separate program from lerobot-rollout: no strategies, no return-to-start-pose. See docs/source/async.mdx for the full flag list.