← Back to the Log Index · Previous day: (none, this is the first log) · Next day: 2026-10-10

Summary of the day

  • We added an explanation of the docker run / docker start / docker exec commands to the OpenArm guide, which became a new minor version, Guide V2.1.
  • We agreed a versioning convention for the guides (see 1. Documentation changes).
  • We started a step-by-step test of the ROS 2 setup, beginning with the environment (host, Docker container, GPU, CAN).
  • The container and workspace are fine. The test was blocked because the laptop’s NVIDIA GPU crashed (driver error Xid 154, GPU Reset Required) after a power-source change, and only a reboot fixes that. See 3. Issue: NVIDIA GPU crashed (Xid 154).
  • Afternoon: after a reboot (15:59) the GPU works again, and test step 1 (GPU and display) passed: RViz runs on the NVIDIA card. See 3.6 Resolution (after the 15:59 reboot) and 5.1 Test step 1: GPU and display.
  • Test step 2 (simulation bringup) passed, run by Manuel himself. A first attempt ran two launches at once by mistake; see 5.2 Test step 2: simulation bringup.
  • New working rule: Manuel runs every command himself; Claude explains, gives the commands and checks results for safety.
  • Evening, hardware stage: CAN checked (PCAN-USB Pro FD, 8 motors per arm), right joint 5 zero fixed (was off by 80°), motor driver patched (slow 10 s start-up, no kick, refuses to move if a motor doesn’t reply or a joint is >0.5 rad from zero), new container with real-time scheduling, first real launch of both arms and first slow moves (left joint1 and joint7, matching RViz). Found that Guide §3’s joint limits are wrong for joint1 (left) and joint2 (both); joint direction map worked out. See 5.4 Hardware stage: plan, and H1 host checks → 5.9 H6: first commanded moves, and the joint direction map.
  • Where to pick up tomorrow: 7. Next session.
  • Late evening: a new major guide version, Guide V3 (a major number because Part III is a large addition), adds Part III, a beginner-level, step-by-step guide to running trained policies with LeRobot (lerobot-rollout). It was written from the installed LeRobot source and hasn’t been tested on the robot yet, since no policy has been trained. While checking the source, three problems with V2.2 came up. Rollout datasets must be named rollout_…, so V2.2’s eval_…/hil_… examples would fail. lerobot-record appends a date-time tag to dataset names unless --dataset.no_stamp=true. And zeus runs a Wayland session, where LeRobot’s keyboard controls probably don’t work.
  • Test step 3 found a safety problem: the trajectory controller doesn’t move smoothly to a target, it waits for time_from_start and then jumps (interpolation_method: none). The guide said the opposite. Corrected in Guide V2.2. Manuel switched the controllers to splines and confirmed smooth motion in sim: goals now reach their target in exactly time_from_start. See 5.3 Trajectory controller snaps instead of moving smoothly (issue).

Contents of this log


1. Documentation changes

log/documentation

1.1 What changed

The ROS 2 guide previously only showed the three Docker commands used every day without explaining how they differ. A new subsection now explains them, in §5 Environment and container of Guide V2.1. In short:

CommandWhat it acts onWhat it does
docker runan image (a frozen template, here openarm-jazzy:built)Creates a brand-new container and starts it. All the settings (network, GPU, display) are fixed at this moment. Only ever run it once for a given container.
docker startan existing stopped containerRestarts the container, with everything previously installed or built inside it still there.
docker execan existing running containerOpens an extra process (usually a bash shell) inside it. Use it once for every terminal you need.

1.2 Versioning convention (agreed today)

Rule: never overwrite a guide version's content

Small additions or corrections to a guide’s content are saved as a new file with the minor version bumped, for example …_V2.md → …_V2.1.md → …_V2.2.md. The previous file’s content stays exactly as it was. Major rewrites get a new major number (V3).

Exception: changes that only affect formatting or navigation, such as converting links to Obsidian format or adding navigation links, are made in place in every existing version, without a new version number. They don’t change what the guide says. See 1.4 Obsidian link conversion (all versions, in place).

Why: the guides are shared, so people may be reading or referring to an older version. Keeping each version as its own file means nothing they rely on changes without warning, and the history stays visible.

Mistake made today (by Claude), already corrected

The first edit went directly into V2, overwriting it. That was undone: the changes were moved into a new V2.1 file and V2 was restored to its original content (105 947 bytes at that point, before the link conversion in §1.4). This is why the convention above was written down.

1.3 Files at end of day

All guides now live in Prandium/Documents/OpenArm Docs/. They were moved there during the session; earlier they sat directly in Prandium/Documents/.

Why: the guides were written with GitHub-style links, such as [Environment and container](#5-environment-and-container). Those work on GitHub but not in Obsidian, which uses its own link format and doesn’t recognise GitHub’s automatically generated “anchor” names. Cross-references in the text were also written as plain §12 with no link, so you had to scroll to find them.

What was done (to V1, V2 and V2.1, edited in place, see the exception in 1.2 Versioning convention (agreed today)):

ChangeBeforeAfterCount (V1 / V2 / V2.1)
Table-of-contents links[Big picture](#1-big-picture)[[#1. Big picture|Big picture]]29 / 29 / 29
§ cross-references in the textsee §12see [[#12. The physical robot: bringing it up with ROS 2|§12]]40 / 69 / 69
Navigation box under the title(none)Links to the other guide versions and to the Log Index1 / 1 / 1

What it looks like when reading: the §12 text looks the same but is now clickable. Hovering shows a preview of that section if Obsidian’s built-in Page preview core plugin is enabled (it is by default; hold Ctrl while hovering).

How it was done (so it can be repeated on future guides): a small Python script read every heading in the file, matched each old GitHub anchor and each §N / §N.M reference to the real heading text, and rewrote it as an Obsidian heading link. Notes on the details:

  • Text inside code blocks and inline code was left alone, so commands that contain § or [[ (e.g. the Python list [[0, 0, …]] in §8.3) weren’t changed.
  • Inside tables, the | separating a link’s target from its display text has to be written \|, otherwise Markdown treats it as a new table column. The script did this automatically.
  • A range such as §24.4–§24.6 became two separate links.
  • Afterwards, a check confirmed that every link points to a heading that exists, that every §N.M points to section N.M and not just to N, and that no GitHub-style links are left. Two table-of-contents entries whose titles contain inline code (§16, §28) were missed by the script and fixed by hand.
  • A backup of the files from before the conversion was kept in Claude’s temporary session folder. It is not permanent.

Tip: linking to a heading in Obsidian

Inside the same file: [[#Heading text]]. In another file: [[File name#Heading text]]. To show different text: [[#Heading text|shown text]]. Typing [[# in Obsidian opens an autocomplete list of the headings.


2. Test step 0: environment check

log/testing topic/docker

Goal of this step: before launching any ROS nodes, confirm that what ROS runs on is correct: the host machine, the Docker container and its settings, the graphics card (needed by RViz, the 3-D visualiser) and the CAN interfaces (needed later for the real robot). If one of these is wrong, the ROS tests fail in confusing ways, so it’s quicker to check them first.

2.1 Results

CheckCommand usedResultOK?
Display variable on hostecho $DISPLAY:1 (Wayland session with XWayland)✅
User can run Docker without sudoid -nGuser is in the docker group✅
Image existsdocker imagesopenarm-jazzy:built (7.36 GB) and an older openarm-jazzy:latest✅
Container exists and is runningdocker ps -a --filter name=openarmUp 24 hours✅
Container was created with the right flagsdocker inspect openarmnetwork = host, IPC = host, GPUs = all, DISPLAY=:1, __NV_PRIME_RENDER_OFFLOAD=1, __GLX_VENDOR_LIBRARY_NAME=nvidia, NVIDIA_DRIVER_CAPABILITIES=all✅
Shell setup inside containergrep source /root/.bashrcsources /opt/ros/jazzy/setup.bash and ~/ros2_ws/install/setup.bash; no leftover humble lines✅
Workspace sources presentls /root/ros2_ws/srcopenarm_can, openarm_description, openarm_ros2✅
Workspace builtls /root/ros2_ws/installopenarm, openarm_bimanual_moveit_config, openarm_bringup, openarm_can, openarm_description, openarm_hardware✅
NVIDIA GPU usable (host)nvidia-smiUnable to determine the device handle for GPU0: 0000:01:00.0: Unknown Error❌
NVIDIA GPU usable (container)docker exec openarm nvidia-smi -Lsame error❌
CAN interfaces existip -br linkcan0 and can1 present, both DOWN✅ (down is expected, we aren’t using hardware yet)

You no longer need to source by hand in each shell

The guide tells you to run the two source lines in every new shell. On zeus, /root/.bashrc inside the container already does this, so every docker exec -it openarm bash shell is ready to use. You only need to source by hand again after rebuilding the workspace (source /root/ros2_ws/install/setup.bash) in a shell that was already open.


3. Issue: NVIDIA GPU crashed (Xid 154)

issue/gpu status/resolved

3.1 Symptom

nvidia-smi (the NVIDIA status tool) fails on the host and in the container:

Unable to determine the device handle for GPU0: 0000:01:00.0: Unknown Error
No devices were found

3.2 Why this matters for ROS

zeus is a hybrid-graphics laptop: it has an Intel integrated GPU and an NVIDIA RTX 5050. The container is deliberately set up to force OpenGL programs onto the NVIDIA card (the __NV_PRIME_RENDER_OFFLOAD=1 and __GLX_VENDOR_LIBRARY_NAME=nvidia variables), because RViz failed on the Intel/Mesa driver in earlier attempts (see Guide §15). With the NVIDIA card in an error state, RViz will fail to open, and so will both launch files in the guide, since both start RViz. Nodes without a GUI (controllers, ros2 topic, etc.) don’t need the GPU.

3.3 Diagnosis

The kernel log (journalctl -k -b | grep -i nvrm) shows what happened, in this order:

14:36:46  NVRM: rm_power_source_change_event: Failed to handle Power Source change event, status=0xf
14:36:46  NVRM: Xid (PCI:0000:01:00): 154, GPU recovery action changed from 0x0 (None) to 0x1 (GPU Reset Required)
15:08:30  NVRM: rm_power_source_change_event: … Failed …  (repeats several times)
15:24:42  WARNING: nvidia/nv.c:5388 at nvidia_dev_put … (driver stack trace)

What this means:

  • A “power source change” is the laptop switching between mains power and battery, i.e. the charger being plugged in or unplugged (or a loose charger connection).
  • The NVIDIA driver (version 595.91.07, the “open” kernel module) failed to handle that switch and marked the GPU as needing a reset. Xid codes are NVIDIA’s numbered error categories, and 154 means “GPU recovery action required”. Here the action is GPU Reset Required.
  • The kernel module itself is still loaded (lsmod shows nvidia, nvidia_drm, etc.), and the PCI device reports as active, so the hardware is still there. The driver just won’t talk to it until it’s reset.

3.4 Fix

Fix: reboot the laptop

On a laptop GPU, a full reboot is the reliable way to perform the “GPU reset” the driver asks for. Afterwards:

  1. On the host, run nvidia-smi. It should list the RTX 5050 with no errors.
  2. Run xhost +SI:localuser:root (needed after every login, so the container may open windows).
  3. Run docker start openarm. The container stops during a reboot, which is normal.
  4. Run docker exec openarm nvidia-smi -L to confirm the container sees the GPU too.

3.5 Prevention

Pitfall: don't plug or unplug the charger while working with the GPU

On this machine and driver version, switching between battery and mains power while the NVIDIA GPU is in use can crash the GPU driver until the next reboot. Plug the charger in before starting and leave it in. If the GPU does crash: nvidia-smi errors → reboot.

If this keeps happening even with the charger left alone, it’s worth checking for a newer NVIDIA driver, or trying the proprietary (closed) kernel module instead of the open one. Not done yet.

Status: open resolved, see below.

3.6 Resolution (after the 15:59 reboot)

GPU working again

The laptop was rebooted at 15:59 with the charger plugged in. Checks at 16:08:

CheckResult
nvidia-smi on the hostRTX 5050 listed, no errors, 12 MiB used
Kernel log since boot (journalctl -k -b | grep -i xid)no Xid or power-source errors
Charger (/sys/class/power_supply/AC*/online)1, plugged in
xhostSI:localuser:root already allowed
ContainerUp, already started after the reboot
docker exec openarm nvidia-smi -LGPU 0: NVIDIA GeForce RTX 5050 Laptop GPU

4. Observations to follow up later

log/follow-up

4.1 The CAN adapters are PEAK, not CANable/gs_usb

topic/can topic/hardware

The kernel module list includes peak_usb, the driver for PEAK-System USB-CAN adapters (e.g. PCAN-USB FD). Guide §10.1 assumes gs_usb-class adapters (CANable 2.0 / candleLight). Both kinds work with Linux SocketCAN (the standard Linux CAN interface that can0 and can1 are part of), so the ROS side doesn’t care. But:

  • The bit-rate setup in §10.4 may need different options for PEAK adapters, especially for CAN-FD data bit-rate and sample point.
  • The udev naming rules in §10.3 (which make sure the same physical adapter always gets the same name, can0 or can1) match on USB IDs, which differ between vendors.

To do before the hardware tests: confirm the adapter model (lsusb), then check and adjust §10.3–10.4 in a new guide version.

Update (evening): Guide §22 already describes zeus’s adapter as a single PEAK PCAN-USB Pro FD with two channels, giving stable can0/can1 names, and says the cables are the reverse of the ROS defaults (left arm = can0). So the udev naming in §10.3 isn’t needed on zeus; the §10.4 bit-rate options for PEAK still need checking. The §12.1 checklist contradicted §22 on the left/right mapping; corrected in V2.2.

4.2 Two Docker images exist

There are two images, openarm-jazzy:latest (6.87 GB) and openarm-jazzy:built (7.36 GB). The container uses :built. Most likely :latest is the base image from before the workspace was built and :built was committed afterwards, but this hasn’t been verified. Don’t delete either until we know.

4.3 Extra line in .bashrc

/root/.bashrc line 101 is [ -f /ws/install/setup.bash ] && source /ws/install/setup.bash, a leftover from an earlier setup that used /ws as the workspace. It only runs if that file exists, so it’s harmless, but it could cause confusion if a /ws folder ever appears. It can be removed when convenient.

4.4 Container not created with real-time or CAN-admin flags

Resolved in the evening: see 5.6 H4: new container with real-time permissions.

docker inspect shows CapAdd=[], so the container does not have --cap-add SYS_NICE, --ulimit rtprio=99 or --cap-add NET_ADMIN. That’s fine for simulation. For the real robot, Guide §5 recommends the real-time flags at 750 Hz. Since flags can only be set at docker run time, this will mean creating a new container (with docker commit first, so nothing installed inside is lost). That decision is for the hardware stage.


5. Test plan and where we stopped

log/testing

flowchart TD
    S0["0. Environment check<br/>(host, container, CAN)"]:::done --> G["GPU working?"]:::done
    G -->|"after reboot"| S1["1. GPU and display test<br/>nvidia-smi + OpenGL window"]:::done
    S1 --> S2["2. Sim bringup (Guide §7A)<br/>controllers active, v1.0 model, RViz"]:::done
    S2 --> S3["3. Commanding (Guide §8.1–8.4)<br/>each arm and gripper, check /joint_states"]:::done
    S3 --> S3b["3b. Switch JTC to splines<br/>and test smooth motion in sim"]:::done
    S3b --> S4["4. Controller switching (Guide §8.5–8.6)"]
    S4 --> S5["5. MoveIt (Guide §9)<br/>incl. gripper caveat"]
    S5 --> S6["6. Hardware (Guide §10–12)<br/>CAN checks first, motors unpowered"]
    classDef done fill:#2e7d32,color:#fff
    classDef blocked fill:#c62828,color:#fff

5.1 Test step 1: GPU and display

Goal: prove that a program inside the container can open an OpenGL window on the host’s screen and that it renders on the NVIDIA card, not on Intel/Mesa. This is the precondition for both launch files, which start RViz.

Safety check first: ip -br link showed can0 and can1 both DOWN, and no ROS processes were running in the container. Nothing could reach the motors.

How: glxgears/glxinfo aren’t installed in the container (mesa-utils is missing), so instead of installing anything we used RViz on its own (no robot, no controllers) as the OpenGL test, for 20 seconds:

docker exec -d openarm bash -c 'source /opt/ros/jazzy/setup.bash && timeout 20 rviz2 > /tmp/rviz_test.log 2>&1'
nvidia-smi        # on the host, while RViz is open
CheckResultOK?
Window openedRViz started, no display/authorisation errors✅
OpenGL version in RViz logOpenGl version: 4.6 (GLSL 4.6), no MESA/iris messages✅
Running on the NVIDIA GPUnvidia-smi lists rviz2 using 21 MiB✅
Shutdownclean (SIGINT/SIGTERM from timeout)✅

The only other message, XDG_RUNTIME_DIR not set, is on the guide’s list of harmless messages (Guide §7).

Quick GPU check for RViz

While RViz is open, nvidia-smi on the host should list rviz2 under Processes. If it doesn’t, RViz has fallen back to the Intel GPU.

5.2 Test step 2: simulation bringup

topic/ros2-control pitfall

Goal: start the robot-only launch file with simulated motors and confirm the controllers, the hardware type, the joint-state topic and the RViz model (Guide §7A).

Working rule from here on

Manuel types and runs every command himself, to learn the procedure. Claude explains each step, gives the commands and what to expect, and checks the results for safety. Claude does not launch or stop anything.

What went wrong on the first attempt

Manuel had already started the launch file in his own terminal when Claude, without noticing, started a second copy in the background. The safety check Claude ran just before (ps aux | grep …) did show Manuel’s processes, but Claude went ahead anyway. Result: two /controller_manager nodes, two robot_state_publishers, two hardware interfaces with the same names. The second copy’s spawners failed (Controller already loaded, Failed to configure controller), and list_controllers showed a single unconfigured controller. ROS warned: “there are nodes in the graph that share an exact name”. No risk to the robot (mock hardware, CAN down), but the results were meaningless.

Clean-up (done by Manuel): Ctrl+C in his terminal; docker exec openarm pkill -INT -f "ros2 launch" for the background copy (-INT = the same signal as Ctrl+C); then checked with ps that only the ros2cli.daemon helper was left, and that can0/can1 were still DOWN.

Clean run

Terminal 1:

docker exec -it openarm bash
ros2 launch openarm_bringup openarm.bimanual.launch.py arm_type:=v10 use_fake_hardware:=true

Terminal 2 (checks):

CheckCommandExpectedOK?
Controllersros2 control list_controllers5 controllers, all active✅
Hardware typeros2 control list_hardware_components | grep -E "name:|plugin|state"two components, mock_components/GenericSystem, active✅
Joint statesros2 topic echo --once /joint_states16 joints (joint1–7 + finger_joint1 per arm), all 0.0✅
RVizlook at itboth arms fully drawn✅

Results reported by Manuel as all correct.

Pitfall: check nothing is running before you launch

Run docker exec openarm bash -c 'ps aux | grep -E "ros2|rviz|controller|spawner|robot_state" | grep -v grep' first. Only a ros2cli.daemon line is OK. Two launches at once give two controller managers, which conflict (Guide §6). On the real robot, two hardware interfaces on the same CAN bus would be dangerous.

What the 5 controllers are and where control runs

Arms and grippers each have their own JointTrajectoryController (same type, different joints), so a gripper can be commanded without sending a goal for all 7 arm joints. The controllers and the hardware interface run on the laptop (in ros2_control_node, 750 Hz). The loop that actually produces torque runs inside each motor on the robot (“MIT mode”: the laptop sends target, kp, kd every cycle). See Guide §4.

5.3 Trajectory controller snaps instead of moving smoothly (issue)

issue/ros topic/ros2-control status/resolved pitfall

Symptom

Test step 3, part 1 (Guide §8.1): the left elbow (openarm_left_joint4) was sent to 0.5 rad with time_from_start: {sec: 3}. The goal was accepted, returned SUCCEEDED, and /joint_states showed 0.5 rad. But in RViz the arm didn’t move smoothly over 3 s, it snapped to the new position.

Why it matters

Guide V2.1 (§8.1, §12.3) told you to use a long time_from_start so that the first moves on the real robot are slow. If the controller jumps instead, the motor gets a sudden step in its target and, with the stiff MIT gains (kp = 60 on the elbow), moves at full speed, however long time_from_start is. Following the guide as written, the very first real motion would have been a jerk.

Diagnosis

  • The bringup config (openarm_bringup/config/controllers/openarm_bimanual_controllers.yaml) sets interpolation_method: "none" on all four trajectory controllers (both arms, both grippers). The comment says it’s for streaming high-frequency commands.
  • Claude read the controller’s source (ros2_controllers, branch jazzy, joint_trajectory_controller/src/trajectory.cpp, function Trajectory::sample()). With none: before the first point’s time, the output is the position held when the goal arrived; between two points, the output is the next point. So a one-point goal = hold for time_from_start, then jump. The guide’s claim that it “ramps linearly” was wrong.
  • Confirmed by Manuel in sim: the same move with time_from_start: {sec: 5} did nothing for 5 s, then snapped at the moment SUCCEEDED was printed. Installed package: ros-jazzy-joint-trajectory-controller 4.42.1-1noble.20260924.190511.
  • The MoveIt demo uses a different file (openarm_bimanual_moveit_controllers.yaml) that doesn’t set interpolation_method, so it gets the default, splines, and isn’t affected.
  • Not affected either: the ~2 s move to zero when the motors are switched on. That’s done by the hardware interface (OpenArmHW), not by this controller.

Fix (applied and tested by Manuel)

Decision: switch the bringup trajectory controllers to interpolation_method: splines

Manuel’s decision. The setting is read-only at runtime, so it’s changed in the config file inside the container (4 lines, "none" → "splines"), then the bringup is relaunched. Manuel applies it himself. Steps and checks: Guide V2.2 §4. The edit is tracked by git in /root/ros2_ws/src/openarm_ros2, so git diff shows it and git checkout -- <file> undoes it. A git pull/re-clone would also undo it.

With positions-only goals, splines gives a linear move (constant speed, abrupt start/stop). Adding velocities: [0, …] to the points should give a cubic, eased move; still to verify.

Documentation: corrected in Guide V2.2 (new version, per the convention in 1.2 Versioning convention (agreed today)): §4 new subsection, §8.1, §12.1 (new pre-flight item), §12.3, §14, §16, and the Part II comparison table.

Fix applied and tested

All commands run by Manuel inside the container, with the bringup stopped:

StepCommandResultOK?
Right file?ros2 pkg prefix openarm_bringup/root/ros2_ws/install/openarm_bringup✅
Old /ws workspace in the way?ls -la /wsempty folder (so the .bashrc line in 4.3 Extra line in `.bashrc` does nothing)✅
Editsed -i 's/interpolation_method: "none"/interpolation_method: "splines"/' openarm_bringup/config/controllers/openarm_bimanual_controllers.yaml—
Exactly what changedgit diffexactly 4 lines, "none" → "splines" (left arm, left gripper, right arm, right gripper); no other files modified✅
Rebuild needed?ls -l …/install/openarm_bringup/share/openarm_bringup/config/controllers/openarm_bimanual_controllers.yamlsymlink to the src file, so no rebuild✅
Running settingros2 param get /<controller> interpolation_method for all 4 (after relaunch)splines ×4✅
Arm motionleft joint4 → 0.5 rad, sec: 5, and back to 0moves right away and reaches the target in the time_from_start set✅
Gripper motionleft gripper → 0.044 (open), sec: 2, and back to 0same: reaches the target in the time set✅

ros2 only exists inside the container

Running a ros2 … command in a host terminal gives ros2: command not found. Use a docker exec -it openarm bash shell.

Status: resolved (2026-10-09, evening). Re-check with ros2 param get before the first real-robot launch, since a git pull/re-clone of openarm_ros2 would undo it.

Pitfall: a long time_from_start is not a safety measure on its own

It only slows a move if the controller interpolates. Check ros2 param get /left_joint_trajectory_controller interpolation_method (must be splines) before commanding real motors. And don’t trust the simulator to show you speed problems unless you watch for them: the mock hardware made the jump visible only as a “snap”.

5.4 Hardware stage: plan, and H1 host checks

log/testing topic/can topic/hardware

Plan for the hardware stage (Guide V2.2 §10–12), ordered from “nothing can move” to “the robot moves”:

#Sub-stepMotor powerCan it move?
H0Physical safety setup, plan for cutting poweroffno
H1Host checks: adapter, CAN tools, interfacesoffno
H2Bring up CAN, power on the motors, discover / monitor (read-only)on, torque offonly by hand
H3Check which arm is which, joint signs and the zero pose, by handon, torque offonly by hand
H4Decide on the container’s real-time flags (4.4 Container not created with real-time or CAN-admin flags)offno
H5First ROS launch, one arm only, then test shutdown behaviour with the arm heldonyes: on start it moves to zero in ~2 s
H6First motions (Guide §12.3)onyes

H1 results (run by Manuel on the host, motors unpowered):

CheckCommandResultOK?
Adapter modellsusb | grep -i -E "peak|pcan"0c72:0011 PEAK System PCAN-USB Pro FD✅
CAN tools on hostwhich openarm-can-cli candump/usr/bin/openarm-can-cli, /usr/bin/candump✅
can0 / can1ip -details link show can0 (and can1)both DOWN, can state STOPPED, error counters 0; driver pcan_usb_pro_fd; same USB parent device (3-2.4:1.0), i.e. the two channels of one adapter✅
No LeRobot on the busesps aux | grep -i -E "lerobot|safe_teleop"nothing✅

This answers follow-up 4.1 The CAN adapters are PEAK, not CANable/gs_usb: one PCAN-USB Pro FD, two channels, as Guide §22 says. mtu 16 means the interfaces are currently in classic-CAN mode; that’s normal while unconfigured. The LeRobot setup script (~/Documents/openarm_lerobot/setup_openarm_lerobot.sh) configures them as CAN-FD, 1 Mbit/s nominal / 5 Mbit/s data, with can0 = left and can1 = right, the same as openarm-can-cli can_configure’s defaults and ROS’s can_fd:=true.

5.5 H0 answers, and what really moves the robot (findings)

topic/hardware issue/ros status/open pitfall

H0 answers (Manuel):

  • Power cut: by removing the motor power, from the laptop position.
  • Base: not clamped or bolted.
  • Space around the arms: clear.
  • Holding the arm at shutdown: Manuel himself.
  • Previous use: the robot has been moved before. Last time it snapped to zero positions that were completely off, and the robot fell over. Requirement from Manuel: every position command must be slow.

Requirement before any torque-on (H5): clamp the base

It already fell over once. Clamp or bolt it before the motors are ever enabled under ROS.

Finding 1: the ROS start-up move is not covered by the splines fix. From openarm_hardware/src/openarm_simple_hardware.cpp (on_activate() → return_to_zero(), commit b9d7a67):

  1. enable_all(): torque on.
  2. One full-stiffness command to position 0 for every arm joint ({kp, kd, 0.0, 0, 0}, kp up to 70), then ~1 ms later it reads the current positions. A brief kick toward zero.
  3. A 2 s linear ramp from the current position to 0 at full stiffness (200 steps × 10 ms, hard-coded, no launch argument).
  4. The gripper is commanded straight to a fixed position, with no ramp.
  5. pos_commands_ starts at 0 and write() sends it with full stiffness every cycle, so the driver relies on the arm having reached 0 before the controllers start.

With a wrong zero, step 3 drives the arm 2 s toward a pose that may be far away, maybe into the table or the torso. That matches what happened last time (to confirm which program was running then). The trajectory controller’s splines setting has no effect on this.

Finding 2: openarm-can-cli behaves differently from what Guide §10.5 says (source: enactic/openarm_can f340d4b, setup/cli/commands/):

CommandWhat it actually doesTorque?
discoverruns sudo ip link set … to reconfigure the interface at 1M, 5M, 8M and 10M data rates in turn, and only queries a parameter (MST_ID). Leaves the interface at the last rate tried Correction (H2): at the end it restores the interface to the can_configure defaults (1M/5M FD, sample points 0.75), as its output shows. Claude had read only part of the source.no
monitorenable_all(), then only state requests (no position command), then disable_all() on exit. Decodes every motor as a DM4310: positions are right, but velocity and torque for joints 1–4 (DM8009/DM4340) are scaled wrong.enabled, no command sent
diagnose, motor_statusalso call enable_all()enabled

Guide §10.5 says that during monitor the motors are “disabled = free to move”. In the code they’re enabled; with no position command sent they should produce no torque (⚠ not verified). For the first monitor run: motors freshly powered on, one arm, hold it.

Options for the start-up move (decision pending):

  • A, recommended: patch OpenArmHW so activation holds the current position (sets pos_commands_ to the measured positions, no kick, no ramp). The arm doesn’t move when ROS starts; it’s then moved to 0 with a slow trajectory goal (splines, e.g. 20–30 s). Can’t be tested in sim, because mock doesn’t use OpenArmHW.
  • B: patch it to a slower ramp (e.g. 15 s) and remove the kick. It still moves on start.
  • C: no patch; rely on a verified zero and placing the arm at zero by hand. Doesn’t meet “every command slow”.

Guide corrections for §4, §10.5 and §12 to follow in a new version once these are tested.

Decision and more answers (Manuel)

  • Power cut: a physical switch on the motor supply.
  • The fall last time happened under a LeRobot script, not ROS. Per Guide §21, LeRobot and ROS use the same zero, stored in the motors, so a zero that was “completely off” for LeRobot is most likely off for ROS too. H3 (verify the zero) is mandatory.
  • Start-up move: option B chosen, a slow ramp to zero.

The patch (written by Claude, to be reviewed and applied by Manuel)

File: Prandium/openarm_patches/0001-openarm_hardware-slow-return-to-zero.patch, against openarm_ros2 b9d7a67. It changes OpenArmHW::return_to_zero():

Before (upstream)After (patch)
Reading the start positionfirst sends a full-stiffness command to 0, then readsreads with a state request only, no command
Ramp2 s, linear10 s, eased (smoothstep: gentle start and stop), RETURN_TO_ZERO_SECONDS
Grippersent straight to its targetramped from its current position, same target
Safety check (added)noneif any arm joint is more than 0.5 rad (≈29°) from zero, it doesn’t move: logs which joints, disables the motors, and activation fails (MAX_START_OFFSET)
Logging—prints each joint’s start position
After the ramppos_commands_ left as they wereset to the zero pose, so write() can’t send a stale target after a re-activation

Why 10 s and not longer: the hardware activation blocks the controller manager, and the controller spawners wait on a lock with a 20 s timeout. Upstream’s 2 s is no problem; 10 s leaves a margin. If the spawners do time out, the arm just holds zero and the controllers can be spawned by hand.

The safety check also guards against the zero problem above: with the arm hanging in the zero pose, a wrong motor zero shows up as a large start offset, and the arm won’t move.

git apply --check on the laptop’s copy of the repo (same commit): applies cleanly. Not compiled yet. Can’t be tested in sim (mock doesn’t use OpenArmHW); the first real test is H5.

Applied and built by Manuel inside the container (docker cp the patch to /root/, git apply, colcon build --symlink-install --packages-select openarm_hardware). Undo: git checkout -- openarm_hardware in /root/ros2_ws/src/openarm_ros2, then rebuild.

H2 part 1, CAN interfaces configured (Manuel, host, motor power off):

sudo ip link set canX type can bitrate 1000000 dbitrate 5000000 fd on
sudo ip link set canX up

Both can0 and can1: state UP, can <FD> state ERROR-ACTIVE, error counters 0, bitrate 1000000 sample-point 0.750, dbitrate 5000000 dsample-point 0.750, mtu 72 (CAN-FD frames). timeout 3 candump canX printed nothing (bus silent with motors off). The sample points match openarm-can-cli can_configure’s defaults (0.75/0.75, dsjw 2, from setup/cli/cli.hpp) and what LeRobot’s lerobot-setup-can gets (kernel defaults).

Why bit rates must be set by hand

CAN has no clock wire and no auto-negotiation: every device on a bus must already agree on the timing of each bit, or frames turn into errors and the bus shuts down (bus-off). The motors’ rates are stored in their firmware (1 Mbit/s for the arbitration phase, 5 Mbit/s for the CAN-FD data phase); Linux can’t ask them, so the adapter is told. The setting is lost on reboot or unplug.

H2 part 2, motors powered, discover (Manuel, host; torque stays off):

CheckResultOK?
openarm-can-cli -i can0 discover8 motors: send 0x01–0x08 → receive 0x11–0x18, all at 5 Mbps (FD) (code 9)✅
openarm-can-cli -i can1 discoversame, 8 motors at 5 Mbps (FD)✅
Interfaces afterwardsdiscover restored both to 1M/5M FD by itself; Manuel’s manual reset was redundant. Both UP, ERROR-ACTIVE, error counters 0✅

H3, zero check by hand (Manuel; openarm-can-cli -i canX monitor -d 60000, run inside the container, which works because of --network host). Arms hanging straight down, grippers closed. Motors stayed limp while monitor had them enabled.

IDJointLeft (can0)Right (can1)
0x01joint1−0.000+0.002
0x02joint2−0.009+0.005
0x03joint3−0.042−0.063
0x04joint4−0.002−0.025
0x05joint5−0.015+1.394 ❌
0x06joint6−0.078−0.003
0x07joint7−0.060+0.064
0x08gripper+0.004+0.003
  • Left arm = can0, right arm = can1, as in Guide §22.
  • Left arm: zero OK. The largest offsets (0.04–0.08 rad, 2–4.5°) are on joints 3, 6 and 7, which gravity doesn’t hold in one exact position when hanging, so they’re within hand-placement precision. The patched start-up would move them to 0 slowly over 10 s.
  • Right joint 5: zero wrong by 1.394 rad (≈ 80°). Manuel confirms it’s physically in the same pose as the left arm’s joint 5. With the patch, ROS start-up of the right arm would refuse to move (1.394 > 0.5). Without the patch, upstream would have swung the wrist ~80° in 2 s. Most likely cause: a LeRobot zeroing with that wrist rotated (the motor zero is shared, see Guide §21).

Pitfall: openarm-can-cli set_zero with no --id zeroes all 8 motors of that bus

Its default is --arm (IDs 1–8). --id overrides it (source setup/cli/openarm_cli.cpp, installed openarm-can-utils 1.4.0). To fix one motor: openarm-can-cli -i can1 set_zero --no-arm --id 5. It sends Disable → Set-Zero (0xFE) → Disable, so no torque.

H3b, right joint 5 re-zeroed (Manuel): joint placed by hand to match the left arm’s joint 5, then openarm-can-cli -i can1 set_zero --no-arm --id 5. monitor afterwards: 0x05 ≈ 0.00. Still ≈ 0.00 after a motor power cycle, so the zero is stored in the motor. (Optional end-stop symmetry check skipped.) Since the zero is shared, this also fixes it for LeRobot. ✅

The base can't be clamped (Manuel)

Accepted risk. Mitigations for H5–H6: ballast the base with heavy objects; first power-on with one arm only; arm starts hanging (near zero, so the start-up check passes and the ramp is short); only small, slow moves close to the hanging pose; motor-power switch within reach at all times.

5.6 H4: new container with real-time permissions

topic/docker topic/ros2-control log/testing

Why real-time is needed

The controller manager (ros2_control_node) runs one loop 750 times a second: read the motors over CAN → run the controllers → send the next MIT command to every motor. That leaves 1.33 ms per cycle. A normal Linux process is scheduled “fairly” (SCHED_OTHER): whenever RViz, the browser, Docker or the GPU driver want the CPU, the loop may have to wait. When it misses its 1.33 ms slot, the log says Overrun detected! … missed cycles: 2, which we saw regularly in sim.

  • In sim that’s harmless: mock hardware has no physics.
  • On the real robot, a late cycle means the motors get their next target late, and then a bigger step to catch up. With stiff MIT gains that’s felt as small jerks or buzzing. Motors that don’t get a frame for a while may also hit their CAN timeout, if one is configured.

Real-time scheduling (SCHED_FIFO) fixes that: a FIFO thread with a priority always runs before every normal process, as soon as it’s ready. ros2_control asks for SCHED_FIFO priority 50 for its loop automatically, and only needs permission. It also locks its memory (mlockall) so the loop never waits for memory to be paged in from disk.

Why Docker doesn’t allow it by default

Docker runs containers with as few permissions as possible, so a program in a container can’t harm the host:

  • Capabilities. Linux splits root’s powers into ~40 “capabilities”. Docker gives containers only a small default set, and CAP_SYS_NICE (raise priorities, use real-time scheduling) isn’t in it. Even root inside the container can’t use SCHED_FIFO without it.
  • Why that’s the safe default: a SCHED_FIFO thread is never interrupted by normal programs. If it gets stuck in a loop, it can freeze a CPU core and make the whole host unresponsive. So Docker leaves this to the user to enable on purpose.
  • Limits (ulimit). rtprio (the highest real-time priority a process may use) defaults to 0, i.e. none. memlock (how much memory a process may lock in RAM) defaults to a few MB, too little for mlockall.

That’s why the controller manager printed “Could not enable FIFO RT scheduling” in the old container. The three flags give exactly those permissions:

FlagGives
--cap-add SYS_NICEthe right to raise priority / use SCHED_FIFO
--ulimit rtprio=99allows RT priorities up to 99 (ros2_control uses 50)
--ulimit memlock=-1no limit on locked memory

The Linux kernel keeps its own safety net: by default real-time tasks may use at most 95 % of each second (/proc/sys/kernel/sched_rt_runtime_us = 950000), so the host can’t be frozen completely.

What was done (Manuel, host terminal)

Flags can only be set when a container is created (docker run), so a new container was made from a snapshot of the old one:

StepCommandResult
Current flags recordeddocker inspect openarm --format …Network=host IPC=host, bind /tmp/.X11-unix, GPU request (--gpus all), CapAdd=null, Ulimits=null; env DISPLAY=:1, NVIDIA_DRIVER_CAPABILITIES=all, __NV_PRIME_RENDER_OFFLOAD=1, __GLX_VENDOR_LIBRARY_NAME=nvidia (plus image defaults)
Nothing runningdocker exec openarm bash -c 'ps aux | grep …'only ros2cli.daemon
Stopdocker stop openarm—
Snapshotdocker commit openarm openarm-jazzy:rtnew image openarm-jazzy:rt (7.37 GB on disk, shares layers with :built)
Keep the old onedocker rename openarm openarm-deprecatedold container kept as a backup, untouched
New containerdocker run -it --name openarm --network host --ipc host --gpus all -e NVIDIA_DRIVER_CAPABILITIES=all -e __NV_PRIME_RENDER_OFFLOAD=1 -e __GLX_VENDOR_LIBRARY_NAME=nvidia -e DISPLAY=$DISPLAY -v /tmp/.X11-unix:/tmp/.X11-unix --cap-add SYS_NICE --ulimit rtprio=99 --ulimit memlock=-1 openarm-jazzy:rt—

Checks in the new container:

CheckResultOK?
ulimit -r; ulimit -l99, unlimited✅
git diff --stat in openarm_ros2same 3 files (splines config + the patch): the snapshot kept everything✅
nvidia-smi -LRTX 5050✅
ros2 pkg prefix openarm_hardware/root/ros2_ws/install/openarm_hardware✅
Sim launch, grep -i -E "FIFO|RT scheduling|overrun"Successful set up FIFO RT scheduling policy with priority 50. One overrun (1.51 ms) in the run, instead of regular ones✅

The remaining occasional overrun is expected on a normal (non-PREEMPT_RT) laptop kernel with RViz running (Guide §12.5). groups: cannot find name for group ID 992 on every shell is harmless (Guide §7).

Containers and images from now on

  • openarm (image openarm-jazzy:rt): the one to use. Daily commands unchanged: docker start openarm, docker exec -it openarm bash. Its main shell is the docker run -it terminal; closing it stops the container.
  • openarm-deprecated: the old container, kept as a backup until the robot has run well on the new one. Don’t start both at once: both use host networking and would fight over ROS and CAN.
  • Images: openarm-jazzy:latest, :built, :rt. :rt = :built + today’s changes.

This resolves 4.4 Container not created with real-time or CAN-admin flags (except NET_ADMIN, not needed: CAN is configured on the host).

Guide update pending (next version, after the hardware tests): §5 container command and image name, §10.5 (monitor enables motors), §11 (set_zero default = all 8), §4/§12 (the start-up patch). Done in Guide V3.1 (2026-10-10).

5.7 H5: first ROS launch on the real robot (left arm only)

log/testing topic/hardware topic/ros2-control

Setup. Right arm unpowered. There’s no single-arm launch file, so the right arm’s driver was pointed at a virtual CAN bus (vcan1), so its commands can’t reach anything physical and there are no bus errors from unanswered frames:

sudo modprobe vcan
sudo ip link add dev vcan1 type vcan
sudo ip link set vcan1 mtu 72      # CAN-FD frame size
sudo ip link set vcan1 up

A motor that has never answered reads 0.0 in openarm_can (Motor constructor), so the right driver sees “all joints at 0”, passes the start-up check and ramps 0 → 0. Pre-checks: only the openarm container running, no ROS processes, splines ×4, can0 up at 1M/5M FD, left arm hanging, weight on the base.

Launch (container):

ros2 launch openarm_bringup openarm.bimanual.launch.py arm_type:=v10 use_fake_hardware:=false left_can_interface:=can0 right_can_interface:=vcan1 can_fd:=true 2>&1 | tee /root/h5_launch.log

Results:

  • Successful set up FIFO RT scheduling policy with priority 50. All controllers report Using 'splines' interpolation method.
  • Both hardware components loaded OpenArmHW (CAN=can0, arm_prefix=left_ and CAN=vcan1, arm_prefix=right_). The right arm was activated first (start values all +0.000, 10 s ramp, Reached zero position), then the left (lines not captured in the paste; still to check in /root/h5_launch.log).
  • Nothing visibly moved. Expected: the left arm hung at zero, with the largest offset ≈ 0.08 rad (4.5°) on the wrist, spread over 10 s. Manuel stayed clear of the robot.
  • All controllers spawned and activated despite the ~20 s of activation ramps.
  • Constant overruns: each loop took 2.0–2.6 ms instead of 1.33 ms (read 1.25–1.57 ms, write 0.5–0.75 ms), so it actually ran at about 400–450 Hz. The main cause is the right driver waiting for replies that never come on vcan1 (recv_all waits up to 500 µs in read() and 100 µs in write()). Not dangerous (trajectories are timed by the clock), but to measure again with both arms powered.

Pitfall: pressing Ctrl+C several times

Manuel stopped the launch with Ctrl+C ×3. ros2 launch escalates on repeated Ctrl+C (SIGINT → SIGTERM → SIGKILL). A killed controller manager never runs the driver’s on_deactivate (motors off), so the motors may keep holding their last command with torque on. Press Ctrl+C once and wait for the shutdown messages; only repeat if nothing happens for ~10 s. Here the arm was hanging at zero, so it was harmless; the motor power was switched off and on afterwards to get a clean state.

Still unknown: what Ctrl+C does to the motors on this setup (torque off, or holding). Test it next time with the arm at the zero pose (nothing to fall).

Correction: the left arm was not powered during H5 (or at the start of the H6 relaunch)

Found while preparing H6. The left arm’s start values in /root/h5_launch.log were exactly +0.000 on all 7 joints, while monitor had read −0.042 / −0.078 / −0.060 on joints 3/6/7 in the same pose. 0.0 is the default of a motor that has never replied (openarm_can Motor constructor). Manuel confirmed the left arm’s power was off. So H5 tested nothing on the real arm: “nothing moved” because nothing was powered. The constant overruns fit too: both drivers waited the full reply timeout every cycle.

Manuel then switched the left arm on while ROS was running. /joint_states immediately showed real values matching monitor (e.g. joint3 −0.0422, joint6 −0.0784, joint7 −0.0597, non-zero efforts). But the motors stayed disabled (the enable command was sent at start-up, when they were off), and ROS’s stored targets were all 0 (from the empty readings). Inconsistent state, so no goals were sent; the launch was to be stopped and redone with the arm powered first.

Flaw in patch 0001, and patch 0002

The start-up check in patch 0001 compared the read positions with zero, but a motor that hasn’t replied reads 0.0, so the check passed without meaning anything. If the arm had been far from zero (and powered just after the read), it would have been driven to 0 with no ramp: the original snap.

Patch 0002 (Prandium/openarm_patches/0002-openarm_hardware-check-motors-replied.patch, applies on top of 0001): after the torque is switched on, every arm motor and the gripper must report is_enabled() with no error code (openarm_can sets that flag only from the motor’s own state replies, dm_motor_device.cpp), otherwise it logs which ones, switches the motors off and refuses to start, the same as the distance check. Dry-run on a clean copy of the repo + 0001: applies cleanly.

Consequence: an arm on the virtual vcan1 bus (no replies) now fails start-up by design, so the one-arm trick needs revisiting.

Patch 0002 applied and built by Manuel (docker cp, git apply, colcon build --symlink-install --packages-select openarm_hardware): Finished <<< openarm_hardware, only the expected on_init(HardwareInfo) deprecation warning.

Pitfall: Ctrl+C on ros2 launch … | tee file hides the shutdown

Ctrl+C goes to every program in the pipe. tee exits at once, so the prompt returns while ROS is still shutting down out of sight, and programs writing to the broken pipe can be killed before switching the motors off. Use tee -i (ignore interrupts) so the output keeps flowing until ROS has exited. Node messages are also saved in /root/.ros/log/.

Pitfall: power the arms before launching ROS, never while it's running

The driver sends “torque on” once, at start-up. An arm powered later stays limp while ROS believes it’s in control, with stored targets that may be far from the real pose.

5.8 H5 (for real): both arms powered, both patches

log/testing topic/hardware topic/ros2-control

Option A (Manuel’s choice): both arms powered before the launch, hanging at zero; only the left arm is to be commanded. Previous launch had stopped cleanly. Launch (container):

ros2 launch openarm_bringup openarm.bimanual.launch.py arm_type:=v10 use_fake_hardware:=false left_can_interface:=can0 right_can_interface:=can1 can_fd:=true 2>&1 | tee -i /root/h5b_launch.log

Start-up (from the log), all real readings, all inside the 0.5 rad limit, no no reply / rad from zero errors:

JointRight (can1, activated first)Left (can0)
1−0.012−0.000
2+0.016−0.009
3−0.009−0.014
4−0.026−0.002
5−0.004 (re-zeroed in H3b ✅)−0.015
6+0.053−0.022
7+0.060−0.015

Each arm: Reached zero position ~10.5 s after its start. (Values differ a little from the H3 monitor snapshot because the arms were re-hung by hand in between.)

Checks: both hardware components openarm_hardware/OpenArmHW, active. All 5 controllers active. /joint_states shows real values for both arms.

Observation: the joints stop 0.01–0.02 rad short of zero. After the ramp, e.g. left joint6 −0.0216, joint7 −0.0143, joint3 −0.0132, i.e. almost exactly their start values. Right joint6 went +0.053 → +0.0147. The reported efforts match the MIT formula kp·(target − position) (target 0): left joint3 70 × 0.0132 ≈ 0.92 Nm (reported 0.87), left joint6 10 × 0.0216 ≈ 0.22 Nm (reported 0.20), right joint4 60 × 0.012 ≈ 0.72 Nm (reported 0.68). So the motors are pushing toward zero, but a small error × the gain gives too little torque to overcome gearbox friction. It’s normal for a PD controller with no integral term (Guide §4: no gravity compensation, no feed-forward). Expect commanded moves to land ~0.01–0.02 rad short on the wrist joints.

5.9 H6: first commanded moves, and the joint direction map

log/testing topic/hardware topic/ros2-control

First moves (Manuel): left joint7 → +0.2 rad over 10 s, and left joint1 → +0.3 rad. Both moved slowly and matched RViz, but went “backwards”, not “forwards”. That’s correct: ”+” is defined by each joint’s axis in the model, not by “forward”. Matching RViz is the real test, since it shows the motor direction and the model agree.

Joint direction map (OpenArm v1.0, from the model)

How it was worked out: openarm_description at commit c103041 (the container’s version) was expanded with xacro (arm_type:=v10 bimanual:=true) on the laptop, in Claude’s scratch folder. Each joint’s axis was then transformed into the robot’s frame at the zero pose (arms hanging), and the direction the hand moves for a small + rotation was computed. Both of Manuel’s tests (left joint1, left joint7 = backward) agree with it.

Robot frame: forward = the direction the robot faces (the way the elbows bend), left/right = the robot’s own left/right, up. Directions are for small moves from the hanging zero pose; once other joints have moved, each axis moves with the arm.

JointWhat it isLeft arm: + moves…Right arm: + moves…
joint1shoulder, swing forward/backhand backwardhand forward
joint2shoulder, swing sidewayshand inward (toward the body)hand outward (away from the body)
joint3upper-arm twistclockwise seen from above (with the elbow bent, the forearm swings toward the robot’s right)same: clockwise seen from above
joint4elbowforearm forward/up (elbow bends)same: forward/up
joint5forearm twistclockwise seen from abovesame
joint6wrist, sidewayshand outwardhand inward
joint7wrist, forward/backhand backwardhand forward
finger_joint1gripper (metres)0 = closed, 0.044 = opensame

Pattern: the arms are mirror images, so joint1, joint2, joint6 and joint7 have opposite meanings on the two arms. joint3 and joint5 turn the same way in the room (clockwise from above), so in body terms they’re mirrored (inward on one arm = outward on the other). joint4 is the same on both.

The joint-limit table in Guide §3 is wrong for joint1 (left) and joint2 (both arms)

Guide §3 copies config/arm/joint_limits.yaml. But the model adds a per-arm offset and mirroring for joints 1 and 2 (urdf/arm/openarm_arm.xacro: joint1 limits - 2.094 on the left; joint2 ± π/2 plus mirroring), so the real limits are:

JointGuide §3 (yaml)Left arm (actual)Right arm (actual)
joint1−1.396 … +3.491−3.491 … +1.396 (forward to 200°, back to 80°)−1.396 … +3.491 (back to 80°, forward to 200°)
joint2−1.745 … +1.745−3.316 … +0.175 (outward to 190°, inward only 10°)−0.175 … +3.316 (inward only 10°, outward to 190°)
joints 3–7as in §3same as §3same as §3

So e.g. left joint2 = +1.0 rad, “inside the limits” per the guide, actually drives the arm into the body. The trajectory controller does not enforce limits (Guide §12.3). Always check targets against this table, per arm.

Guide update pending: add this direction map, and replace the §3 limit table with the per-arm one. Done in Guide V3.1 (2026-10-10).


6. Lessons and pitfalls from today

pitfall

Pitfall: charger and the NVIDIA GPU

Changing power source (charger in or out) crashed the NVIDIA driver (Xid 154). Keep the charger plugged in. Details in 3. Issue: NVIDIA GPU crashed (Xid 154).

Pitfall: docker run is not how you get back into the container

docker run makes a new container, without anything you installed or built in the old one. Use docker start openarm (if stopped) and then docker exec -it openarm bash. See 1.1 What changed.

Pitfall: two launches at once

Check with ps that nothing is running before launching. See 5.2 Test step 2: simulation bringup.

Pitfall: time_from_start doesn't slow a move with interpolation_method: none

The joint waits, then jumps. Check the setting is splines before real-robot tests. See 5.3 Trajectory controller snaps instead of moving smoothly (issue).

Convention: guide versions are never overwritten


7. Next session

log/follow-up

State at end of day: both arms tested on real hardware with patches 0001 + 0002, splines, real-time container openarm (image openarm-jazzy:rt). Manuel to stop for the night: arms back to zero, Ctrl+C once on the launch (watch for Deactivating OpenArm V10... ×2 = the shutdown test), motor power off.

After a reboot: xhost +SI:localuser:root → docker start openarm (the new one, not openarm-deprecated) → can0/can1 set up again (1M/5M FD, see H2) → power the arms → launch with tee -i.

To do, in order:

  • Record the shutdown-test result (H5d): did Deactivating OpenArm V10... appear twice, and did the arms go limp?
  • Test every joint of both arms once: ≤0.3 rad, 10 s, one at a time, back to zero each time. Check the direction against RViz and against Joint direction map (OpenArm v1.0, from the model); check each target against the per-arm limits in that section (joint2 especially).
  • Grippers: 0.022 (half) first, then 0.044 / 0.0 (Guide §12.3).
  • Look at the overruns now that both arms answer (grep -c "Overrun detected" <launch log>, and the read/write times).
  • Guide V2.3 → done as Guide V3.1 (2026-10-10): container (openarm-jazzy:rt, RT flags), §10.5 (monitor enables motors), §11 (set_zero default = all 8 motors), §4/§12 (the start-up patches, power before launch, tee -i), §3 (per-arm limits) + the direction map.
  • Later: steps 4 (controller switching) and 5 (MoveIt), skipped today; delete openarm-deprecated once the new container has proven itself.