← Back to the Log Index · Previous day: (none, this is the first log) · Next day: 2026-10-10
Summary of the day
- We added an explanation of the
docker run/docker start/docker execcommands to the OpenArm guide, which became a new minor version, Guide V2.1.- We agreed a versioning convention for the guides (see 1. Documentation changes).
- We started a step-by-step test of the ROS 2 setup, beginning with the environment (host, Docker container, GPU, CAN).
- The container and workspace are fine. The test was blocked because the laptop’s NVIDIA GPU crashed (driver error Xid 154, GPU Reset Required) after a power-source change, and only a reboot fixes that. See 3. Issue: NVIDIA GPU crashed (Xid 154).
- Afternoon: after a reboot (15:59) the GPU works again, and test step 1 (GPU and display) passed: RViz runs on the NVIDIA card. See 3.6 Resolution (after the 15:59 reboot) and 5.1 Test step 1: GPU and display.
- Test step 2 (simulation bringup) passed, run by Manuel himself. A first attempt ran two launches at once by mistake; see 5.2 Test step 2: simulation bringup.
- New working rule: Manuel runs every command himself; Claude explains, gives the commands and checks results for safety.
- Evening, hardware stage: CAN checked (PCAN-USB Pro FD, 8 motors per arm), right joint 5 zero fixed (was off by 80°), motor driver patched (slow 10 s start-up, no kick, refuses to move if a motor doesn’t reply or a joint is >0.5 rad from zero), new container with real-time scheduling, first real launch of both arms and first slow moves (left joint1 and joint7, matching RViz). Found that Guide §3’s joint limits are wrong for joint1 (left) and joint2 (both); joint direction map worked out. See 5.4 Hardware stage: plan, and H1 host checks → 5.9 H6: first commanded moves, and the joint direction map.
- Where to pick up tomorrow: 7. Next session.
- Late evening: a new major guide version, Guide V3 (a major number because Part III is a large addition), adds Part III, a beginner-level, step-by-step guide to running trained policies with LeRobot (
lerobot-rollout). It was written from the installed LeRobot source and hasn’t been tested on the robot yet, since no policy has been trained. While checking the source, three problems with V2.2 came up. Rollout datasets must be namedrollout_…, so V2.2’seval_…/hil_…examples would fail.lerobot-recordappends a date-time tag to dataset names unless--dataset.no_stamp=true. And zeus runs a Wayland session, where LeRobot’s keyboard controls probably don’t work.- Test step 3 found a safety problem: the trajectory controller doesn’t move smoothly to a target, it waits for
time_from_startand then jumps (interpolation_method: none). The guide said the opposite. Corrected in Guide V2.2. Manuel switched the controllers tosplinesand confirmed smooth motion in sim: goals now reach their target in exactlytime_from_start. See 5.3 Trajectory controller snaps instead of moving smoothly (issue).
Contents of this log
- 1. Documentation changes
- 2. Test step 0: environment check
- 3. Issue: NVIDIA GPU crashed (Xid 154)
- 4. Observations to follow up later
- 5. Test plan and where we stopped
- 6. Lessons and pitfalls from today
- 7. Next session
1. Documentation changes
1.1 What changed
The ROS 2 guide previously only showed the three Docker commands used every day without explaining how they differ. A new subsection now explains them, in §5 Environment and container of Guide V2.1. In short:
| Command | What it acts on | What it does |
|---|---|---|
docker run | an image (a frozen template, here openarm-jazzy:built) | Creates a brand-new container and starts it. All the settings (network, GPU, display) are fixed at this moment. Only ever run it once for a given container. |
docker start | an existing stopped container | Restarts the container, with everything previously installed or built inside it still there. |
docker exec | an existing running container | Opens an extra process (usually a bash shell) inside it. Use it once for every terminal you need. |
1.2 Versioning convention (agreed today)
Rule: never overwrite a guide version's content
Small additions or corrections to a guide’s content are saved as a new file with the minor version bumped, for example
…_V2.md→…_V2.1.md→…_V2.2.md. The previous file’s content stays exactly as it was. Major rewrites get a new major number (V3).Exception: changes that only affect formatting or navigation, such as converting links to Obsidian format or adding navigation links, are made in place in every existing version, without a new version number. They don’t change what the guide says. See 1.4 Obsidian link conversion (all versions, in place).
Why: the guides are shared, so people may be reading or referring to an older version. Keeping each version as its own file means nothing they rely on changes without warning, and the history stays visible.
Mistake made today (by Claude), already corrected
The first edit went directly into
V2, overwriting it. That was undone: the changes were moved into a newV2.1file andV2was restored to its original content (105 947 bytes at that point, before the link conversion in §1.4). This is why the convention above was written down.
1.3 Files at end of day
All guides now live in Prandium/Documents/OpenArm Docs/. They were moved there during the session; earlier they sat directly in Prandium/Documents/.
- OpenArm_ROS2_and_teleop_Guide_V1: first version.
- OpenArm_ROS2_and_teleop_Guide_V2: second version, unchanged.
- OpenArm_ROS2_and_teleop_Guide_V2.1: V2 plus the Docker command descriptions.
- OpenArm_ROS2_and_teleop_Guide_V2.2: V2.1 plus the interpolation correction and the §12.1 CAN-mapping correction (added in the evening, see 5.3 Trajectory controller snaps instead of moving smoothly (issue)).
- OpenArm_ROS2_and_teleop_Guide_V3: V2.2 plus Part III, a step-by-step guide to running a trained policy with
lerobot-rollout, and corrections to the LeRobot dataset names (see the summary at the top). This is the current version.
1.4 Obsidian link conversion (all versions, in place)
Why: the guides were written with GitHub-style links, such as [Environment and container](#5-environment-and-container). Those work on GitHub but not in Obsidian, which uses its own link format and doesn’t recognise GitHub’s automatically generated “anchor” names. Cross-references in the text were also written as plain §12 with no link, so you had to scroll to find them.
What was done (to V1, V2 and V2.1, edited in place, see the exception in 1.2 Versioning convention (agreed today)):
| Change | Before | After | Count (V1 / V2 / V2.1) |
|---|---|---|---|
| Table-of-contents links | [Big picture](#1-big-picture) | [[#1. Big picture|Big picture]] | 29 / 29 / 29 |
§ cross-references in the text | see §12 | see [[#12. The physical robot: bringing it up with ROS 2|§12]] | 40 / 69 / 69 |
| Navigation box under the title | (none) | Links to the other guide versions and to the Log Index | 1 / 1 / 1 |
What it looks like when reading: the §12 text looks the same but is now clickable. Hovering shows a preview of that section if Obsidian’s built-in Page preview core plugin is enabled (it is by default; hold Ctrl while hovering).
How it was done (so it can be repeated on future guides): a small Python script read every heading in the file, matched each old GitHub anchor and each §N / §N.M reference to the real heading text, and rewrote it as an Obsidian heading link. Notes on the details:
- Text inside code blocks and
inline codewas left alone, so commands that contain§or[[(e.g. the Python list[[0, 0, …]]in §8.3) weren’t changed. - Inside tables, the
|separating a link’s target from its display text has to be written\|, otherwise Markdown treats it as a new table column. The script did this automatically. - A range such as
§24.4–§24.6became two separate links. - Afterwards, a check confirmed that every link points to a heading that exists, that every
§N.Mpoints to section N.M and not just to N, and that no GitHub-style links are left. Two table-of-contents entries whose titles contain inline code (§16, §28) were missed by the script and fixed by hand. - A backup of the files from before the conversion was kept in Claude’s temporary session folder. It is not permanent.
Tip: linking to a heading in Obsidian
Inside the same file:
[[#Heading text]]. In another file:[[File name#Heading text]]. To show different text:[[#Heading text|shown text]]. Typing[[#in Obsidian opens an autocomplete list of the headings.
2. Test step 0: environment check
Goal of this step: before launching any ROS nodes, confirm that what ROS runs on is correct: the host machine, the Docker container and its settings, the graphics card (needed by RViz, the 3-D visualiser) and the CAN interfaces (needed later for the real robot). If one of these is wrong, the ROS tests fail in confusing ways, so it’s quicker to check them first.
2.1 Results
| Check | Command used | Result | OK? |
|---|---|---|---|
| Display variable on host | echo $DISPLAY | :1 (Wayland session with XWayland) | ✅ |
User can run Docker without sudo | id -nG | user is in the docker group | ✅ |
| Image exists | docker images | openarm-jazzy:built (7.36 GB) and an older openarm-jazzy:latest | ✅ |
| Container exists and is running | docker ps -a --filter name=openarm | Up 24 hours | ✅ |
| Container was created with the right flags | docker inspect openarm | network = host, IPC = host, GPUs = all, DISPLAY=:1, __NV_PRIME_RENDER_OFFLOAD=1, __GLX_VENDOR_LIBRARY_NAME=nvidia, NVIDIA_DRIVER_CAPABILITIES=all | ✅ |
| Shell setup inside container | grep source /root/.bashrc | sources /opt/ros/jazzy/setup.bash and ~/ros2_ws/install/setup.bash; no leftover humble lines | ✅ |
| Workspace sources present | ls /root/ros2_ws/src | openarm_can, openarm_description, openarm_ros2 | ✅ |
| Workspace built | ls /root/ros2_ws/install | openarm, openarm_bimanual_moveit_config, openarm_bringup, openarm_can, openarm_description, openarm_hardware | ✅ |
| NVIDIA GPU usable (host) | nvidia-smi | Unable to determine the device handle for GPU0: 0000:01:00.0: Unknown Error | ❌ |
| NVIDIA GPU usable (container) | docker exec openarm nvidia-smi -L | same error | ❌ |
| CAN interfaces exist | ip -br link | can0 and can1 present, both DOWN | ✅ (down is expected, we aren’t using hardware yet) |
You no longer need to
sourceby hand in each shellThe guide tells you to run the two
sourcelines in every new shell. On zeus,/root/.bashrcinside the container already does this, so everydocker exec -it openarm bashshell is ready to use. You only need to source by hand again after rebuilding the workspace (source /root/ros2_ws/install/setup.bash) in a shell that was already open.
3. Issue: NVIDIA GPU crashed (Xid 154)
3.1 Symptom
nvidia-smi (the NVIDIA status tool) fails on the host and in the container:
Unable to determine the device handle for GPU0: 0000:01:00.0: Unknown Error
No devices were found3.2 Why this matters for ROS
zeus is a hybrid-graphics laptop: it has an Intel integrated GPU and an NVIDIA RTX 5050. The container is deliberately set up to force OpenGL programs onto the NVIDIA card (the __NV_PRIME_RENDER_OFFLOAD=1 and __GLX_VENDOR_LIBRARY_NAME=nvidia variables), because RViz failed on the Intel/Mesa driver in earlier attempts (see Guide §15). With the NVIDIA card in an error state, RViz will fail to open, and so will both launch files in the guide, since both start RViz. Nodes without a GUI (controllers, ros2 topic, etc.) don’t need the GPU.
3.3 Diagnosis
The kernel log (journalctl -k -b | grep -i nvrm) shows what happened, in this order:
14:36:46 NVRM: rm_power_source_change_event: Failed to handle Power Source change event, status=0xf
14:36:46 NVRM: Xid (PCI:0000:01:00): 154, GPU recovery action changed from 0x0 (None) to 0x1 (GPU Reset Required)
15:08:30 NVRM: rm_power_source_change_event: … Failed … (repeats several times)
15:24:42 WARNING: nvidia/nv.c:5388 at nvidia_dev_put … (driver stack trace)What this means:
- A “power source change” is the laptop switching between mains power and battery, i.e. the charger being plugged in or unplugged (or a loose charger connection).
- The NVIDIA driver (version 595.91.07, the “open” kernel module) failed to handle that switch and marked the GPU as needing a reset. Xid codes are NVIDIA’s numbered error categories, and 154 means “GPU recovery action required”. Here the action is GPU Reset Required.
- The kernel module itself is still loaded (
lsmodshowsnvidia,nvidia_drm, etc.), and the PCI device reports asactive, so the hardware is still there. The driver just won’t talk to it until it’s reset.
3.4 Fix
Fix: reboot the laptop
On a laptop GPU, a full reboot is the reliable way to perform the “GPU reset” the driver asks for. Afterwards:
- On the host, run
nvidia-smi. It should list the RTX 5050 with no errors.- Run
xhost +SI:localuser:root(needed after every login, so the container may open windows).- Run
docker start openarm. The container stops during a reboot, which is normal.- Run
docker exec openarm nvidia-smi -Lto confirm the container sees the GPU too.
3.5 Prevention
Pitfall: don't plug or unplug the charger while working with the GPU
On this machine and driver version, switching between battery and mains power while the NVIDIA GPU is in use can crash the GPU driver until the next reboot. Plug the charger in before starting and leave it in. If the GPU does crash:
nvidia-smierrors → reboot.If this keeps happening even with the charger left alone, it’s worth checking for a newer NVIDIA driver, or trying the proprietary (closed) kernel module instead of the open one. Not done yet.
Status: open resolved, see below.
3.6 Resolution (after the 15:59 reboot)
GPU working again
The laptop was rebooted at 15:59 with the charger plugged in. Checks at 16:08:
Check Result nvidia-smion the hostRTX 5050 listed, no errors, 12 MiB used Kernel log since boot ( journalctl -k -b | grep -i xid)no Xid or power-source errors Charger ( /sys/class/power_supply/AC*/online)1, plugged inxhostSI:localuser:rootalready allowedContainer Up, already started after the rebootdocker exec openarm nvidia-smi -LGPU 0: NVIDIA GeForce RTX 5050 Laptop GPU
4. Observations to follow up later
4.1 The CAN adapters are PEAK, not CANable/gs_usb
The kernel module list includes peak_usb, the driver for PEAK-System USB-CAN adapters (e.g. PCAN-USB FD). Guide §10.1 assumes gs_usb-class adapters (CANable 2.0 / candleLight). Both kinds work with Linux SocketCAN (the standard Linux CAN interface that can0 and can1 are part of), so the ROS side doesn’t care. But:
- The bit-rate setup in §10.4 may need different options for PEAK adapters, especially for CAN-FD data bit-rate and sample point.
- The udev naming rules in §10.3 (which make sure the same physical adapter always gets the same name,
can0orcan1) match on USB IDs, which differ between vendors.
To do before the hardware tests: confirm the adapter model (lsusb), then check and adjust §10.3–10.4 in a new guide version.
Update (evening): Guide §22 already describes zeus’s adapter as a single PEAK PCAN-USB Pro FD with two channels, giving stable can0/can1 names, and says the cables are the reverse of the ROS defaults (left arm = can0). So the udev naming in §10.3 isn’t needed on zeus; the §10.4 bit-rate options for PEAK still need checking. The §12.1 checklist contradicted §22 on the left/right mapping; corrected in V2.2.
4.2 Two Docker images exist
There are two images, openarm-jazzy:latest (6.87 GB) and openarm-jazzy:built (7.36 GB). The container uses :built. Most likely :latest is the base image from before the workspace was built and :built was committed afterwards, but this hasn’t been verified. Don’t delete either until we know.
4.3 Extra line in .bashrc
/root/.bashrc line 101 is [ -f /ws/install/setup.bash ] && source /ws/install/setup.bash, a leftover from an earlier setup that used /ws as the workspace. It only runs if that file exists, so it’s harmless, but it could cause confusion if a /ws folder ever appears. It can be removed when convenient.
4.4 Container not created with real-time or CAN-admin flags
Resolved in the evening: see 5.6 H4: new container with real-time permissions.
docker inspect shows CapAdd=[], so the container does not have --cap-add SYS_NICE, --ulimit rtprio=99 or --cap-add NET_ADMIN. That’s fine for simulation. For the real robot, Guide §5 recommends the real-time flags at 750 Hz. Since flags can only be set at docker run time, this will mean creating a new container (with docker commit first, so nothing installed inside is lost). That decision is for the hardware stage.
5. Test plan and where we stopped
flowchart TD S0["0. Environment check<br/>(host, container, CAN)"]:::done --> G["GPU working?"]:::done G -->|"after reboot"| S1["1. GPU and display test<br/>nvidia-smi + OpenGL window"]:::done S1 --> S2["2. Sim bringup (Guide §7A)<br/>controllers active, v1.0 model, RViz"]:::done S2 --> S3["3. Commanding (Guide §8.1–8.4)<br/>each arm and gripper, check /joint_states"]:::done S3 --> S3b["3b. Switch JTC to splines<br/>and test smooth motion in sim"]:::done S3b --> S4["4. Controller switching (Guide §8.5–8.6)"] S4 --> S5["5. MoveIt (Guide §9)<br/>incl. gripper caveat"] S5 --> S6["6. Hardware (Guide §10–12)<br/>CAN checks first, motors unpowered"] classDef done fill:#2e7d32,color:#fff classDef blocked fill:#c62828,color:#fff
- Step 0: environment check, done, see 2. Test step 0: environment check
- Reboot to clear the GPU error, see 3.6 Resolution (after the 15:59 reboot)
- Step 1: GPU and display test, done, see 5.1 Test step 1: GPU and display
- Step 2: simulation bringup, done, see 5.2 Test step 2: simulation bringup
- Step 3: commanding arms and grippers (single-joint arm move and gripper), which revealed the interpolation issue, see 5.3 Trajectory controller snaps instead of moving smoothly (issue)
- Step 3b: switch the trajectory controllers to
splinesand test in sim, done - Step 6 (hardware): H1 done, see 5.4 Hardware stage: plan, and H1 host checks; H0 waiting for answers
- Steps 4 (controller switching) and 5 (MoveIt): skipped for now by Manuel’s choice; going to the hardware stage next.
- Step 4: controller switching
- Step 5: MoveIt
- Step 6: real hardware
5.1 Test step 1: GPU and display
Goal: prove that a program inside the container can open an OpenGL window on the host’s screen and that it renders on the NVIDIA card, not on Intel/Mesa. This is the precondition for both launch files, which start RViz.
Safety check first: ip -br link showed can0 and can1 both DOWN, and no ROS processes were running in the container. Nothing could reach the motors.
How: glxgears/glxinfo aren’t installed in the container (mesa-utils is missing), so instead of installing anything we used RViz on its own (no robot, no controllers) as the OpenGL test, for 20 seconds:
docker exec -d openarm bash -c 'source /opt/ros/jazzy/setup.bash && timeout 20 rviz2 > /tmp/rviz_test.log 2>&1'
nvidia-smi # on the host, while RViz is open| Check | Result | OK? |
|---|---|---|
| Window opened | RViz started, no display/authorisation errors | ✅ |
| OpenGL version in RViz log | OpenGl version: 4.6 (GLSL 4.6), no MESA/iris messages | ✅ |
| Running on the NVIDIA GPU | nvidia-smi lists rviz2 using 21 MiB | ✅ |
| Shutdown | clean (SIGINT/SIGTERM from timeout) | ✅ |
The only other message, XDG_RUNTIME_DIR not set, is on the guide’s list of harmless messages (Guide §7).
Quick GPU check for RViz
While RViz is open,
nvidia-smion the host should listrviz2under Processes. If it doesn’t, RViz has fallen back to the Intel GPU.
5.2 Test step 2: simulation bringup
Goal: start the robot-only launch file with simulated motors and confirm the controllers, the hardware type, the joint-state topic and the RViz model (Guide §7A).
Working rule from here on
Manuel types and runs every command himself, to learn the procedure. Claude explains each step, gives the commands and what to expect, and checks the results for safety. Claude does not launch or stop anything.
What went wrong on the first attempt
Manuel had already started the launch file in his own terminal when Claude, without noticing, started a second copy in the background. The safety check Claude ran just before (ps aux | grep …) did show Manuel’s processes, but Claude went ahead anyway. Result: two /controller_manager nodes, two robot_state_publishers, two hardware interfaces with the same names. The second copy’s spawners failed (Controller already loaded, Failed to configure controller), and list_controllers showed a single unconfigured controller. ROS warned: “there are nodes in the graph that share an exact name”. No risk to the robot (mock hardware, CAN down), but the results were meaningless.
Clean-up (done by Manuel): Ctrl+C in his terminal; docker exec openarm pkill -INT -f "ros2 launch" for the background copy (-INT = the same signal as Ctrl+C); then checked with ps that only the ros2cli.daemon helper was left, and that can0/can1 were still DOWN.
Clean run
Terminal 1:
docker exec -it openarm bash
ros2 launch openarm_bringup openarm.bimanual.launch.py arm_type:=v10 use_fake_hardware:=trueTerminal 2 (checks):
| Check | Command | Expected | OK? |
|---|---|---|---|
| Controllers | ros2 control list_controllers | 5 controllers, all active | ✅ |
| Hardware type | ros2 control list_hardware_components | grep -E "name:|plugin|state" | two components, mock_components/GenericSystem, active | ✅ |
| Joint states | ros2 topic echo --once /joint_states | 16 joints (joint1–7 + finger_joint1 per arm), all 0.0 | ✅ |
| RViz | look at it | both arms fully drawn | ✅ |
Results reported by Manuel as all correct.
Pitfall: check nothing is running before you launch
Run
docker exec openarm bash -c 'ps aux | grep -E "ros2|rviz|controller|spawner|robot_state" | grep -v grep'first. Only aros2cli.daemonline is OK. Two launches at once give two controller managers, which conflict (Guide §6). On the real robot, two hardware interfaces on the same CAN bus would be dangerous.
What the 5 controllers are and where control runs
Arms and grippers each have their own
JointTrajectoryController(same type, different joints), so a gripper can be commanded without sending a goal for all 7 arm joints. The controllers and the hardware interface run on the laptop (inros2_control_node, 750 Hz). The loop that actually produces torque runs inside each motor on the robot (“MIT mode”: the laptop sends target,kp,kdevery cycle). See Guide §4.
5.3 Trajectory controller snaps instead of moving smoothly (issue)
issue/ros topic/ros2-control status/resolved pitfall
Symptom
Test step 3, part 1 (Guide §8.1): the left elbow (openarm_left_joint4) was sent to 0.5 rad with time_from_start: {sec: 3}. The goal was accepted, returned SUCCEEDED, and /joint_states showed 0.5 rad. But in RViz the arm didn’t move smoothly over 3 s, it snapped to the new position.
Why it matters
Guide V2.1 (§8.1, §12.3) told you to use a long time_from_start so that the first moves on the real robot are slow. If the controller jumps instead, the motor gets a sudden step in its target and, with the stiff MIT gains (kp = 60 on the elbow), moves at full speed, however long time_from_start is. Following the guide as written, the very first real motion would have been a jerk.
Diagnosis
- The bringup config (
openarm_bringup/config/controllers/openarm_bimanual_controllers.yaml) setsinterpolation_method: "none"on all four trajectory controllers (both arms, both grippers). The comment says it’s for streaming high-frequency commands. - Claude read the controller’s source (
ros2_controllers, branchjazzy,joint_trajectory_controller/src/trajectory.cpp, functionTrajectory::sample()). Withnone: before the first point’s time, the output is the position held when the goal arrived; between two points, the output is the next point. So a one-point goal = hold fortime_from_start, then jump. The guide’s claim that it “ramps linearly” was wrong. - Confirmed by Manuel in sim: the same move with
time_from_start: {sec: 5}did nothing for 5 s, then snapped at the momentSUCCEEDEDwas printed. Installed package:ros-jazzy-joint-trajectory-controller4.42.1-1noble.20260924.190511. - The MoveIt demo uses a different file (
openarm_bimanual_moveit_controllers.yaml) that doesn’t setinterpolation_method, so it gets the default,splines, and isn’t affected. - Not affected either: the ~2 s move to zero when the motors are switched on. That’s done by the hardware interface (
OpenArmHW), not by this controller.
Fix (applied and tested by Manuel)
Decision: switch the bringup trajectory controllers to
interpolation_method: splinesManuel’s decision. The setting is read-only at runtime, so it’s changed in the config file inside the container (4 lines,
"none"→"splines"), then the bringup is relaunched. Manuel applies it himself. Steps and checks: Guide V2.2 §4. The edit is tracked by git in/root/ros2_ws/src/openarm_ros2, sogit diffshows it andgit checkout -- <file>undoes it. Agit pull/re-clone would also undo it.
With positions-only goals, splines gives a linear move (constant speed, abrupt start/stop). Adding velocities: [0, …] to the points should give a cubic, eased move; still to verify.
Documentation: corrected in Guide V2.2 (new version, per the convention in 1.2 Versioning convention (agreed today)): §4 new subsection, §8.1, §12.1 (new pre-flight item), §12.3, §14, §16, and the Part II comparison table.
Fix applied and tested
All commands run by Manuel inside the container, with the bringup stopped:
| Step | Command | Result | OK? |
|---|---|---|---|
| Right file? | ros2 pkg prefix openarm_bringup | /root/ros2_ws/install/openarm_bringup | ✅ |
Old /ws workspace in the way? | ls -la /ws | empty folder (so the .bashrc line in 4.3 Extra line in `.bashrc` does nothing) | ✅ |
| Edit | sed -i 's/interpolation_method: "none"/interpolation_method: "splines"/' openarm_bringup/config/controllers/openarm_bimanual_controllers.yaml | — | |
| Exactly what changed | git diff | exactly 4 lines, "none" → "splines" (left arm, left gripper, right arm, right gripper); no other files modified | ✅ |
| Rebuild needed? | ls -l …/install/openarm_bringup/share/openarm_bringup/config/controllers/openarm_bimanual_controllers.yaml | symlink to the src file, so no rebuild | ✅ |
| Running setting | ros2 param get /<controller> interpolation_method for all 4 (after relaunch) | splines ×4 | ✅ |
| Arm motion | left joint4 → 0.5 rad, sec: 5, and back to 0 | moves right away and reaches the target in the time_from_start set | ✅ |
| Gripper motion | left gripper → 0.044 (open), sec: 2, and back to 0 | same: reaches the target in the time set | ✅ |
ros2only exists inside the containerRunning a
ros2 …command in a host terminal givesros2: command not found. Use adocker exec -it openarm bashshell.
Status: resolved (2026-10-09, evening). Re-check with ros2 param get before the first real-robot launch, since a git pull/re-clone of openarm_ros2 would undo it.
Pitfall: a long
time_from_startis not a safety measure on its ownIt only slows a move if the controller interpolates. Check
ros2 param get /left_joint_trajectory_controller interpolation_method(must besplines) before commanding real motors. And don’t trust the simulator to show you speed problems unless you watch for them: the mock hardware made the jump visible only as a “snap”.
5.4 Hardware stage: plan, and H1 host checks
log/testing topic/can topic/hardware
Plan for the hardware stage (Guide V2.2 §10–12), ordered from “nothing can move” to “the robot moves”:
| # | Sub-step | Motor power | Can it move? |
|---|---|---|---|
| H0 | Physical safety setup, plan for cutting power | off | no |
| H1 | Host checks: adapter, CAN tools, interfaces | off | no |
| H2 | Bring up CAN, power on the motors, discover / monitor (read-only) | on, torque off | only by hand |
| H3 | Check which arm is which, joint signs and the zero pose, by hand | on, torque off | only by hand |
| H4 | Decide on the container’s real-time flags (4.4 Container not created with real-time or CAN-admin flags) | off | no |
| H5 | First ROS launch, one arm only, then test shutdown behaviour with the arm held | on | yes: on start it moves to zero in ~2 s |
| H6 | First motions (Guide §12.3) | on | yes |
H1 results (run by Manuel on the host, motors unpowered):
| Check | Command | Result | OK? |
|---|---|---|---|
| Adapter model | lsusb | grep -i -E "peak|pcan" | 0c72:0011 PEAK System PCAN-USB Pro FD | ✅ |
| CAN tools on host | which openarm-can-cli candump | /usr/bin/openarm-can-cli, /usr/bin/candump | ✅ |
can0 / can1 | ip -details link show can0 (and can1) | both DOWN, can state STOPPED, error counters 0; driver pcan_usb_pro_fd; same USB parent device (3-2.4:1.0), i.e. the two channels of one adapter | ✅ |
| No LeRobot on the buses | ps aux | grep -i -E "lerobot|safe_teleop" | nothing | ✅ |
This answers follow-up 4.1 The CAN adapters are PEAK, not CANable/gs_usb: one PCAN-USB Pro FD, two channels, as Guide §22 says. mtu 16 means the interfaces are currently in classic-CAN mode; that’s normal while unconfigured. The LeRobot setup script (~/Documents/openarm_lerobot/setup_openarm_lerobot.sh) configures them as CAN-FD, 1 Mbit/s nominal / 5 Mbit/s data, with can0 = left and can1 = right, the same as openarm-can-cli can_configure’s defaults and ROS’s can_fd:=true.
5.5 H0 answers, and what really moves the robot (findings)
topic/hardware issue/ros status/open pitfall
H0 answers (Manuel):
- Power cut: by removing the motor power, from the laptop position.
- Base: not clamped or bolted.
- Space around the arms: clear.
- Holding the arm at shutdown: Manuel himself.
- Previous use: the robot has been moved before. Last time it snapped to zero positions that were completely off, and the robot fell over. Requirement from Manuel: every position command must be slow.
Requirement before any torque-on (H5): clamp the base
It already fell over once. Clamp or bolt it before the motors are ever enabled under ROS.
Finding 1: the ROS start-up move is not covered by the splines fix. From openarm_hardware/src/openarm_simple_hardware.cpp (on_activate() → return_to_zero(), commit b9d7a67):
enable_all(): torque on.- One full-stiffness command to position 0 for every arm joint (
{kp, kd, 0.0, 0, 0}, kp up to 70), then ~1 ms later it reads the current positions. A brief kick toward zero. - A 2 s linear ramp from the current position to 0 at full stiffness (200 steps × 10 ms, hard-coded, no launch argument).
- The gripper is commanded straight to a fixed position, with no ramp.
pos_commands_starts at 0 andwrite()sends it with full stiffness every cycle, so the driver relies on the arm having reached 0 before the controllers start.
With a wrong zero, step 3 drives the arm 2 s toward a pose that may be far away, maybe into the table or the torso. That matches what happened last time (to confirm which program was running then). The trajectory controller’s splines setting has no effect on this.
Finding 2: openarm-can-cli behaves differently from what Guide §10.5 says (source: enactic/openarm_can f340d4b, setup/cli/commands/):
| Command | What it actually does | Torque? |
|---|---|---|
discover | runs sudo ip link set … to reconfigure the interface at 1M, 5M, 8M and 10M data rates in turn, and only queries a parameter (MST_ID). can_configure defaults (1M/5M FD, sample points 0.75), as its output shows. Claude had read only part of the source. | no |
monitor | enable_all(), then only state requests (no position command), then disable_all() on exit. Decodes every motor as a DM4310: positions are right, but velocity and torque for joints 1–4 (DM8009/DM4340) are scaled wrong. | enabled, no command sent |
diagnose, motor_status | also call enable_all() | enabled |
Guide §10.5 says that during monitor the motors are “disabled = free to move”. In the code they’re enabled; with no position command sent they should produce no torque (⚠ not verified). For the first monitor run: motors freshly powered on, one arm, hold it.
Options for the start-up move (decision pending):
- A, recommended: patch
OpenArmHWso activation holds the current position (setspos_commands_to the measured positions, no kick, no ramp). The arm doesn’t move when ROS starts; it’s then moved to 0 with a slow trajectory goal (splines, e.g. 20–30 s). Can’t be tested in sim, because mock doesn’t useOpenArmHW. - B: patch it to a slower ramp (e.g. 15 s) and remove the kick. It still moves on start.
- C: no patch; rely on a verified zero and placing the arm at zero by hand. Doesn’t meet “every command slow”.
Guide corrections for §4, §10.5 and §12 to follow in a new version once these are tested.
Decision and more answers (Manuel)
- Power cut: a physical switch on the motor supply.
- The fall last time happened under a LeRobot script, not ROS. Per Guide §21, LeRobot and ROS use the same zero, stored in the motors, so a zero that was “completely off” for LeRobot is most likely off for ROS too. H3 (verify the zero) is mandatory.
- Start-up move: option B chosen, a slow ramp to zero.
The patch (written by Claude, to be reviewed and applied by Manuel)
File: Prandium/openarm_patches/0001-openarm_hardware-slow-return-to-zero.patch, against openarm_ros2 b9d7a67. It changes OpenArmHW::return_to_zero():
| Before (upstream) | After (patch) | |
|---|---|---|
| Reading the start position | first sends a full-stiffness command to 0, then reads | reads with a state request only, no command |
| Ramp | 2 s, linear | 10 s, eased (smoothstep: gentle start and stop), RETURN_TO_ZERO_SECONDS |
| Gripper | sent straight to its target | ramped from its current position, same target |
| Safety check (added) | none | if any arm joint is more than 0.5 rad (≈29°) from zero, it doesn’t move: logs which joints, disables the motors, and activation fails (MAX_START_OFFSET) |
| Logging | — | prints each joint’s start position |
| After the ramp | pos_commands_ left as they were | set to the zero pose, so write() can’t send a stale target after a re-activation |
Why 10 s and not longer: the hardware activation blocks the controller manager, and the controller spawners wait on a lock with a 20 s timeout. Upstream’s 2 s is no problem; 10 s leaves a margin. If the spawners do time out, the arm just holds zero and the controllers can be spawned by hand.
The safety check also guards against the zero problem above: with the arm hanging in the zero pose, a wrong motor zero shows up as a large start offset, and the arm won’t move.
git apply --check on the laptop’s copy of the repo (same commit): applies cleanly. Not compiled yet. Can’t be tested in sim (mock doesn’t use OpenArmHW); the first real test is H5.
Applied and built by Manuel inside the container (docker cp the patch to /root/, git apply, colcon build --symlink-install --packages-select openarm_hardware). Undo: git checkout -- openarm_hardware in /root/ros2_ws/src/openarm_ros2, then rebuild.
H2 part 1, CAN interfaces configured (Manuel, host, motor power off):
sudo ip link set canX type can bitrate 1000000 dbitrate 5000000 fd on
sudo ip link set canX upBoth can0 and can1: state UP, can <FD> state ERROR-ACTIVE, error counters 0, bitrate 1000000 sample-point 0.750, dbitrate 5000000 dsample-point 0.750, mtu 72 (CAN-FD frames). timeout 3 candump canX printed nothing (bus silent with motors off). The sample points match openarm-can-cli can_configure’s defaults (0.75/0.75, dsjw 2, from setup/cli/cli.hpp) and what LeRobot’s lerobot-setup-can gets (kernel defaults).
Why bit rates must be set by hand
CAN has no clock wire and no auto-negotiation: every device on a bus must already agree on the timing of each bit, or frames turn into errors and the bus shuts down (bus-off). The motors’ rates are stored in their firmware (1 Mbit/s for the arbitration phase, 5 Mbit/s for the CAN-FD data phase); Linux can’t ask them, so the adapter is told. The setting is lost on reboot or unplug.
H2 part 2, motors powered, discover (Manuel, host; torque stays off):
| Check | Result | OK? |
|---|---|---|
openarm-can-cli -i can0 discover | 8 motors: send 0x01–0x08 → receive 0x11–0x18, all at 5 Mbps (FD) (code 9) | ✅ |
openarm-can-cli -i can1 discover | same, 8 motors at 5 Mbps (FD) | ✅ |
| Interfaces afterwards | discover restored both to 1M/5M FD by itself; Manuel’s manual reset was redundant. Both UP, ERROR-ACTIVE, error counters 0 | ✅ |
H3, zero check by hand (Manuel; openarm-can-cli -i canX monitor -d 60000, run inside the container, which works because of --network host). Arms hanging straight down, grippers closed. Motors stayed limp while monitor had them enabled.
| ID | Joint | Left (can0) | Right (can1) |
|---|---|---|---|
| 0x01 | joint1 | −0.000 | +0.002 |
| 0x02 | joint2 | −0.009 | +0.005 |
| 0x03 | joint3 | −0.042 | −0.063 |
| 0x04 | joint4 | −0.002 | −0.025 |
| 0x05 | joint5 | −0.015 | +1.394 ❌ |
| 0x06 | joint6 | −0.078 | −0.003 |
| 0x07 | joint7 | −0.060 | +0.064 |
| 0x08 | gripper | +0.004 | +0.003 |
- Left arm =
can0, right arm =can1, as in Guide §22. - Left arm: zero OK. The largest offsets (0.04–0.08 rad, 2–4.5°) are on joints 3, 6 and 7, which gravity doesn’t hold in one exact position when hanging, so they’re within hand-placement precision. The patched start-up would move them to 0 slowly over 10 s.
- Right joint 5: zero wrong by 1.394 rad (≈ 80°). Manuel confirms it’s physically in the same pose as the left arm’s joint 5. With the patch, ROS start-up of the right arm would refuse to move (1.394 > 0.5). Without the patch, upstream would have swung the wrist ~80° in 2 s. Most likely cause: a LeRobot zeroing with that wrist rotated (the motor zero is shared, see Guide §21).
Pitfall:
openarm-can-cli set_zerowith no--idzeroes all 8 motors of that busIts default is
--arm(IDs 1–8).--idoverrides it (sourcesetup/cli/openarm_cli.cpp, installedopenarm-can-utils 1.4.0). To fix one motor:openarm-can-cli -i can1 set_zero --no-arm --id 5. It sends Disable → Set-Zero (0xFE) → Disable, so no torque.
H3b, right joint 5 re-zeroed (Manuel): joint placed by hand to match the left arm’s joint 5, then openarm-can-cli -i can1 set_zero --no-arm --id 5. monitor afterwards: 0x05 ≈ 0.00. Still ≈ 0.00 after a motor power cycle, so the zero is stored in the motor. (Optional end-stop symmetry check skipped.) Since the zero is shared, this also fixes it for LeRobot. ✅
The base can't be clamped (Manuel)
Accepted risk. Mitigations for H5–H6: ballast the base with heavy objects; first power-on with one arm only; arm starts hanging (near zero, so the start-up check passes and the ramp is short); only small, slow moves close to the hanging pose; motor-power switch within reach at all times.
5.6 H4: new container with real-time permissions
topic/docker topic/ros2-control log/testing
Why real-time is needed
The controller manager (ros2_control_node) runs one loop 750 times a second: read the motors over CAN → run the controllers → send the next MIT command to every motor. That leaves 1.33 ms per cycle. A normal Linux process is scheduled “fairly” (SCHED_OTHER): whenever RViz, the browser, Docker or the GPU driver want the CPU, the loop may have to wait. When it misses its 1.33 ms slot, the log says Overrun detected! … missed cycles: 2, which we saw regularly in sim.
- In sim that’s harmless: mock hardware has no physics.
- On the real robot, a late cycle means the motors get their next target late, and then a bigger step to catch up. With stiff MIT gains that’s felt as small jerks or buzzing. Motors that don’t get a frame for a while may also hit their CAN timeout, if one is configured.
Real-time scheduling (SCHED_FIFO) fixes that: a FIFO thread with a priority always runs before every normal process, as soon as it’s ready. ros2_control asks for SCHED_FIFO priority 50 for its loop automatically, and only needs permission. It also locks its memory (mlockall) so the loop never waits for memory to be paged in from disk.
Why Docker doesn’t allow it by default
Docker runs containers with as few permissions as possible, so a program in a container can’t harm the host:
- Capabilities. Linux splits root’s powers into ~40 “capabilities”. Docker gives containers only a small default set, and
CAP_SYS_NICE(raise priorities, use real-time scheduling) isn’t in it. Evenrootinside the container can’t useSCHED_FIFOwithout it. - Why that’s the safe default: a
SCHED_FIFOthread is never interrupted by normal programs. If it gets stuck in a loop, it can freeze a CPU core and make the whole host unresponsive. So Docker leaves this to the user to enable on purpose. - Limits (
ulimit).rtprio(the highest real-time priority a process may use) defaults to 0, i.e. none.memlock(how much memory a process may lock in RAM) defaults to a few MB, too little formlockall.
That’s why the controller manager printed “Could not enable FIFO RT scheduling” in the old container. The three flags give exactly those permissions:
| Flag | Gives |
|---|---|
--cap-add SYS_NICE | the right to raise priority / use SCHED_FIFO |
--ulimit rtprio=99 | allows RT priorities up to 99 (ros2_control uses 50) |
--ulimit memlock=-1 | no limit on locked memory |
The Linux kernel keeps its own safety net: by default real-time tasks may use at most 95 % of each second (/proc/sys/kernel/sched_rt_runtime_us = 950000), so the host can’t be frozen completely.
What was done (Manuel, host terminal)
Flags can only be set when a container is created (docker run), so a new container was made from a snapshot of the old one:
| Step | Command | Result |
|---|---|---|
| Current flags recorded | docker inspect openarm --format … | Network=host IPC=host, bind /tmp/.X11-unix, GPU request (--gpus all), CapAdd=null, Ulimits=null; env DISPLAY=:1, NVIDIA_DRIVER_CAPABILITIES=all, __NV_PRIME_RENDER_OFFLOAD=1, __GLX_VENDOR_LIBRARY_NAME=nvidia (plus image defaults) |
| Nothing running | docker exec openarm bash -c 'ps aux | grep …' | only ros2cli.daemon |
| Stop | docker stop openarm | — |
| Snapshot | docker commit openarm openarm-jazzy:rt | new image openarm-jazzy:rt (7.37 GB on disk, shares layers with :built) |
| Keep the old one | docker rename openarm openarm-deprecated | old container kept as a backup, untouched |
| New container | docker run -it --name openarm --network host --ipc host --gpus all -e NVIDIA_DRIVER_CAPABILITIES=all -e __NV_PRIME_RENDER_OFFLOAD=1 -e __GLX_VENDOR_LIBRARY_NAME=nvidia -e DISPLAY=$DISPLAY -v /tmp/.X11-unix:/tmp/.X11-unix --cap-add SYS_NICE --ulimit rtprio=99 --ulimit memlock=-1 openarm-jazzy:rt | — |
Checks in the new container:
| Check | Result | OK? |
|---|---|---|
ulimit -r; ulimit -l | 99, unlimited | ✅ |
git diff --stat in openarm_ros2 | same 3 files (splines config + the patch): the snapshot kept everything | ✅ |
nvidia-smi -L | RTX 5050 | ✅ |
ros2 pkg prefix openarm_hardware | /root/ros2_ws/install/openarm_hardware | ✅ |
Sim launch, grep -i -E "FIFO|RT scheduling|overrun" | Successful set up FIFO RT scheduling policy with priority 50. One overrun (1.51 ms) in the run, instead of regular ones | ✅ |
The remaining occasional overrun is expected on a normal (non-PREEMPT_RT) laptop kernel with RViz running (Guide §12.5). groups: cannot find name for group ID 992 on every shell is harmless (Guide §7).
Containers and images from now on
openarm(imageopenarm-jazzy:rt): the one to use. Daily commands unchanged:docker start openarm,docker exec -it openarm bash. Its main shell is thedocker run -itterminal; closing it stops the container.openarm-deprecated: the old container, kept as a backup until the robot has run well on the new one. Don’t start both at once: both use host networking and would fight over ROS and CAN.- Images:
openarm-jazzy:latest,:built,:rt.:rt=:built+ today’s changes.
This resolves 4.4 Container not created with real-time or CAN-admin flags (except NET_ADMIN, not needed: CAN is configured on the host).
Guide update pending (next version, after the hardware tests): §5 container command and image name, §10.5 (monitor enables motors), §11 (set_zero default = all 8), §4/§12 (the start-up patch). Done in Guide V3.1 (2026-10-10).
5.7 H5: first ROS launch on the real robot (left arm only)
log/testing topic/hardware topic/ros2-control
Setup. Right arm unpowered. There’s no single-arm launch file, so the right arm’s driver was pointed at a virtual CAN bus (vcan1), so its commands can’t reach anything physical and there are no bus errors from unanswered frames:
sudo modprobe vcan
sudo ip link add dev vcan1 type vcan
sudo ip link set vcan1 mtu 72 # CAN-FD frame size
sudo ip link set vcan1 upA motor that has never answered reads 0.0 in openarm_can (Motor constructor), so the right driver sees “all joints at 0”, passes the start-up check and ramps 0 → 0. Pre-checks: only the openarm container running, no ROS processes, splines ×4, can0 up at 1M/5M FD, left arm hanging, weight on the base.
Launch (container):
ros2 launch openarm_bringup openarm.bimanual.launch.py arm_type:=v10 use_fake_hardware:=false left_can_interface:=can0 right_can_interface:=vcan1 can_fd:=true 2>&1 | tee /root/h5_launch.logResults:
Successful set up FIFO RT scheduling policy with priority 50. All controllers reportUsing 'splines' interpolation method.- Both hardware components loaded
OpenArmHW(CAN=can0, arm_prefix=left_andCAN=vcan1, arm_prefix=right_). The right arm was activated first (start values all+0.000, 10 s ramp,Reached zero position), then the left (lines not captured in the paste; still to check in/root/h5_launch.log). - Nothing visibly moved. Expected: the left arm hung at zero, with the largest offset ≈ 0.08 rad (4.5°) on the wrist, spread over 10 s. Manuel stayed clear of the robot.
- All controllers spawned and activated despite the ~20 s of activation ramps.
- Constant overruns: each loop took 2.0–2.6 ms instead of 1.33 ms (read 1.25–1.57 ms, write 0.5–0.75 ms), so it actually ran at about 400–450 Hz. The main cause is the right driver waiting for replies that never come on
vcan1(recv_allwaits up to 500 µs inread()and 100 µs inwrite()). Not dangerous (trajectories are timed by the clock), but to measure again with both arms powered.
Pitfall: pressing Ctrl+C several times
Manuel stopped the launch with Ctrl+C ×3.
ros2 launchescalates on repeated Ctrl+C (SIGINT → SIGTERM → SIGKILL). A killed controller manager never runs the driver’son_deactivate(motors off), so the motors may keep holding their last command with torque on. Press Ctrl+C once and wait for the shutdown messages; only repeat if nothing happens for ~10 s. Here the arm was hanging at zero, so it was harmless; the motor power was switched off and on afterwards to get a clean state.
Still unknown: what Ctrl+C does to the motors on this setup (torque off, or holding). Test it next time with the arm at the zero pose (nothing to fall).
Correction: the left arm was not powered during H5 (or at the start of the H6 relaunch)
Found while preparing H6. The left arm’s start values in
/root/h5_launch.logwere exactly+0.000on all 7 joints, whilemonitorhad read −0.042 / −0.078 / −0.060 on joints 3/6/7 in the same pose.0.0is the default of a motor that has never replied (openarm_canMotorconstructor). Manuel confirmed the left arm’s power was off. So H5 tested nothing on the real arm: “nothing moved” because nothing was powered. The constant overruns fit too: both drivers waited the full reply timeout every cycle.Manuel then switched the left arm on while ROS was running.
/joint_statesimmediately showed real values matchingmonitor(e.g. joint3 −0.0422, joint6 −0.0784, joint7 −0.0597, non-zero efforts). But the motors stayed disabled (the enable command was sent at start-up, when they were off), and ROS’s stored targets were all 0 (from the empty readings). Inconsistent state, so no goals were sent; the launch was to be stopped and redone with the arm powered first.
Flaw in patch 0001, and patch 0002
The start-up check in patch 0001 compared the read positions with zero, but a motor that hasn’t replied reads 0.0, so the check passed without meaning anything. If the arm had been far from zero (and powered just after the read), it would have been driven to 0 with no ramp: the original snap.
Patch 0002 (Prandium/openarm_patches/0002-openarm_hardware-check-motors-replied.patch, applies on top of 0001): after the torque is switched on, every arm motor and the gripper must report is_enabled() with no error code (openarm_can sets that flag only from the motor’s own state replies, dm_motor_device.cpp), otherwise it logs which ones, switches the motors off and refuses to start, the same as the distance check. Dry-run on a clean copy of the repo + 0001: applies cleanly.
Consequence: an arm on the virtual vcan1 bus (no replies) now fails start-up by design, so the one-arm trick needs revisiting.
Patch 0002 applied and built by Manuel (docker cp, git apply, colcon build --symlink-install --packages-select openarm_hardware): Finished <<< openarm_hardware, only the expected on_init(HardwareInfo) deprecation warning.
Pitfall: Ctrl+C on
ros2 launch … | tee filehides the shutdownCtrl+C goes to every program in the pipe.
teeexits at once, so the prompt returns while ROS is still shutting down out of sight, and programs writing to the broken pipe can be killed before switching the motors off. Usetee -i(ignore interrupts) so the output keeps flowing until ROS has exited. Node messages are also saved in/root/.ros/log/.
Pitfall: power the arms before launching ROS, never while it's running
The driver sends “torque on” once, at start-up. An arm powered later stays limp while ROS believes it’s in control, with stored targets that may be far from the real pose.
5.8 H5 (for real): both arms powered, both patches
log/testing topic/hardware topic/ros2-control
Option A (Manuel’s choice): both arms powered before the launch, hanging at zero; only the left arm is to be commanded. Previous launch had stopped cleanly. Launch (container):
ros2 launch openarm_bringup openarm.bimanual.launch.py arm_type:=v10 use_fake_hardware:=false left_can_interface:=can0 right_can_interface:=can1 can_fd:=true 2>&1 | tee -i /root/h5b_launch.logStart-up (from the log), all real readings, all inside the 0.5 rad limit, no no reply / rad from zero errors:
| Joint | Right (can1, activated first) | Left (can0) |
|---|---|---|
| 1 | −0.012 | −0.000 |
| 2 | +0.016 | −0.009 |
| 3 | −0.009 | −0.014 |
| 4 | −0.026 | −0.002 |
| 5 | −0.004 (re-zeroed in H3b ✅) | −0.015 |
| 6 | +0.053 | −0.022 |
| 7 | +0.060 | −0.015 |
Each arm: Reached zero position ~10.5 s after its start. (Values differ a little from the H3 monitor snapshot because the arms were re-hung by hand in between.)
Checks: both hardware components openarm_hardware/OpenArmHW, active. All 5 controllers active. /joint_states shows real values for both arms.
Observation: the joints stop 0.01–0.02 rad short of zero. After the ramp, e.g. left joint6 −0.0216, joint7 −0.0143, joint3 −0.0132, i.e. almost exactly their start values. Right joint6 went +0.053 → +0.0147. The reported efforts match the MIT formula kp·(target − position) (target 0): left joint3 70 × 0.0132 ≈ 0.92 Nm (reported 0.87), left joint6 10 × 0.0216 ≈ 0.22 Nm (reported 0.20), right joint4 60 × 0.012 ≈ 0.72 Nm (reported 0.68). So the motors are pushing toward zero, but a small error × the gain gives too little torque to overcome gearbox friction. It’s normal for a PD controller with no integral term (Guide §4: no gravity compensation, no feed-forward). Expect commanded moves to land ~0.01–0.02 rad short on the wrist joints.
5.9 H6: first commanded moves, and the joint direction map
log/testing topic/hardware topic/ros2-control
First moves (Manuel): left joint7 → +0.2 rad over 10 s, and left joint1 → +0.3 rad. Both moved slowly and matched RViz, but went “backwards”, not “forwards”. That’s correct: ”+” is defined by each joint’s axis in the model, not by “forward”. Matching RViz is the real test, since it shows the motor direction and the model agree.
Joint direction map (OpenArm v1.0, from the model)
How it was worked out: openarm_description at commit c103041 (the container’s version) was expanded with xacro (arm_type:=v10 bimanual:=true) on the laptop, in Claude’s scratch folder. Each joint’s axis was then transformed into the robot’s frame at the zero pose (arms hanging), and the direction the hand moves for a small + rotation was computed. Both of Manuel’s tests (left joint1, left joint7 = backward) agree with it.
Robot frame: forward = the direction the robot faces (the way the elbows bend), left/right = the robot’s own left/right, up. Directions are for small moves from the hanging zero pose; once other joints have moved, each axis moves with the arm.
| Joint | What it is | Left arm: + moves… | Right arm: + moves… |
|---|---|---|---|
| joint1 | shoulder, swing forward/back | hand backward | hand forward |
| joint2 | shoulder, swing sideways | hand inward (toward the body) | hand outward (away from the body) |
| joint3 | upper-arm twist | clockwise seen from above (with the elbow bent, the forearm swings toward the robot’s right) | same: clockwise seen from above |
| joint4 | elbow | forearm forward/up (elbow bends) | same: forward/up |
| joint5 | forearm twist | clockwise seen from above | same |
| joint6 | wrist, sideways | hand outward | hand inward |
| joint7 | wrist, forward/back | hand backward | hand forward |
| finger_joint1 | gripper (metres) | 0 = closed, 0.044 = open | same |
Pattern: the arms are mirror images, so joint1, joint2, joint6 and joint7 have opposite meanings on the two arms. joint3 and joint5 turn the same way in the room (clockwise from above), so in body terms they’re mirrored (inward on one arm = outward on the other). joint4 is the same on both.
The joint-limit table in Guide §3 is wrong for joint1 (left) and joint2 (both arms)
Guide §3 copies
config/arm/joint_limits.yaml. But the model adds a per-arm offset and mirroring for joints 1 and 2 (urdf/arm/openarm_arm.xacro: joint1 limits- 2.094on the left; joint2± π/2plus mirroring), so the real limits are:
Joint Guide §3 (yaml) Left arm (actual) Right arm (actual) joint1 −1.396 … +3.491 −3.491 … +1.396 (forward to 200°, back to 80°) −1.396 … +3.491 (back to 80°, forward to 200°) joint2 −1.745 … +1.745 −3.316 … +0.175 (outward to 190°, inward only 10°) −0.175 … +3.316 (inward only 10°, outward to 190°) joints 3–7 as in §3 same as §3 same as §3 So e.g. left
joint2= +1.0 rad, “inside the limits” per the guide, actually drives the arm into the body. The trajectory controller does not enforce limits (Guide §12.3). Always check targets against this table, per arm.
Guide update pending: add this direction map, and replace the §3 limit table with the per-arm one. Done in Guide V3.1 (2026-10-10).
6. Lessons and pitfalls from today
Pitfall: charger and the NVIDIA GPU
Changing power source (charger in or out) crashed the NVIDIA driver (Xid 154). Keep the charger plugged in. Details in 3. Issue: NVIDIA GPU crashed (Xid 154).
Pitfall:
docker runis not how you get back into the container
docker runmakes a new container, without anything you installed or built in the old one. Usedocker start openarm(if stopped) and thendocker exec -it openarm bash. See 1.1 What changed.
Pitfall: two launches at once
Check with
psthat nothing is running before launching. See 5.2 Test step 2: simulation bringup.
Pitfall:
time_from_startdoesn't slow a move withinterpolation_method: noneThe joint waits, then jumps. Check the setting is
splinesbefore real-robot tests. See 5.3 Trajectory controller snaps instead of moving smoothly (issue).
Convention: guide versions are never overwritten
7. Next session
State at end of day: both arms tested on real hardware with patches 0001 + 0002, splines, real-time container openarm (image openarm-jazzy:rt). Manuel to stop for the night: arms back to zero, Ctrl+C once on the launch (watch for Deactivating OpenArm V10... ×2 = the shutdown test), motor power off.
After a reboot: xhost +SI:localuser:root → docker start openarm (the new one, not openarm-deprecated) → can0/can1 set up again (1M/5M FD, see H2) → power the arms → launch with tee -i.
To do, in order:
- Record the shutdown-test result (H5d): did
Deactivating OpenArm V10...appear twice, and did the arms go limp? - Test every joint of both arms once: ≤0.3 rad, 10 s, one at a time, back to zero each time. Check the direction against RViz and against Joint direction map (OpenArm v1.0, from the model); check each target against the per-arm limits in that section (joint2 especially).
- Grippers: 0.022 (half) first, then 0.044 / 0.0 (Guide §12.3).
- Look at the overruns now that both arms answer (
grep -c "Overrun detected" <launch log>, and the read/write times). -
Guide V2.3→ done as Guide V3.1 (2026-10-10): container (openarm-jazzy:rt, RT flags), §10.5 (monitorenables motors), §11 (set_zerodefault = all 8 motors), §4/§12 (the start-up patches, power before launch,tee -i), §3 (per-arm limits) + the direction map. - Later: steps 4 (controller switching) and 5 (MoveIt), skipped today; delete
openarm-deprecatedonce the new container has proven itself.