Robotics · Control study

RoboLLM: diagnosing failed reaching, measuring a repair

Correct targeting was not enough. Repairing the simulated arm’s geometry, reset and gravity support improved reaching from 0/100 to 35/100 on fresh paired scenes, with substantial limits.

On this page

The failure: a correct plan can drive the wrong action

A manipulation system must preserve an object’s identity from selection through execution and evaluation. RoboLLM’s executor grounded the requested object, then replaced its position with the first object in the scene when preparing controller input. Its pick/place checks also inspected that first object. The interfaces appeared connected, but their meanings differed.

A regression fixture puts the requested object to the right of the end effector and the first object to its left. The old executor commands motion left; the repaired executor commands motion right. This test isolates the interface defect without depending on whether the simulated arm can complete a physical task.

The repair: retain identity and read the live position

The executor now carries the grounded object’s name, resolves its current position at every step and scores that same object. A task already satisfied before acting is recorded separately and earns no new success credit. Failures and episode endings stop later actions; the environment’s original time limit is restored even if execution raises an exception.

Executor implementation · Target and lifecycle regressions. Pick/place also retain the selected source and check relative lift, gripper state and release. Those checks remain heuristics; the experiment below evaluates reaching only.

What counts as success

The custom Franka-inspired arm runs in MuJoCo with three randomized objects. Success requires the end-effector site to become strictly closer than 3 cm to the requested object’s current body center after an action, within 200 steps at 20 Hz. The condition must not hold initially. All scenes remain in the denominator.

This measures instantaneous proximity, with no grasp, orientation or stable hold requirement. There is no settling phase, and objects may move. The traces therefore retain the target’s motion as well as the arm’s. The simulation is not a calibrated physical-robot result.

First study: the frozen four-control comparison

The protocol and controller code were committed before evaluating seeds 260926000–260926099. Development used seeds 17 and 18. The requested target alternates between scene indices one and two; object zero is always a distractor. Each of the four controls starts from an identical, checked scene for each seed.

Every control keeps the gripper open and uses the same executor, live-target success predicate and stopping rule. Grounding reads privileged simulator metadata; there is no VLM inference or learned policy in this study. The evaluation scenes were held out from this repair’s development, not drawn from an independent robotics benchmark.

Results: correct targeting is not sufficient for reaching

Action target / controlSuccessesWilson 95% intervalMean final distance
Requested object0 / 1000–3.70%33.13 cm
First object instead0 / 1000–3.70%35.24 cm
Random translation0 / 1000–3.70%71.14 cm
Zero translation0 / 1000–3.70%67.81 cm

All 400 episodes reached the 200-step limit. None started inside the success threshold. The repaired controller’s mean final distance was 33.13 cm, against a 3 cm criterion. Its smaller mean distance than the first-object control does not establish successful reaching: both failed on all 100 paired scenes, giving a paired success-rate difference of zero.

These are new measurements from the frozen code revision, not a reproduction of the old one-object benchmark. The historical 20/100 scripted result used a different task and lacks complete episode provenance, so it is not a comparable before/after baseline.

Raw traces, summary and replay record retain every scene, action, target identity, position and outcome. An exact rerun on the recorded local runtime reproduced all semantic fields. That replay checks repeatability; it adds no independent trials. Wilson intervals describe the individual controls, not the paired difference or performance on real robots.

What the first study could not explain

The identity bug was fixed and tested, but motor control still failed. Every corrected run in that first study ended short along the X axis. Across those runs, positive X commands reach their maximum in 99.4% of all steps. The requested target’s displacement stays below 0.117 mm across all four controls, so moving targets do not explain the failure. These are observations from the traces, not an identified cause.

That result did not distinguish arm geometry, actuator behavior, inverse kinematics or controller settings as the cause. The follow-up therefore added joint, force, contact and Jacobian diagnostics on separate development scenes before another evaluation.

Follow-up: repair the simulated arm before more training

The original reset placed the first arm link 10 mm into the table and folded two other links into one another. The positional Jacobian was full rank at reset; initial singularity did not explain those collisions. Without gravity support, zero incremental commands also allowed the arm to drift.

The repair shortens two collision capsules while preserving the robot’s original masses, inertias and kinematic link lengths. It uses a collision-free home, aligns the actuator targets with that pose and adds ideal gravity compensation to robot bodies. Objects retain gravity; collisions remain active. The proportional policy, inverse kinematics, gain, timing and success threshold are unchanged.

Ideal MuJoCo body gravity compensation supports the robot’s weight in simulation; it is not a validated hardware controller. Twenty-six physical tests check initialization, inertial preservation, position holding, falling objects, retained collisions and nearby reachable targets. Eight development configurations cross the collider, reset and gravity changes on two scenes. These diagnostics guide the repair but do not isolate each component’s effect over the whole workspace.

Fresh comparison: 35/100 versus 0/100

The new protocol and implementation were committed before evaluating seeds 260927000–260927099. Development used seeds 17, 18 and 101–110. Each fresh object scene runs both the frozen original system and the repaired configuration, with exact named-target grounding and the same strict post-action 3 cm criterion. No scene is filtered or resampled.

SystemSuccessesWilson 95% intervalMean final distance
Frozen original0 / 1000–3.70%32.65 cm
Repaired configuration35 / 10026.36–44.75%13.01 cm

35 paired scenes succeed only after the repair; 65 fail in both configurations. The observed difference is +35 percentage points. There are 200 episodes and 33,927 control steps, with no initially satisfied scene. A second execution reproduces every recorded semantic field exactly on the same runtime. That verifies repeatability, not another 100 independent trials.

This compares a geometry/reset/gravity-support bundle. The objects start identically, but the arm does not: the new home reduces mean initial target distance from 52.46 to 40.67 cm. The score cannot be attributed to one component alone. The Wilson intervals describe each configuration, not the paired difference. Inspect all trajectories, model snapshots, summary and replay record.

Why proximity is not reliable manipulation

The repaired arm moves the target by more than 3 cm during 42 of the 100 episodes, including 6 of the 35 successful ones. The largest displacement is 1.79 m in a failed episode. Contacts can disturb or eject objects; no such episode was removed from the result. Entering a sphere around a moving object is a much weaker requirement than reaching safely and holding position.

This experiment uses privileged simulator state and scripted actions. It does not demonstrate visual grounding, a learned policy, a secure grasp, stable placement or physical-robot performance. Mock planning and SAC training code remain separate from the measured result.

Reproduction and the next useful question

The recorded runtime is Python 3.11.14, MuJoCo 3.14.0, NumPy 2.4.2 and Gymnasium 1.3.0 on macOS ARM64, without Torch, rendering or a GPU. The repair method and reproduction guide identify the frozen revision, commands and exact metric. The first study’s protocol and its negative results remain available separately.

Focused GitHub validation passed 230 tests and both development simulation smokes. The code and evidence are on a review branch; the main branch is not merged and the broader training/rendering suite is not certified.

A separate stationary-goal study now examines approaching a fixed point above the object and holding there, with explicit object-motion and contact constraints. Its goal and success contract differ from the live-center proximity task reported here, so the scores are not directly comparable. Both evaluation sets in this article remain observed evidence; reusing them is reproduction, not a new held-out result.

Implementation and regression tests were developed with Codex assistance. The inspected code, recorded executions and explicit limits support the result.

Project scope and historical evidence →

Keep exploring

Working on a similar problem? Let’s discuss it.

Get new articles via RSS