The failure: a correct plan can drive the wrong action
A manipulation system must preserve an object’s identity from selection through execution and evaluation. RoboLLM’s executor grounded the requested object, then replaced its position with the first object in the scene when preparing controller input. Its pick/place checks also inspected that first object. The interfaces appeared connected, but their meanings differed.
A regression fixture puts the requested object to the right of the end effector and the first object to its left. The old executor commands motion left; the repaired executor commands motion right. This test isolates the interface defect without depending on whether the simulated arm can complete a physical task.
The repair: retain identity and read the live position
The executor now carries the grounded object’s name, resolves its current position at every step and scores that same object. A task already satisfied before acting is recorded separately and earns no new success credit. Failures and episode endings stop later actions; the environment’s original time limit is restored even if execution raises an exception.
Executor implementation · Target and lifecycle regressions. Pick/place also retain the selected source and check relative lift, gripper state and release. Those checks remain heuristics; the experiment below evaluates reaching only.
What counts as success
The custom Franka-inspired arm runs in MuJoCo with three randomized objects. Success requires the end-effector site to become strictly closer than 3 cm to the requested object’s current body center after an action, within 200 steps at 20 Hz. The condition must not hold initially. All scenes remain in the denominator.
This measures instantaneous proximity, with no grasp, orientation or stable hold requirement. There is no settling phase, and objects may move. The traces therefore retain the target’s motion as well as the arm’s. The simulation is not a calibrated physical-robot result.
First study: the frozen four-control comparison
The protocol and controller code were committed before evaluating seeds 260926000–260926099. Development used seeds 17 and 18. The requested target alternates between scene indices one and two; object zero is always a distractor. Each of the four controls starts from an identical, checked scene for each seed.
- Requested object: the existing proportional controller follows the named live target, with gain 8.
- First object instead: the same controller follows object zero, while success is still measured against the requested target. This isolates incorrect action targeting, rather than reproducing every old executor behavior.
- Random translation: seeded random XYZ actions provide a baseline.
- Zero translation: zero XYZ commands reveal possible passive successes from simulator dynamics.
Every control keeps the gripper open and uses the same executor, live-target success predicate and stopping rule. Grounding reads privileged simulator metadata; there is no VLM inference or learned policy in this study. The evaluation scenes were held out from this repair’s development, not drawn from an independent robotics benchmark.
Results: correct targeting is not sufficient for reaching
| Action target / control | Successes | Wilson 95% interval | Mean final distance |
|---|---|---|---|
| Requested object | 0 / 100 | 0–3.70% | 33.13 cm |
| First object instead | 0 / 100 | 0–3.70% | 35.24 cm |
| Random translation | 0 / 100 | 0–3.70% | 71.14 cm |
| Zero translation | 0 / 100 | 0–3.70% | 67.81 cm |
All 400 episodes reached the 200-step limit. None started inside the success threshold. The repaired controller’s mean final distance was 33.13 cm, against a 3 cm criterion. Its smaller mean distance than the first-object control does not establish successful reaching: both failed on all 100 paired scenes, giving a paired success-rate difference of zero.
These are new measurements from the frozen code revision, not a reproduction of the old one-object benchmark. The historical 20/100 scripted result used a different task and lacks complete episode provenance, so it is not a comparable before/after baseline.
Raw traces, summary and replay record retain every scene, action, target identity, position and outcome. An exact rerun on the recorded local runtime reproduced all semantic fields. That replay checks repeatability; it adds no independent trials. Wilson intervals describe the individual controls, not the paired difference or performance on real robots.
What the first study could not explain
The identity bug was fixed and tested, but motor control still failed. Every corrected run in that first study ended short along the X axis. Across those runs, positive X commands reach their maximum in 99.4% of all steps. The requested target’s displacement stays below 0.117 mm across all four controls, so moving targets do not explain the failure. These are observations from the traces, not an identified cause.
That result did not distinguish arm geometry, actuator behavior, inverse kinematics or controller settings as the cause. The follow-up therefore added joint, force, contact and Jacobian diagnostics on separate development scenes before another evaluation.
Follow-up: repair the simulated arm before more training
The original reset placed the first arm link 10 mm into the table and folded two other links into one another. The positional Jacobian was full rank at reset; initial singularity did not explain those collisions. Without gravity support, zero incremental commands also allowed the arm to drift.
The repair shortens two collision capsules while preserving the robot’s original masses, inertias and kinematic link lengths. It uses a collision-free home, aligns the actuator targets with that pose and adds ideal gravity compensation to robot bodies. Objects retain gravity; collisions remain active. The proportional policy, inverse kinematics, gain, timing and success threshold are unchanged.
Ideal MuJoCo body gravity compensation supports the robot’s weight in simulation; it is not a validated hardware controller. Twenty-six physical tests check initialization, inertial preservation, position holding, falling objects, retained collisions and nearby reachable targets. Eight development configurations cross the collider, reset and gravity changes on two scenes. These diagnostics guide the repair but do not isolate each component’s effect over the whole workspace.
Fresh comparison: 35/100 versus 0/100
The new protocol and implementation were committed before evaluating seeds 260927000–260927099. Development used seeds 17, 18 and 101–110. Each fresh object scene runs both the frozen original system and the repaired configuration, with exact named-target grounding and the same strict post-action 3 cm criterion. No scene is filtered or resampled.
| System | Successes | Wilson 95% interval | Mean final distance |
|---|---|---|---|
| Frozen original | 0 / 100 | 0–3.70% | 32.65 cm |
| Repaired configuration | 35 / 100 | 26.36–44.75% | 13.01 cm |
35 paired scenes succeed only after the repair; 65 fail in both configurations. The observed difference is +35 percentage points. There are 200 episodes and 33,927 control steps, with no initially satisfied scene. A second execution reproduces every recorded semantic field exactly on the same runtime. That verifies repeatability, not another 100 independent trials.
This compares a geometry/reset/gravity-support bundle. The objects start identically, but the arm does not: the new home reduces mean initial target distance from 52.46 to 40.67 cm. The score cannot be attributed to one component alone. The Wilson intervals describe each configuration, not the paired difference. Inspect all trajectories, model snapshots, summary and replay record.
Why proximity is not reliable manipulation
The repaired arm moves the target by more than 3 cm during 42 of the 100 episodes, including 6 of the 35 successful ones. The largest displacement is 1.79 m in a failed episode. Contacts can disturb or eject objects; no such episode was removed from the result. Entering a sphere around a moving object is a much weaker requirement than reaching safely and holding position.
This experiment uses privileged simulator state and scripted actions. It does not demonstrate visual grounding, a learned policy, a secure grasp, stable placement or physical-robot performance. Mock planning and SAC training code remain separate from the measured result.
Reproduction and the next useful question
The recorded runtime is Python 3.11.14, MuJoCo 3.14.0, NumPy 2.4.2 and Gymnasium 1.3.0 on macOS ARM64, without Torch, rendering or a GPU. The repair method and reproduction guide identify the frozen revision, commands and exact metric. The first study’s protocol and its negative results remain available separately.
Focused GitHub validation passed 230 tests and both development simulation smokes. The code and evidence are on a review branch; the main branch is not merged and the broader training/rendering suite is not certified.
A separate stationary-goal study now examines approaching a fixed point above the object and holding there, with explicit object-motion and contact constraints. Its goal and success contract differ from the live-center proximity task reported here, so the scores are not directly comparable. Both evaluation sets in this article remain observed evidence; reusing them is reproduction, not a new held-out result.
Implementation and regression tests were developed with Codex assistance. The inspected code, recorded executions and explicit limits support the result.