Problem and scope
A successful plan is only useful if the robot can execute it. RoboLLM separates instruction parsing, object grounding and motor control so each component can be inspected independently. Its existing mock backend makes integration tests inexpensive; it does not measure a vision-language model’s understanding of a camera image.
Task progression
| Task | What it exercises |
|---|---|
| Pick and place | Grasping, transport and release |
| Color-conditioned pick | Object selection before manipulation |
| Stacking | Placement and sequential precision |
| Sorting | Repeated object-to-zone assignment |
| Language instructions | Composition of task conditions |
| Move to | Reaching as a simpler control baseline |
Implemented and experimental components
- Environment: a Franka-inspired seven-joint arm in MuJoCo with a tabletop, randomized objects and bounded episodes.
- Tested planning path: MockVLM performs keyword matching; SimGrounder reads privileged simulator state. Passing these tests is not visual grounding accuracy.
- Motor control: scripted controllers support pipeline tests; a bounded model-based joint planner is evaluated separately on stationary reaching. SAC training and checkpoint evaluation remain separate from these results.
- Extension points: a Transformers VLM adapter exists. DINOv2-based visual grounding remains a proposed extension in the reviewed source.
Planner and grounding source · Hierarchical executor · Policy implementations
Stationary reaching and holding · 26 September 2026
A new controller comparison fixes the goal 10 cm above the object’s initial top. Success requires strict 2 cm proximity for one second, a surface-motion bound of at most 5 mm for every object, and no forbidden generated contacts involving moving robot bodies. Execution is checked at 500 Hz. Both controllers start from identical complete simulator states with the same repaired physics and fixed goal.
| Controller on 100 fresh paired scenes | Successes | Wilson 95% interval |
|---|---|---|
| Direct DLS | 29 / 100 | 21.01–38.54% |
| Collision-aware joint plan | 88 / 100 | 80.19–93.00% |
Both succeed on 29 scenes, the joint planner alone on 59, and neither on 12. Direct DLS fails on 71 forbidden-contact episodes; the planner fails to find an accepted path on 12. A bounded search failure does not prove that a goal is unreachable. This fixed-goal task differs from the earlier moving-center proximity criterion; its score is not comparable with the earlier 35/100.
The 200 episodes contain 253,150 recorded samples. Independent checking confirms the holding condition, and the full semantic record and provenance replay exactly. The largest observed all-object motion bound remains below 0.118 mm. Privileged state, ideal gravity support, discrete checks and collision-pair exclusions limit the claim; this is not learned, hardware or continuously certified safe manipulation.
Read the contract, controllers and results · Frozen method · Raw evidence and replay record.
Focused GitHub validation passed 293 tests and three development simulation smokes. The source remains on a review branch; these checks do not certify the broader training/rendering suite.
Earlier proximity studies · 26 September 2026
Initial study: the executor could select the requested object but drive toward the first object in the scene. Retaining the named target through control and scoring fixed that defect, but the corrected controller and each of three controls still scored 0/100. Its mean final distance was 33.13 cm against a 3 cm threshold. The 400 original traces and replay record remain available.
Follow-up: development diagnostics identified penetrating arm colliders and gravity-related drift. Two collider corrections preserve the original inertials and kinematics; a collision-free home, aligned actuator targets and ideal robot-only gravity support change the simulated system. The gain-8 policy and strict 3 cm criterion remain unchanged.
| System on 100 fresh paired scenes | Successes | Wilson 95% interval |
|---|---|---|
| Frozen original system | 0 / 100 | 0–3.70% |
| Repaired geometry/reset/support bundle | 35 / 100 | 26.36–44.75% |
Thirty-five pairs succeed only after the repair; 65 fail in both systems. The changed home also starts closer on average, so this measures the complete system bundle. The 200 episodes replayed exactly on the recorded runtime; replay adds no independent scenes.
Limits: success is instantaneous proximity to the live target body center, without dwell, grasp or stable placement. The target moves more than 3 cm in 42 repaired episodes, including six of the 35 successes. These outcomes remain in the score. Ideal gravity support is a simulation assumption. The study does not establish safe manipulation, learned-policy performance, VLM understanding or hardware capability.
Read the follow-up and its limits · Frozen method · Raw evidence and replay record.
Focused GitHub validation passed 230 tests and both development simulation smokes. This is a review-branch delivery, not a main-branch merge or a pass of the repository’s full training/rendering suite.
Historical baseline evidence
The archived evaluation dated 22 February 2026 records 100 episodes per task/policy pair, with a 200-step limit. The intervals below are Wilson 95% intervals recalculated from the recorded counts (z = 1.96), not new simulation results.
| Task | Policy | Successes | 95% interval | Mean return |
|---|---|---|---|---|
| Pick & place | Random | 0 / 100 | 0–3.70% | −142.37 |
| Pick & place | Scripted | 0 / 100 | 0–3.70% | −53.17 |
| Move to | Random | 0 / 100 | 0–3.70% | −715.76 |
| Move to | Scripted | 20 / 100 | 13.34–28.88% | −260.33 |
| Language task | Random | 4 / 100 | 1.57–9.84% | −955.20 |
Language-task success checks final scene predicates, which can already hold at reset; 4/100 does not demonstrate language understanding. The other archived random baselines—color pick, stacking and sorting—each report 0/100 successes. A dash or an absent run is not a zero result. Less-negative shaped return does not imply a multiplicative improvement or successful manipulation.
The original archive does not contain per-episode traces or complete environment/version provenance, and the evaluator did not seed the random action space. These records establish what was reported, not an exactly reproducible run or a learned-policy benchmark.
Current test pipeline
The hierarchical pipeline below remains separate from the stationary study. That controller comparison uses named simulator targets directly; its score does not validate the language planner or the full instruction-to-action chain.
Reproduction and next investigation
The stationary-goal protocol and reproduction guide specify the fixed task, two controllers and recorded runtime. The earlier physical-system comparison and four-control study retain their own protocols and results.
Next: investigate the bounded planner’s 12 no-plan cases without relabeling them unreachable or reusing observed scenes as a fresh holdout. Any planner changes need separate development and newly reserved evaluation scenes. Broader repository CI remains a separate requirement before a main-branch release; grasping, placement and language understanding still need their own evidence.