Language × Robotics

RoboLLM

A MuJoCo testbed for decomposing instructions, grounding objects and executing manipulation primitives. Recorded studies trace interface and physical-model defects, then compare controllers on a stationary-goal reaching task.

MuJoCo · PyTorch · SAC · Simulation

Problem and scope

A successful plan is only useful if the robot can execute it. RoboLLM separates instruction parsing, object grounding and motor control so each component can be inspected independently. Its existing mock backend makes integration tests inexpensive; it does not measure a vision-language model’s understanding of a camera image.

Task progression

TaskWhat it exercises
Pick and placeGrasping, transport and release
Color-conditioned pickObject selection before manipulation
StackingPlacement and sequential precision
SortingRepeated object-to-zone assignment
Language instructionsComposition of task conditions
Move toReaching as a simpler control baseline

Implemented and experimental components

  • Environment: a Franka-inspired seven-joint arm in MuJoCo with a tabletop, randomized objects and bounded episodes.
  • Tested planning path: MockVLM performs keyword matching; SimGrounder reads privileged simulator state. Passing these tests is not visual grounding accuracy.
  • Motor control: scripted controllers support pipeline tests; a bounded model-based joint planner is evaluated separately on stationary reaching. SAC training and checkpoint evaluation remain separate from these results.
  • Extension points: a Transformers VLM adapter exists. DINOv2-based visual grounding remains a proposed extension in the reviewed source.

Planner and grounding source · Hierarchical executor · Policy implementations

Stationary reaching and holding · 26 September 2026

A new controller comparison fixes the goal 10 cm above the object’s initial top. Success requires strict 2 cm proximity for one second, a surface-motion bound of at most 5 mm for every object, and no forbidden generated contacts involving moving robot bodies. Execution is checked at 500 Hz. Both controllers start from identical complete simulator states with the same repaired physics and fixed goal.

Controller on 100 fresh paired scenesSuccessesWilson 95% interval
Direct DLS29 / 10021.01–38.54%
Collision-aware joint plan88 / 10080.19–93.00%

Both succeed on 29 scenes, the joint planner alone on 59, and neither on 12. Direct DLS fails on 71 forbidden-contact episodes; the planner fails to find an accepted path on 12. A bounded search failure does not prove that a goal is unreachable. This fixed-goal task differs from the earlier moving-center proximity criterion; its score is not comparable with the earlier 35/100.

The 200 episodes contain 253,150 recorded samples. Independent checking confirms the holding condition, and the full semantic record and provenance replay exactly. The largest observed all-object motion bound remains below 0.118 mm. Privileged state, ideal gravity support, discrete checks and collision-pair exclusions limit the claim; this is not learned, hardware or continuously certified safe manipulation.

Read the contract, controllers and results · Frozen method · Raw evidence and replay record.

Focused GitHub validation passed 293 tests and three development simulation smokes. The source remains on a review branch; these checks do not certify the broader training/rendering suite.

Earlier proximity studies · 26 September 2026

Initial study: the executor could select the requested object but drive toward the first object in the scene. Retaining the named target through control and scoring fixed that defect, but the corrected controller and each of three controls still scored 0/100. Its mean final distance was 33.13 cm against a 3 cm threshold. The 400 original traces and replay record remain available.

Follow-up: development diagnostics identified penetrating arm colliders and gravity-related drift. Two collider corrections preserve the original inertials and kinematics; a collision-free home, aligned actuator targets and ideal robot-only gravity support change the simulated system. The gain-8 policy and strict 3 cm criterion remain unchanged.

System on 100 fresh paired scenesSuccessesWilson 95% interval
Frozen original system0 / 1000–3.70%
Repaired geometry/reset/support bundle35 / 10026.36–44.75%

Thirty-five pairs succeed only after the repair; 65 fail in both systems. The changed home also starts closer on average, so this measures the complete system bundle. The 200 episodes replayed exactly on the recorded runtime; replay adds no independent scenes.

Limits: success is instantaneous proximity to the live target body center, without dwell, grasp or stable placement. The target moves more than 3 cm in 42 repaired episodes, including six of the 35 successes. These outcomes remain in the score. Ideal gravity support is a simulation assumption. The study does not establish safe manipulation, learned-policy performance, VLM understanding or hardware capability.

Read the follow-up and its limits · Frozen method · Raw evidence and replay record.

Focused GitHub validation passed 230 tests and both development simulation smokes. This is a review-branch delivery, not a main-branch merge or a pass of the repository’s full training/rendering suite.

Historical baseline evidence

The archived evaluation dated 22 February 2026 records 100 episodes per task/policy pair, with a 200-step limit. The intervals below are Wilson 95% intervals recalculated from the recorded counts (z = 1.96), not new simulation results.

TaskPolicySuccesses95% intervalMean return
Pick & placeRandom0 / 1000–3.70%−142.37
Pick & placeScripted0 / 1000–3.70%−53.17
Move toRandom0 / 1000–3.70%−715.76
Move toScripted20 / 10013.34–28.88%−260.33
Language taskRandom4 / 1001.57–9.84%−955.20

Language-task success checks final scene predicates, which can already hold at reset; 4/100 does not demonstrate language understanding. The other archived random baselines—color pick, stacking and sorting—each report 0/100 successes. A dash or an absent run is not a zero result. Less-negative shaped return does not imply a multiplicative improvement or successful manipulation.

The original archive does not contain per-episode traces or complete environment/version provenance, and the evaluator did not seed the random action space. These records establish what was reported, not an exactly reproducible run or a learned-policy benchmark.

Current test pipeline

The hierarchical pipeline below remains separate from the stationary study. That controller comparison uses named simulator targets directly; its score does not validate the language planner or the full instruction-to-action chain.

flowchart TB Instruction["Instruction"] --> Mock["MockVLM keyword planner"] Mock --> Parser["Validated task sequence"] Scene["Privileged simulator state"] --> Ground["SimGrounder"] Parser --> Ground Ground --> Exec["Hierarchical executor + scripted policies"] Exec --> Env["MuJoCo tabletop"] Env --> Scene Future["Real VLM, visual grounding and learned control: separate validation"]

Reproduction and next investigation

The stationary-goal protocol and reproduction guide specify the fixed task, two controllers and recorded runtime. The earlier physical-system comparison and four-control study retain their own protocols and results.

Next: investigate the bounded planner’s 12 no-plan cases without relabeling them unreachable or reusing observed scenes as a fresh holdout. Any planner changes need separate development and newly reserved evaluation scenes. Broader repository CI remains a separate requirement before a main-branch release; grasping, placement and language understanding still need their own evidence.

Read the latest control study →

Building a similar system?

Get in touch about your project.