In-Context World Modeling for Robotic Control
- Published
- Source
- arXiv
- Paper number
- 513
- Field
- Robotics
- arXiv ID
- 2606.26025
Key points
- We formulate the problem of restoring missing system configuration psi, such as camera viewpoint and morphology, as context in the VLA setting.
- We prepend N task-independent random-search clips, typically 3 to 5, as context to implicitly infer system dynamics without parameter updates.
- Across six OOD viewpoints, the average success rate is 25.0%, a gain of 5.2 percentage points over Multi-View's 19.8%, and wrong context can reduce performance to 18.9%, proving that the model truly depends on context.
- With 20 to 80 mm spacers attached to the robot gripper, a morphology change is preserved with a 60% margin, with ICWM at 14.4% versus 5.6% for Multi-View.
- Reusing context hidden states with KV caching restores inference latency to the 0.112 s baseline.
- This context adaptation ability does not emerge naturally in standard imitation learning and requires an explicit training incentive.
Paper links
External research summaries. These are not HDATF publications or measured product results.