INTACT: Isomorphic Intent-to-Action Learning for Search-Free World Models

Published
Source
arXiv
Paper number
751
Field
Robotics
arXiv ID
2607.26056

Key points

  • It integrates a JEPA prediction architecture with an intent-to-action interface so it can generate actions directly without retrieval.
  • It processes physical next states and target intents with one shared predictor, while designing the endpoint gradient asymmetrically.
  • In direct mode without retrieval, it reaches 85.78% to 100% success on four LeWM tasks, with 2.9 to 5.5 ms inference time.
  • With selective CEM search, it achieves 96.86% macro success while using 23.44x fewer samples than pure CEM.
  • Joint training with a single encoder over four tasks yields 89.39% direct macro success.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)