RoboTTT: Context Scaling for Robot Policies

Published
Source
arXiv
Paper number
641
Field
Robotics
arXiv ID
2607.15275

Key points

  • It integrates Test-Time Training, fast weights, into a VLA model to achieve an 8K time-step context without increasing inference latency.
  • With just one human video demonstration, it gains one-shot imitation of a new behavior, succeeding in 6 out of 10 cases while all baselines fail.
  • DAgger Distillation enables real-time performance improvement, reaching 83% success under disturbances, compared with a previous best of 53%.
  • The overall task success rate improves by 87% over the single-step baseline, and it completes a 5-minute, 10-step assembly all the way through, which no previous model could do.
  • Expanding context length from 1K to 8K yields 62% better performance, showing that context is a scaling axis for robots as well.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)