Hy-Embodied-VLM-1.0: Efficient Physical-World Agents

Published
Source
arXiv
Paper number
620
Field
Computer Vision
arXiv ID
2607.12894

Key points

  • It organizes pretraining and post-training data around a three-stage capability taxonomy: understanding action-relevant states, reasoning about action transitions, and sequential and adaptive reasoning.
  • It limits active parameters to 3B using the Hy3-A3B language backbone, the Hy-ViT2 visual encoder, and a mixture-of-experts architecture.
  • It achieved the highest performance among similarly sized models on 19 of 38 benchmarks and improved by an average of 8.4% over the previous Hy-Embodied-0.5.
  • By supporting multi-turn interaction and long-horizon reasoning with few active parameters, it can serve as a foundation model for latency-sensitive physical-world agents.
  • The model with 3B active parameters only approaches the previous generation's model with 32B active parameters rather than surpassing it on every evaluation, and its top performance was limited to half of the 38 benchmarks.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)