Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing

Published
Source
arXiv
Paper number
820
Field
Computer Vision
arXiv ID
2608.02711

Key points

  • It unifies 3D understanding, text-to-3D generation, instruction-based 3D editing, and text-based part generation in one architecture.
  • It builds one of the largest multimodal 3D training corpora, with 87 million examples in total, including 25 million for understanding, 50 million for generation, and 12 million for editing.
  • The Nano3D-v2 agent-based pipeline mass-produces geometrically consistent 3D editing data.
  • The combination of Hunyuan3D-VLM for understanding and DiT for generation achieves state of the art on both generation and editing benchmarks.
  • Experiments show cross-task synergy, where better generation and understanding improve editing quality.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)