HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enchancement

Published
Source
arXiv
Paper number
671
Field
Computer Vision
arXiv ID
2607.18217

Key points

  • The problem setting unifies inter-subject cases, where the reference images are of different targets, and intra-subject cases, where they are multiple views of the same target.
  • It points out the limits of two existing approaches: aligning multimodal model features to the text-embedding space weakens the control power of the text encoder, while replacing the text encoder entirely with a multimodal model introduces a large realignment cost.
  • GMG, the first proposal, adds global representations extracted from the multimodal model to the video tokens during query and key computation, which injects knowledge into self-attention.
  • MRE, the second proposal, uses learnable embeddings that distinguish token modality and reference index, which prevents errors such as duplicating the same object or text map as separate entities.
  • In performance, text recognition accuracy improves by 21.8 percent relative to SkyReels-V3, DINOrec reaches 0.696 and is best among the compared methods, and 40 human raters prefer the overall quality 64.7 percent of the time.
  • Ablation studies show that removing GMG drops face similarity from 0.786 to 0.697, and removing MRE drops text accuracy from 0.452 to 0.376.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)