AVA-Encoder: Towards Agent-Native Video Representation Learning
- Published
- Source
- arXiv
- Paper number
- 885
- Field
- Computer Vision
- arXiv ID
- 2608.12313
Key points
- It organizes people, backgrounds, objects, style, camera, and audio states under story, event, and scene hierarchies, and links the related assets.
- It turns reconstruction errors into natural-language correction instructions, then improves both the shared encoding policy and the video-specific knowledge graph in two stages.
- Policy improvement used 6 videos, and evaluation was done on 18 non-overlapping videos, 129 scenes, and 246 key frames.
- The overall reconstruction score is 49.0 percent, which is 20.7 points higher than the strongest external baseline at 28.3 percent.
- When comparing policies only, it reaches 45.8 percent, beating the human-tuned policy at 44.4 percent, and it reduces the system prompt from 31,336 tokens to 8,052 tokens.
Paper links
External research summaries. These are not HDATF publications or measured product results.