Watch, Remember, Reason: Human-View Video Understanding with MLLMs
- Published
- Source
- arXiv
- Paper number
- 369
- Field
- Computer Vision
- arXiv ID
- 2606.07433
Key points
- It formalizes video understanding as the functional process of observing, remembering, and reasoning, and proposes an integrated formulation for perceptual representations, memory state, reasoning traces, and final prediction.
- It systematically organizes the viewing domain into four subareas: fine-grained perception, holistic perception, audiovisual perception, and efficient processing.
- In memory, it distinguishes offline memory from streaming memory to analyze the tension between redundancy and evidence sparsity in long videos.
- In reasoning, it integrates and classifies different paradigms such as agentic versus non-agentic and text-only versus video-involved reasoning.
- It covers five application domains, including egocentric, sports, instructional, medical, and narrative, along with major training datasets and evaluation benchmarks.
Paper links
External research summaries. These are not HDATF publications or measured product results.