Watch, Remember, Reason: Human-View Video Understanding with MLLMs

Published
Source
arXiv
Paper number
369
Field
Computer Vision
arXiv ID
2606.07433

Key points

  • It formalizes video understanding as the functional process of observing, remembering, and reasoning, and proposes an integrated formulation for perceptual representations, memory state, reasoning traces, and final prediction.
  • It systematically organizes the viewing domain into four subareas: fine-grained perception, holistic perception, audiovisual perception, and efficient processing.
  • In memory, it distinguishes offline memory from streaming memory to analyze the tension between redundancy and evidence sparsity in long videos.
  • In reasoning, it integrates and classifies different paradigms such as agentic versus non-agentic and text-only versus video-involved reasoning.
  • It covers five application domains, including egocentric, sports, instructional, medical, and narrative, along with major training datasets and evaluation benchmarks.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)