OCR-Memory: Optical Context Retrieval for Long-Horizon Agent Memory

Published
Source
arXiv
Paper number
167
Field
Agents / Memory / Multimodal
arXiv ID
2604.26622

Key points

  • LLMs have finite context windows, which limits how much past interaction history an agent can preserve and use effectively in long-horizon tasks.
  • Existing text-based memory compression methods often lose fine-grained structural and procedural details that are critical for complex agent tasks such as debugging.
  • Text-only retrieval methods can return fragmented or semantically similar but logically unrelated information, and generative retrieval can hallucinate incorrect evidence.
  • The framework visually encodes the agent's entire interaction trajectory into a dense image and then uses an OCR model to compress it into compact visual tokens.
  • The Locate-and-Transcribe retrieval paradigm predicts specific segment indices in visually marked memory to fetch original text deterministically and reduce hallucination risk.
  • An adaptive multi-resolution strategy dynamically adjusts the visual fidelity of stored memory over time, reducing visual token cost while still allowing high-fidelity details to be recovered on demand through active recall.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)