OCR-Memory: Optical Context Retrieval for Long-Horizon Agent Memory
- Published
- Source
- arXiv
- Paper number
- 167
- Field
- Agents / Memory / Multimodal
- arXiv ID
- 2604.26622
Key points
- LLMs have finite context windows, which limits how much past interaction history an agent can preserve and use effectively in long-horizon tasks.
- Existing text-based memory compression methods often lose fine-grained structural and procedural details that are critical for complex agent tasks such as debugging.
- Text-only retrieval methods can return fragmented or semantically similar but logically unrelated information, and generative retrieval can hallucinate incorrect evidence.
- The framework visually encodes the agent's entire interaction trajectory into a dense image and then uses an OCR model to compress it into compact visual tokens.
- The Locate-and-Transcribe retrieval paradigm predicts specific segment indices in visually marked memory to fetch original text deterministically and reduce hallucination risk.
- An adaptive multi-resolution strategy dynamically adjusts the visual fidelity of stored memory over time, reducing visual token cost while still allowing high-fidelity details to be recovered on demand through active recall.
Paper links
External research summaries. These are not HDATF publications or measured product results.