DeepSeek-OCR: Contexts Optical Compression
- Published
- Source
- arXiv
- Paper number
- 087
- Field
- Vision / Compression
- arXiv ID
- 2510.18234
Key points
- As input sequence length increases, LLMs face major compute and memory challenges because computation grows quadratically, limiting their ability to handle long text contexts.
- Existing vision encoders in vision-language models (VLMs) suffer from limitations such as complex deployment, excessive image splitting, and high activation memory consumption for high-resolution documents.
- Traditional OCR models did not explicitly address the optimal vision-text compression ratio or the minimum number of vision tokens needed for accurate text decoding when integrated with LLMs for long-context understanding.
- DeepSeek-OCR is an end-to-end vision-language model composed of a new 380M-parameter DeepEncoder and a DeepSeek3B-MoE decoder.
- DeepEncoder combines SAM-base and CLIP-large with a core 16x token compressor, a two-layer convolutional neural network, to efficiently process high-resolution images, reduce activation memory, and produce only a small number of vision tokens.
- The model supports multiple resolution modes, including a dynamic "Gundam" mode that handles ultra-high-resolution inputs through tiling, and is trained on diverse datasets spanning traditional OCR, complex synthetic images such as charts, equations, and diagrams, and general vision data.
Paper links
External research summaries. These are not HDATF publications or measured product results.