Coding Agents are Effective Long-Context Processors
- Published
- Source
- arXiv
- Paper number
- 132
- Field
- Code Agents / Long Context
- arXiv ID
- 2603.20432
Key points
- The method is to represent long documents and large corpora as navigable files and directories, then let the coding agent choose executable operations such as grep, sed, Python aggregation, and note files.
- The core idea is that long-context reasoning becomes environment interaction, so the model no longer needs to compress everything into latent attention or a single retrieved snippet.
- As a result, it reports an average improvement of 17.3 percent over the previous best published system, achieves SOTA on four of five benchmarks, and shows strong scaling from 188K-token context to a 3-trillion-token Wikipedia corpus.
- The limitation and implication are that this approach depends on the agent's tool capabilities and on data organization, and that naive retriever addition can suppress useful exploration, so future systems need retrieval that complements rather than replaces filesystem search.
Paper links
External research summaries. These are not HDATF publications or measured product results.