DeepCode: Open Agentic Coding
- Published
- Source
- arXiv
- Paper number
- 095
- Field
- Code Generation / Agents
- arXiv ID
- 2512.07921
Key points
- Existing LLM-based coding agents struggle with information overload and context bottlenecks when converting long, multimodal scientific papers into functional code repositories.
- Common failure modes include preserving fragmented specifications, maintaining global consistency across modules, completing underspecified designs, and ensuring end-to-end execution fidelity.
- Existing general-purpose and specialized scientific code agents have had limited success because their reproduction scores on autonomous software engineering tasks fall far short of human experts.
- DeepCode adopts a multi-stage framework. Blueprint Generation reduces information overload by refining raw documents into structured specifications.
- Code Generation uses CodeMem for stateful structural indexing of the evolving codebase to ensure global consistency, and CodeRAG for conditional knowledge injection from high-quality code corpora.
- Automated Verification and Refinement integrates static analysis with sandbox execution feedback to perform closed-loop error correction, ensuring the functional fidelity of the synthesized repository.
Paper links
External research summaries. These are not HDATF publications or measured product results.