Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements
- Published
- Source
- arXiv
- Paper number
- 941
- Field
- Machine Learning
- arXiv ID
- 2608.17310
Key points
- Using evolution strategies that evaluate random parameter perturbations through forward passes, it adjusted all parameters with only inference-level GPU memory.
- It avoided the reinforcement-learning credit-assignment problem, which becomes harder as interactions lengthen, by attributing a long trajectory's reward to a single policy perturbation rather than distributing it across tokens.
- On 15-turn Sudoku, it exceeded a strong GRPO baseline by 12.5%, and on ReAct mathematics and document question answering it improved over GRPO by an average of 8.3%.
- It fully fine-tuned a Qwen3.5-27B web agent on four H100s, and a heuristic search that changed parameters and prompts together surpassed the baseline in 28 of 36 settings.
- It requires additional tuning values such as perturbation scale, and the many forward passes become a disadvantage when environment evaluation is expensive; the issue of drifting in irrelevant directions during continual training also remains unclear.
Paper links
External research summaries. These are not HDATF publications or measured product results.