Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements

Published
Source
arXiv
Paper number
941
Field
Machine Learning
arXiv ID
2608.17310

Key points

  • Using evolution strategies that evaluate random parameter perturbations through forward passes, it adjusted all parameters with only inference-level GPU memory.
  • It avoided the reinforcement-learning credit-assignment problem, which becomes harder as interactions lengthen, by attributing a long trajectory's reward to a single policy perturbation rather than distributing it across tokens.
  • On 15-turn Sudoku, it exceeded a strong GRPO baseline by 12.5%, and on ReAct mathematics and document question answering it improved over GRPO by an average of 8.3%.
  • It fully fine-tuned a Qwen3.5-27B web agent on four H100s, and a heuristic search that changed parameters and prompts together surpassed the baseline in 28 of 36 settings.
  • It requires additional tuning values such as perturbation scale, and the many forward passes become a disadvantage when environment evaluation is expensive; the issue of drifting in irrelevant directions during continual training also remains unclear.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)