Agent-G$^2$: Gaussian Guidance for Agentic Reinforcement Learning

Published
Source
arXiv
Paper number
998
Field
AI / General
arXiv ID
2608.23318

Key points

  • Diagnostic experiments showed that useful hint depths form a Gaussian-shaped 'band' rather than a single point.
  • It estimates a Gaussian (center + width) online for each task cluster using only existing rollout statistics, then samples depths without additional rollouts.
  • On ALFWorld, it achieved 95.3% (1.5B)/98.4% (7B), outperforming the strongest hint-based baseline by 1.5–2.3 points.
  • It achieved a reward of 92.3 on WebShop with less than one-third of the rollout cost of task-specific search, making it possible to add to existing reinforcement-learning pipelines without a separate search process.
  • However, the paper notes that each training task requires an expert trajectory, so the method cannot be applied as-is in environments where such demonstrations are difficult or expensive to obtain.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)