Agent-G$^2$: Gaussian Guidance for Agentic Reinforcement Learning
- Published
- Source
- arXiv
- Paper number
- 998
- Field
- AI / General
- arXiv ID
- 2608.23318
Key points
- Diagnostic experiments showed that useful hint depths form a Gaussian-shaped 'band' rather than a single point.
- It estimates a Gaussian (center + width) online for each task cluster using only existing rollout statistics, then samples depths without additional rollouts.
- On ALFWorld, it achieved 95.3% (1.5B)/98.4% (7B), outperforming the strongest hint-based baseline by 1.5–2.3 points.
- It achieved a reward of 92.3 on WebShop with less than one-third of the rollout cost of task-specific search, making it possible to add to existing reinforcement-learning pipelines without a separate search process.
- However, the paper notes that each training task requires an expert trajectory, so the method cannot be applied as-is in environments where such demonstrations are difficult or expensive to obtain.
Paper links
External research summaries. These are not HDATF publications or measured product results.