LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks
- Published
- Source
- arXiv
- Paper number
- 668
- Field
- Machine Learning
- arXiv ID
- 2607.18110
Key points
- A single score is an information bottleneck because it cannot distinguish answers that receive the same top score, especially at the upper end.
- The solution is to use the same feedback model as a coach rather than a judge. The coach turns answer evaluation into generalized, reusable advice, or experiential knowledge.
- That advice is given to the teacher model as context and internalized by the policy through on-policy context distillation, which minimizes token-level reverse KL divergence between the teacher and the policy output. At inference time, neither the coach nor the advice context is needed.
- The motivation is illustrated by feedback bandwidth. A 1 to 10 integer score carries 3.3 bits per sample, while a 1024-token piece of advice over a 150,000-word vocabulary theoretically carries 17,600 bits.
- Across two policy families, EL outperformed rubric-based reinforcement learning on held-out and novel benchmarks, whether the feedback model was the policy itself or a commercial API model.
- While RL is prone to reward hacking because it learns to chase reward on the training set, EL is a distribution-matching method that follows the teacher distribution and therefore reduces that failure mode.
Paper links
External research summaries. These are not HDATF publications or measured product results.