Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration
- Published
- Source
- arXiv
- Paper number
- 154
- Field
- Agents / Web / Reasoning
- arXiv ID
- 2604.18131
Key points
- Existing self-evolving LLM agents depend heavily on external human supervision, such as predefined tasks, workflows, or validated reward signals.
- Unlike humans, whose exploration is driven by curiosity, agents lack the ability to learn spontaneously and autonomously in new environments without explicit instructions or rewards.
- Current paradigms often shift the engineering burden into complex agent orchestration instead of enabling genuine workflow-free, reward-free self-evolution.
- The authors introduce Native Evolution, a decoupled agent life cycle in which the agent autonomously explores the environment at inference time without tasks or rewards and summarizes the experience into structured World Knowledge, or K.
- They develop an outcome-based reward function used only during training that measures the usefulness of the generated K by performance on potential downstream tasks.
- They use a two-stage training framework that first applies supervised fine-tuning to internalize meta-evolutionary behavior and then applies reinforcement-based rejection sampling to optimize exploration and information-management strategies.
Paper links
External research summaries. These are not HDATF publications or measured product results.