Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration

Published
Source
arXiv
Paper number
154
Field
Agents / Web / Reasoning
arXiv ID
2604.18131

Key points

  • Existing self-evolving LLM agents depend heavily on external human supervision, such as predefined tasks, workflows, or validated reward signals.
  • Unlike humans, whose exploration is driven by curiosity, agents lack the ability to learn spontaneously and autonomously in new environments without explicit instructions or rewards.
  • Current paradigms often shift the engineering burden into complex agent orchestration instead of enabling genuine workflow-free, reward-free self-evolution.
  • The authors introduce Native Evolution, a decoupled agent life cycle in which the agent autonomously explores the environment at inference time without tasks or rewards and summarizes the experience into structured World Knowledge, or K.
  • They develop an outcome-based reward function used only during training that measures the usefulness of the generated K by performance on potential downstream tasks.
  • They use a two-stage training framework that first applies supervised fine-tuning to internalize meta-evolutionary behavior and then applies reinforcement-based rejection sampling to optimize exploration and information-management strategies.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)