Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Tasks

Published
Source
arXiv
Paper number
169
Field
Agents / Skills / Self-Improvement
arXiv ID
2604.20987

Key points

  • We build autonomous agents that can continuously improve by discovering, preserving, and reusing structured skills in complex, long-horizon environments.
  • Current LLM agents often rely on curated human datasets or extensive external feedback to acquire skills, which limits scalability and adaptability.
  • Organizing, refining, and reliably retrieving skills across diverse and future tasks remains a challenge for existing agent architectures.
  • The COS-PLAY framework introduces a multi-agent coevolution system with a decision agent (LLM) that generates trajectories and a skill-bank agent that converts them into reusable structured skills.
  • Skills are defined as reusable, temporally extended action abstractions that include protocols such as summaries, preconditions, plans, and success or abort criteria, and they are stored in a dynamically evolving skill bank.
  • The two agents are optimized with Group Relative Policy Optimization (GRPO) and their own LoRA adapters, fostering a closed-loop coevolution in which better skills improve decision making and better rollouts refine skill learning.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)