Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Tasks
- Published
- Source
- arXiv
- Paper number
- 169
- Field
- Agents / Skills / Self-Improvement
- arXiv ID
- 2604.20987
Key points
- We build autonomous agents that can continuously improve by discovering, preserving, and reusing structured skills in complex, long-horizon environments.
- Current LLM agents often rely on curated human datasets or extensive external feedback to acquire skills, which limits scalability and adaptability.
- Organizing, refining, and reliably retrieving skills across diverse and future tasks remains a challenge for existing agent architectures.
- The COS-PLAY framework introduces a multi-agent coevolution system with a decision agent (LLM) that generates trajectories and a skill-bank agent that converts them into reusable structured skills.
- Skills are defined as reusable, temporally extended action abstractions that include protocols such as summaries, preconditions, plans, and success or abort criteria, and they are stored in a dynamically evolving skill bank.
- The two agents are optimized with Group Relative Policy Optimization (GRPO) and their own LoRA adapters, fostering a closed-loop coevolution in which better skills improve decision making and better rollouts refine skill learning.
Paper links
External research summaries. These are not HDATF publications or measured product results.