How Well Do Agentic Skills Work in the Wild: Benchmarking LLM Skill Usage in Realistic Settings
- Published
- Source
- arXiv
- Paper number
- 141
- Field
- Agents / Skills
- arXiv ID
- 2604.04323
Key points
- Existing LLM agent skill benchmarks often rely on idealized conditions, where preselected, hand-crafted skills are provided directly to the agent.
- In real-world scenarios, agents face challenges such as independently selecting relevant skills from many options, retrieving skills from large and uncurated repositories, and adapting general skills to specific queries.
- The true utility and limits of agent skills under practical, non-ideal conditions have remained largely unquantified.
- The authors built a large dataset of 34,198 real-world skills to simulate realistic retrieval and adaptation conditions for LLM agents.
- They developed a skill retrieval engine and evaluated different retrieval strategies, including keyword, semantic, hybrid, and agentic retrieval, to identify effective skill discovery methods.
- They evaluate LLM agent performance across six progressively more realistic settings on SKILLSBENCH and TERMINAL-BENCH 2.0, introducing the challenges of selection, retrieval, and adaptation.
Paper links
External research summaries. These are not HDATF publications or measured product results.