KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards
- Published
- Source
- arXiv
- Paper number
- 1153
- Field
- AI / Evaluation
- arXiv ID
- 2610.02206
Key points
- They built a benchmark of 8,504 query-command pairs covering 1,642 Kali Linux tools across 23 capability dimensions.
- In the unrestricted setting without tool hints, no open-weight model exceeded 42% exact-command accuracy.
- They created runtime-free verifiable rewards that can be used directly for SFT and reinforcement learning training.
- An 8B model trained with these rewards performed on par with a 685B MoE model.
- They proposed a fine-grained evaluation scheme that decomposes command errors into tool selection, optional arguments, ordering, and syntax.
Paper links
External research summaries. These are not HDATF publications or measured product results.