KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards

Published
Source
arXiv
Paper number
1153
Field
AI / Evaluation
arXiv ID
2610.02206

Key points

  • They built a benchmark of 8,504 query-command pairs covering 1,642 Kali Linux tools across 23 capability dimensions.
  • In the unrestricted setting without tool hints, no open-weight model exceeded 42% exact-command accuracy.
  • They created runtime-free verifiable rewards that can be used directly for SFT and reinforcement learning training.
  • An 8B model trained with these rewards performed on par with a 685B MoE model.
  • They proposed a fine-grained evaluation scheme that decomposes command errors into tool selection, optional arguments, ordering, and syntax.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)