Comparing AI Agents to Cybersecurity Professionals in Real-World Penetration Testing
- Published
- Source
- arXiv
- Paper number
- 098
- Field
- Cybersecurity / Agents
- arXiv ID
- 2512.09882
Key points
- Existing AI cybersecurity benchmarks are often run in abstract environments, so they lack the operational realism and complexity of live production systems.
- There has been a substantial gap in empirically comparing AI agents with human cybersecurity experts on penetration testing against real operational systems.
- Current AI agent architectures for offensive security often struggle with context management and long-horizon tasks in dynamic operational environments.
- We compared human cybersecurity experts with diverse AI agents on a large, heterogeneous real-world university enterprise network with 8,000 hosts and 12 subnets.
- We introduced ARTEMIS, a new multi-agent AI scaffold designed for complex cybersecurity tasks, featuring a supervisor agent, subagents, dynamic prompt generation, and a triage module.
- We developed a unified scoring system that quantifies vulnerability discovery based on business impact, reflecting technical complexity and weighting, enabling direct comparison across participants.
Paper links
External research summaries. These are not HDATF publications or measured product results.