Hevolve AI: Self-Evolving Multimodal AI Agents

Turn your domain expertise into AI agents that keep learning. Hevolve AI lets experts build multimodal AI systems by talking to them and correcting them in real time, with no code to write.

Key Features

Quick Links

© 2024 Hevolve AI Pvt Ltd. All rights reserved.

← All research
Artificial Intelligence
arXiv (cs.AI) · July 16, 2026

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents

Paul Kassianik, Blaine Nelson, Yaron Singer

Most security-agent leaderboards ask a generous question: given all the compute it wants, can the model find the bug or crack the challenge? Useful, but it misses how real security work runs, where every reasoning step, tool call, and log query costs money and time. This paper re-scores agents on a cost-versus-success basis across offensive challenges (Cybench) and defensive investigations (Splunk BOTS), comparing models at fixed spend rather than at their best-case peak.

The interesting part is that attack and defense scale differently. Offensive capture-the-flag work keeps improving as you add test-time compute, and cheaper open-weight models can get close to the pricey frontier ones. Defensive SOC work doesn't behave that way: success leans more on disciplined tool use and knowing which logs to pull than on raw reasoning budget. Their takeaway is that benchmarks should report economic efficiency, not just wins.

This reflects the abstract's summary of their evaluation, so the paper and their interactive site hold the per-model breakdowns.

From the arXiv (cs.AI) abstract

Security-agent evaluations commonly measure peak offensive capability under generous inference budgets, emphasizing vulnerability discovery, exploit development, penetration testing, and CTF completion. Such measurements are useful but incomplete: in operational security, every reasoning step, tool call, telemetry query, and enrichment request consumes budget. We evaluate language-model security agents through this cost-success lens on offensive Cybench challenges and…


More Artificial Intelligence papers