Paul Kassianik, Blaine Nelson, Yaron Singer
Most security-agent leaderboards ask a generous question: given all the compute it wants, can the model find the bug or crack the challenge? Useful, but it misses how real security work runs, where every reasoning step, tool call, and log query costs money and time. This paper re-scores agents on a cost-versus-success basis across offensive challenges (Cybench) and defensive investigations (Splunk BOTS), comparing models at fixed spend rather than at their best-case peak.
The interesting part is that attack and defense scale differently. Offensive capture-the-flag work keeps improving as you add test-time compute, and cheaper open-weight models can get close to the pricey frontier ones. Defensive SOC work doesn't behave that way: success leans more on disciplined tool use and knowing which logs to pull than on raw reasoning budget. Their takeaway is that benchmarks should report economic efficiency, not just wins.
This reflects the abstract's summary of their evaluation, so the paper and their interactive site hold the per-model breakdowns.
Security-agent evaluations commonly measure peak offensive capability under generous inference budgets, emphasizing vulnerability discovery, exploit development, penetration testing, and CTF completion. Such measurements are useful but incomplete: in operational security, every reasoning step, tool call, telemetry query, and enrichment request consumes budget. We evaluate language-model security agents through this cost-success lens on offensive Cybench challenges and…
CRAFT: Clustering Rubrics to Diagnose Weak LLM Capabilities and Generate Targeted Fine-Tuning Data
arXiv (cs.CL) · July 17, 2026Rate-Utility Frontiers for Language Encodings: Comparing Tokens, Bytes, and Pixels Under Controlled Linguistic Content
arXiv (cs.AI) · July 17, 2026Harmonizing AI Safety Thresholds
arXiv (cs.LG) · July 17, 2026The Honest Quorum Problem: Epistemic Byzantine Fault Tolerance for Agentic Infrastructure