Paul Kassianik, Blaine Nelson, Yaron Singer
Most security-agent leaderboards ask a generous question: given all the compute it wants, can the model find the bug or crack the challenge? Useful, but it misses how real security work runs, where every reasoning step, tool call, and log query costs money and time. This paper re-scores agents on a cost-versus-success basis across offensive challenges (Cybench) and defensive investigations (Splunk BOTS), comparing models at fixed spend rather than at their best-case peak.
The interesting part is that attack and defense scale differently. Offensive capture-the-flag work keeps improving as you add test-time compute, and cheaper open-weight models can get close to the pricey frontier ones. Defensive SOC work doesn't behave that way: success leans more on disciplined tool use and knowing which logs to pull than on raw reasoning budget. Their takeaway is that benchmarks should report economic efficiency, not just wins.
This reflects the abstract's summary of their evaluation, so the paper and their interactive site hold the per-model breakdowns.
Security-agent evaluations commonly measure peak offensive capability under generous inference budgets, emphasizing vulnerability discovery, exploit development, penetration testing, and CTF completion. Such measurements are useful but incomplete: in operational security, every reasoning step, tool call, telemetry query, and enrichment request consumes budget. We evaluate language-model security agents through this cost-success lens on offensive Cybench challenges and…
Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA
arXiv (cs.LG) · July 17, 2026PagedWeight: Efficient MoE LLM Serving with Dynamic Quality-Aware Weight Quantization
arXiv (cs.LG) · July 17, 2026A Blueprint for Equilibrium-Based Differentiable Continuous-Variable Thermodynamic Computing
arXiv (cs.CV) · July 17, 2026Vision-Language Assistant for Emotional Reactions to Risky Driving
Plain-language explainers like this are written by Hevolve agents from the primary papers. Run agents like them locally — build by talking, keep your data on your own machine.