Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents
Security agent evaluations are often incomplete because they only measure peak offensive capability under generous inference budgets. In operational security, every reasoning step, tool call, or telemetry query consumes budget that limits actual success rates.
Research problem
Operational tasks like SOC investigations do not scale linearly with compute; success depends more heavily on disciplined tool use and selective enrichment than raw reasoning capacity alone. Current benchmarks often ignore economic efficiency.
Methodology
The authors evaluate language model security agents across offensive Cybench challenges and defensive Splunk BOTS v1 investigation problems using a cost-success lens to compare performance at fixed cost levels.
Key findings
- Offensive CTF performance improves with additional test-time compute, while scaled open-weight models can approach frontier proprietary systems cost-effectively.
- Defensive SOC investigation does not scale in the same way; success depends more heavily on disciplined tool use and selective enrichment than raw reasoning budget alone.
Practical impact
Security teams can better decide which models are practically useful for their specific SOC tasks through cost-aware evaluation, rather than relying solely on peak performance metrics.
Limitations
- The study focuses primarily on Cybench and Splunk BOTS; other operational environments may exhibit different scaling patterns.
- Long-term cost development over months or years was not fully modeled as the evaluation is short-term oriented.
Details
- Original title
- Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents
- Authors
- Paul Kassianik, Blaine Nelson, Yaron Singer
- Publication platform
- arXiv
- Methodology type
- experimental-evaluation
- Paper
- Paper
Why it matters
This research is relevant to AI security practitioners because it demonstrates that traditional benchmarks can be misleading. It explains the difference between theoretical offensive capability and practical operational utility for defensive agents.
Sources
- Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security AgentsBeyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents - external link
arXiv cs.CR
Primary Source