Research2min read

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents

Security agent evaluations are often incomplete because they only measure peak offensive capability under generous inference budgets. In operational security, every reasoning step, tool call, or telemetry query consumes budget that limits actual success rates.

High relevanceAgent SecurityAI SecurityRed Teaming

Research problem

Operational tasks like SOC investigations do not scale linearly with compute; success depends more heavily on disciplined tool use and selective enrichment than raw reasoning capacity alone. Current benchmarks often ignore economic efficiency.

Methodology

The authors evaluate language model security agents across offensive Cybench challenges and defensive Splunk BOTS v1 investigation problems using a cost-success lens to compare performance at fixed cost levels.

Key findings

  • Offensive CTF performance improves with additional test-time compute, while scaled open-weight models can approach frontier proprietary systems cost-effectively.
  • Defensive SOC investigation does not scale in the same way; success depends more heavily on disciplined tool use and selective enrichment than raw reasoning budget alone.

Practical impact

Security teams can better decide which models are practically useful for their specific SOC tasks through cost-aware evaluation, rather than relying solely on peak performance metrics.

Limitations

  • The study focuses primarily on Cybench and Splunk BOTS; other operational environments may exhibit different scaling patterns.
  • Long-term cost development over months or years was not fully modeled as the evaluation is short-term oriented.

Details

Original title
Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents
Authors
Paul Kassianik, Blaine Nelson, Yaron Singer
Publication platform
arXiv
Methodology type
experimental-evaluation
Paper
Paper

Why it matters

This research is relevant to AI security practitioners because it demonstrates that traditional benchmarks can be misleading. It explains the difference between theoretical offensive capability and practical operational utility for defensive agents.

Sources