KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux
KaliBench is the first fine-grained benchmark for natural-language-to-CLI translation on Kali Linux, comprising 8,504 query-command pairs across 1,642 tools, 23 capability dimensions, and 5 security phases. Built through a manuscript-grounded pipeline with deterministic canonicalization and alias-aware evaluation, it is verified via a three-stage pipeline combining LLM validation, sandboxed execution, and human review. No open-weight model exceeds 42% exact-command accuracy without tool hints, but supervised fine-tuning and reinforcement learning with KaliBench-derived verifiable rewards lift an 8B model to performance comparable with a 685B MoE model.
Background and Problem Definition
Large language models are increasingly inserted into cybersecurity workflows as the translation layer between an analyst's intent and the actual tool invocation that intent requires. Yet the evaluation landscape for this role has remained stuck at two extremes. On one end sit knowledge-based assessments that merely probe whether a model recognizes a tool or concept — multiple-choice trivia dressed up as security competence. On the other end sit end-to-end agentic evaluations that judge success purely by whether a flag was captured or an objective reached, collapsing every intermediate command into an opaque black box. Neither extreme measures the capability that actually determines whether an LLM is useful in a real security operations center: can it reliably turn a one-line natural-language request into an executable command line.
That capability is harder than it looks, precisely because cybersecurity tooling lives and dies by strict CLI syntax. A reordered argument, a flag bound to the wrong value, or a misplaced subcommand does not merely produce a warning — it can silently execute a different operation entirely, which in a live penetration test or red-team engagement is the difference between a clean finding and a damaging false action. KaliBench targets exactly this neglected middle layer: it does not ask whether a model has memorized a tool's manual, nor whether an entire attack chain eventually succeeds. It isolates and measures natural-language-to-CLI translation itself, at the granularity of individual commands.
Architectural Core and Technical Principles
The dataset comprises 8,504 query-command pairs spanning 1,642 real Kali Linux tools, organized across 23 capability dimensions and the 5 canonical security phases from reconnaissance through post-exploitation. Construction follows a manuscript-grounded pipeline: rather than hand-authoring prompts, the dataset is anchored to each tool's official documentation, then passed through deterministic canonicalization that maps every syntactically valid variant of a command to one normalized form. This is paired with alias-aware evaluation, so a model that uses an equivalent shorthand flag or an alternate but semantically identical argument is scored correctly rather than penalized for surface-level string mismatch.
The more consequential design choice is the three-stage verification pipeline. Stage one applies LLM-based validation to filter for semantic plausibility. Stage two actually executes candidate commands inside a sandboxed terminal, capturing real runtime feedback to confirm the command is not just plausible but genuinely executable. Stage three brings in human-in-the-loop refinement to catch edge cases the automated stages miss. Together, this gives KaliBench a dual guarantee — semantic correctness and practical executability — that single-stage LLM-judged benchmarks typically cannot claim.
Practical Evaluation and Applications
The authors ran 24 model configurations across three evaluation modes, spanning both general-purpose and security-specialized open-weight models. The headline result is sobering: in the unrestricted setting, with no explicit tool hints provided, not a single open-weight model exceeded 42% exact-command accuracy. Even the strongest currently available open models fall well short of the reliability a production security operation would require if the LLM were acting autonomously.
KaliBench's second contribution is turning this diagnostic signal into a training asset. Because the verification pipeline yields deterministic, runtime-free rewards, they can be reused directly for reinforcement learning without re-executing a live sandbox on every training step, which otherwise would make RL prohibitively expensive at scale. The paper demonstrates that supervised fine-tuning followed by reinforcement learning with these KaliBench-derived verifiable rewards lifts an 8B-parameter model to performance comparable with a 685B-parameter Mixture-of-Experts model — a striking efficiency result that suggests targeted, verifiable training signal can substitute for raw parameter count in this narrow but critical skill.
Industry Impact and Outlook
For security teams evaluating whether to let an LLM drive tool invocation directly, the 42% ceiling is a clear warning: autonomous CLI generation for penetration-testing tools is not yet trustworthy enough to run unsupervised, and any production pipeline built today needs a human or deterministic verification layer in the loop. The runtime-free verifiable reward design is itself a reusable pattern beyond this specific benchmark — it shows how to build RL training signal for domains with strict, checkable syntax without paying the cost of live execution on every rollout.
Looking forward, fine-grained, capability-dimension-indexed benchmarks like KaliBench are likely to become the standard for security-agent evaluation, shifting the field's conversation away from vague "did it pop the box" narratives and toward precise diagnosis of exactly which capability dimension and which security phase still falls short — giving the industry an actual measurable progress bar for deploying LLMs into real security operations.
Sources
FAQ
How large is the KaliBench dataset and what does it cover?
KaliBench contains 8,504 query-command pairs across 1,642 Kali Linux tools, organized into 23 capability dimensions and 5 security phases from reconnaissance to post-exploitation.
What is the best exact-command accuracy achieved by open-weight models?
Across 24 model configurations in the unrestricted setting with no tool hints, no open-weight model exceeded 42% exact-command accuracy.
How does KaliBench enable training, not just evaluation?
Its three-stage verification pipeline (LLM validation, sandboxed execution, human review) produces deterministic, runtime-free verifiable rewards usable directly for SFT and RL, lifting an 8B model to performance comparable with a 685B MoE model.