RustMizan: A Compilable, Contamination-Aware Benchmarking Framework for Rust Vulnerabilities
For researchers evaluating LLM agents on Rust vulnerability analysis, RustMizan offers a more realistic and contamination-aware benchmark, revealing substantial room for improvement in fine-grained localization.
Existing vulnerability benchmarks rely on non-compilable snippets, binary labels, and risk contamination from public datasets. RustMizan provides compilable Rust code with multi-level annotations and mutation-based contamination testing, finding that frontier LLM agents achieve 56-65% binary classification but only ~20% line-level F1, which drops ~27% under adversarial cues.
LLM agents are increasingly applied to vulnerability analysis, but existing benchmarks have not kept pace. They typically rely on small non-compilable snippets, focus on binary classification (vulnerable or not), and do not account for the risk that publicly-released datasets are part of model training corpora. We introduce RustMizan, a benchmarking framework for Rust vulnerability analysis that addresses these gaps. RustMizan contains compilable code variants at the crate, file, and function levels, with annotations for binary vulnerability detection, CWE classification, and function- and line-level localization. A paired mutation framework produces semantics-preserving code mutants for contamination testing and robustness probing. Across four frontier models in an agentic setup with command-line access, binary classification sits in the 56-65% range, but line localization F1 stays near 20%, and adversarial cues drop line F1 by about 27%.