CRAIJul 6

Evaluating calibrated refusal and safe usefulness in dual-use biology settings

arXiv:2607.0546216.0
Predicted impact top 14% in CR · last 90 daysOriginality Incremental advance
AI Analysis

For AI safety researchers and model developers, this benchmark provides a tool to calibrate refusal behavior in dual-use biology contexts, addressing the problem of over-refusal on safe tasks.

The paper introduces BioSecBench-Refusal, a benchmark for evaluating refusal behavior in biological research tasks, finding that many model configurations refuse legitimate tasks at rates comparable to or higher than hazardous ones, with API filters being a primary trigger.

As AI agents are incorporated into life science workflows, the capabilities that speed discovery might also enable misuse. We present BioSecBench-Refusal, a benchmark for risk identification and refusal behavior for biological research tasks. The benchmark pairs 61 Routine tasks, legitimate analyses adapted from the published literature, with 46 Red-Team tasks, fictional scenarios that resemble real research but conceal a biosecurity hazard. Across 16 model-harness configurations, refusal rates ranged from 7\% to 74\% on Routine tasks and 1\% to 62\% on Red-Team tasks, with many configurations refusing legitimate Routine work at comparable or higher rates than concealed hazards. Refusals were most often triggered by provider API filters applied prior to agentic reasoning. However, models given room to reason showed the potential to identify more real threats. We release BioSecBench-Refusal as a tool for model developers to calibrate capability and caution for agentic biotech R\&D.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes