13.2CYJul 16
BioTIER: A Refusal Benchmark for Targeted Biological Risk MitigationEleanor M. Marshall, Pedro Medeiros, Peter Peneder et al.
As large language models become increasingly capable, concerns about their potential to assist with biological misuse continue to grow. Prioritization of safety differs across the model ecosystem, with some models freely providing high-risk information that could be misused, and others refusing benign scientific content, potentially hindering legitimate research. Both failures stem from a lack of targeted mitigation to distinguish the most dangerous information from broader scientific content. To address this, we introduce BioTIER (Biological Targeted Information for Exclusion and Refusal), a benchmark designed to enable more targeted biological risk mitigation. BioTIER organizes biological content into three risk sets: Catastrophe Avoidance (CA), Biomedical DURC (BD) and Related Biology (RB). These sets represent a spectrum from extremely narrow high-risk topics to a broad range of benign and beneficial biological knowledge. The benchmark consists of 542 expert-curated prompts with rich associated metadata to support differentiated access policies. We release BioTIER to aid in isolating and gating the tiny fraction of information that could engender catastrophic risk from misuse, while ensuring access to the vast wealth of knowledge that is essential for advancing biological science.
2.4AIFeb 26
LLM Novice Uplift on Dual-Use, In Silico Biology TasksChen Bo Calvin Zhang, Christina Q. Knight, Nicholas Kruus et al.
Large language models (LLMs) perform increasingly well on biology benchmarks, but it remains unclear whether they uplift novice users -- i.e., enable humans to perform better than with internet-only resources. This uncertainty is central to understanding both scientific acceleration and dual-use risk. We conducted a multi-model, multi-benchmark human uplift study comparing novices with LLM access versus internet-only access across eight biosecurity-relevant task sets. Participants worked on complex problems with ample time (up to 13 hours for the most involved tasks). We found that LLM access provided substantial uplift: novices with LLMs were 4.16 times more accurate than controls (95% CI [2.63, 6.87]). On four benchmarks with available expert baselines (internet-only), novices with LLMs outperformed experts on three of them. Perhaps surprisingly, standalone LLMs often exceeded LLM-assisted novices, indicating that users were not eliciting the strongest available contributions from the LLMs. Most participants (89.6%) reported little difficulty obtaining dual-use-relevant information despite safeguards. Overall, LLMs substantially uplift novices on biological tasks previously reserved for trained practitioners, underscoring the need for sustained, interactive uplift evaluations alongside traditional benchmarks.