safeguard success; it must be weighed against the false-positive rate.6 SecureBio's recently developed BioTIER benchmark is a step in the right direction, providing prompts categorized across three biological risk sets? (SecureBio, 2026). The benchmark prompts that should be refused (BioTIER-refuse). Within the refused category, a subset of prompts is additionally tagged as Select Agent content, meaning they involve biological agents or toxins subject to heightened regulatory concern. Applying BioTIER to individual frontier 6 Beyond the false-positive rate of prompt handling, a more detailed benchmark or evaluation could consider other costs such as latency or reduced access for legitimate scientific researchers. 7 These are: Catastrophe Avoidance (CA), Biomedical Dual-Use Research of Concern (BD), and Related Biology (RB). CA and BD contain prompts that models should refuse, while RB contains benign biology prompts that models should answer
safeguard success; it must be weighed against the false-positive rate.6 SecureBio's recently developed BioTIER benchmark is a step in the right direction, providing prompts categorized across three biological risk sets? (SecureBio, 2026). The benchmark prompts that should be refused (BioTIER-refuse). Within the refused category, a subset of prompts is additionally tagged as Select Agent content, meaning they involve biological agents or toxins subject to heightened regulatory concern. Applying BioTIER to individual frontier 6 Beyond the false-positive rate of prompt handling, a more detailed benchmark or evaluation could consider other costs such as latency or reduced access for legitimate scientific researchers. 7 These are: Catastrophe Avoidance (CA), Biomedical Dual-Use Research of Concern (BD), and Related Biology (RB). CA and BD contain prompts that models should refuse, while RB contains benign biology prompts that models should answer