Keep pulling the thread on Dan Hendrycks.
Open-weight bio-foundation models could enable malicious actors to develop more deadly bioweapons.
Biohazardous knowledge excluded from bio-foundation models during pre-training can be rapidly recovered via fine-tuning.
Dual-use signals reside in the pretrained representations of bio-foundation models and can be elicited via simple linear probing.
Data filtering as a standalone procedure is insufficient to mitigate the dual-use risks of open-weight bio-foundation models.
Current approaches to mitigate risks from bio-foundation models focus on filtering biohazardous data during pre-training.
The BioRiskEval framework was developed to evaluate the robustness of procedures intended to reduce the dual-use capabilities of bio-foundation models.