Keep pulling the thread on Alexandr Wang.
The RMU unlearning method reduces model performance on the WMDP benchmark.
The RMU unlearning method maintains a model's general capabilities in areas such as biology and computer science after its application.
The White House Executive Order on Artificial Intelligence highlights the risks of large language models empowering malicious actors in developing biological, cyber, and chemical weapons.
The Weapons of Mass Destruction Proxy (WMDP) benchmark is a publicly released dataset containing 3,668 multiple-choice questions.
The Weapons of Mass Destruction Proxy (WMDP) benchmark is designed to serve as a proxy measurement of hazardous knowledge in biosecurity, cybersecurity, and chemical security.
The Weapons of Mass Destruction Proxy (WMDP) benchmark was stringently filtered to eliminate sensitive information prior to its public release.
The Weapons of Mass Destruction Proxy (WMDP) benchmark is designed to evaluate unlearning methods intended to remove hazardous knowledge from large language models.
RMU is a state-of-the-art unlearning method based on controlling model representations.
The WMDP benchmark and its associated code are publicly available at https://wmdp.ai.