Safety evaluations for ASL-4 models will require techniques like mechanistic interpretability to ..., Sonic AI
“Safety evaluations for ASL-4 models will require techniques like mechanistic interpretability to look inside the model, as simple interaction-based tests will no longer be reliable.”