Keep pulling the thread on Ben Mann.
Anthropic informed the United States Senate that current AI models show signs of understanding biology well enough to accelerate bioweapon development.
A key capability marker for an AI Safety Level 4 (ASL-4) model is the ability to recursively improve itself without human oversight.
The most advanced AI models today are at AI Safety Level 2 (ASL-2), showing early indications of dangerous capabilities but not elevating bioweapon risk beyond what is possible with a search engine like Google.
AI models at AI Safety Level 3 (ASL-3) may show early signs of being able to autonomously replicate and spread over the internet.
Anthropic has committed not to develop or deploy models that reach a new AI Safety Level (ASL) until the corresponding safety measures are in place.
The AI industry may need to pause all model development if safety progress falls significantly behind capability advancements.
Anthropic's safety evaluation concluded that Claude 3 did not cross the capability threshold for AI Safety Level 3 (ASL-3).
Anthropic determined there is a 30% chance that Claude 3 could have met the criteria for autonomous replication with additional fine-tuning and improved prompt engineering.
Anthropic determined there is a 30% chance that Claude 3 could have met at least one of the criteria related to chemical, biological, radiological, and nuclear (CBRN) risks with additional fine-tuning.
Anthropic's AI Safety Levels (ASLs) are modeled after the United States government's Biosafety Levels (BSL).
Approximately 8% of all Anthropic employees work in security-adjacent areas.
Anthropic is implementing multi-party authorization and time-bound access controls to reduce the risk of model weight exfiltration.