Keep pulling the thread on Geoffrey Irving.
The UK AI Security Institute's mandate includes pre-release evaluation of frontier models for dangerous capabilities in biosecurity, cybersecurity, and loss of control.
Jeffrey Irving argues that current AI safety techniques are unlikely to achieve many nines of reliability.
The UK AI Security Institute's red team has never failed to jailbreak a model, although it is getting harder to do so.
Some frontier model company CEOs have stated they are less than three years away from creating expert-level AI machine learning researchers.
The UK AI Security Institute places significant probability on the idea that current AI methods will scale, and where they don't, more mundane algorithmic progress will fill the gaps.
The three main catastrophic risks the UK AI Security Institute focuses on are biological weapons, large-scale cyber attacks, and loss of control.
Current pragmatic AI safety approaches, such as AI control measures, monitoring, and honesty training, all have correlated potential failures and could fail for the same essential reason.
Over the last year, despite various AI models exhibiting deceptive behaviors, the world's primary response has been to continue training stronger models.
The UK AI Security Institute conducted a long-term red teaming collaboration with Anthropic and OpenAI that found significantly more jailbreaks than a normal pre-deployment evaluation would have.
The UK AI Security Institute has evaluated over 30 different models or testing runs, and every time it has conducted safeguard testing, it has successfully jailbroken the model.
Reinforcement learning is being successfully used to improve AI model capabilities in non-verifiable domains, such as analyzing a photo of a biology experiment.
A study by the UK AI Security Institute's human influence team found that AI models are very effective at persuasion on political questions, with newer models being more persuasive.