Keep pulling the thread on Jan Leike.
OpenAI has developed an automated interpretability technique where GPT-4 is used to write natural language explanations for the behavior of individual neurons in a neural network.
In a randomized controlled trial, human evaluators assisted by critiques from a GPT-3.5 model were able to find 50% more flaws in a summarization task compared to unassisted humans.
OpenAI's alignment research includes training deceptively aligned models to stress-test its evaluation methods.
If AI capabilities advance significantly without corresponding progress in alignment, all major AGI labs will need to collaborate to slow down capabilities research.
The OpenAI nonprofit board has the ultimate authority to decide whether or not to deploy a new model, and can halt a deployment for safety reasons even if there is a strong commercial incentive.
OpenAI's strategy for scaling alignment research relies on using AI to create the equivalent of millions of virtual full-time employees.
OpenAI's Superalignment team has set a goal to solve the core technical challenges of superintelligence alignment within four years.
Superintelligence could lead to the disempowerment of humanity or even human extinction.
OpenAI has created a new team, Superalignment, with the goal of solving the problem of steering and controlling superintelligence within four years.
OpenAI is dedicating 20% of the compute it has secured to date to its Superalignment effort.
Reinforcement Learning from Human Feedback (RLHF) will not scale as an alignment technique because humans will no longer be able to effectively evaluate the outputs of smarter-than-human AI systems.
The goal of OpenAI's Superalignment project is to be able to align an AI system that is roughly as smart as the smartest human alignment researchers.