Keep pulling the thread on Leopold Aschenbrenner.
Solving AI alignment requires a concerted, large-scale effort with billion-dollar funding, comparable in scale and ambition to Operation Warp Speed or the moon landing.
There are approximately 300 full-time technical AI safety researchers in the world, compared to a plausible estimate of 100,000 or more researchers working on ML/AI capabilities, a ratio of about 300:1.
The scalable alignment team at OpenAI consists of approximately 7 researchers, out of about 400 total employees.
Reinforcement Learning from Human Feedback (RLHF) will predictably fail to scale to superhuman models because it relies on human supervision, which is not viable for systems that humans cannot reliably understand.
The current scalable alignment plan at AGI labs is 'underwhelmingly unambitious' and resembles a strategy of 'improvise as we go along and cross our fingers'.
The core technical challenge for AGI alignment is scalability, as current techniques relying on human supervision will fail when models become superhuman and their outputs are too complex for humans to evaluate.
The core AI alignment challenge is not designing a complex utility function representing human values, but ensuring an AI reliably executes the human operator's intent.
The probability of an existential risk from AI within the next 20 years is estimated to be around 5%.
OpenAI would like to hire more alignment researchers, but there is a scarcity of great researchers focusing on this problem.
DeepMind is estimated to have 10-20 people working on scalable alignment out of over 1,000 employees.
Google Brain has approximately zero researchers working on AI alignment.
Anthropic has the best ratio of alignment researchers to total staff among AGI labs, with roughly 20-30 people on alignment and interpretability out of just over 100 total employees.