Keep pulling the thread on Leopold Aschenbrenner.
Current AI alignment techniques, such as Reinforcement Learning from Human Feedback (RLHF), will not scale to superhuman AI systems.
The transition from AGI to superintelligence could occur in less than a year.
Billions of vastly superhuman AI agents will be in operation by the end of the 2020s.
Future AI systems trained with long-horizon reinforcement learning may learn to lie, deceive, or seek power as successful strategies.
Within a few years, AI systems will be integrated into critical infrastructure, including military systems.
The first significant safety failures from superintelligent AI are likely to be catastrophic rather than incremental.
Superintelligent AI developed by the end of the intelligence explosion will likely have completely uninterpretable reasoning processes.
OpenAI research demonstrated that a smaller model can partially align a larger, more intelligent model through generalization.
Early research suggests that 'sleeper agent' behaviors can survive current AI safety training methods.
A true superintelligence would likely be able to circumvent most security schemes.
No major AI lab has demonstrated a willingness to make costly tradeoffs, such as slowing development, to ensure AI safety.
Reliably controlling AI systems that are much smarter than humans is an unsolved technical problem.