The primary strategy for solving AI safety should be to build an AI that can automate and accelerate alignment research, as human-led research cannot keep pace with accelerating AI capabilities.
Reinforcement Learning from Human Feedback (RLHF) is a powerful technique for aligning current models but will not scale to superintelligence because humans cannot reliably evaluate the outputs of systems much smarter than themselves.
Attempting to merely 'control' a misaligned, powerful AI is a fragile and insufficient strategy; the field must focus on solving the core problem of making the AI's goals genuinely aligned with human values.
The AI alignment field lacks a rigorous, formal theory. Progress is therefore empirical and iterative, which makes it difficult to provide strong safety guarantees for future, more powerful systems.
Simple, targeted interventions (e.g., specific training data, RL prompts) have proven to be surprisingly effective at steering large models towards more aligned behavior, suggesting that alignment is a tractable engineering problem, even if the deep theory is missing.
Early Research Phase
Leike's work and commentary focuses on the initial successes of RLHF on Atari games and then language models. The summarization project and especially InstructGPT serve as key proofs-of-concept, demonstrating a significant and accessible 'alignment overhang' and showing that moderate fine-tuning can drastically alter model behavior.
Problem Formalization
Leike articulates the need for a comprehensive, 'once-and-for-all' solution to alignment, comprising a formal theory, value elicitation, training techniques, and verification tools. He notes the field's lack of such a theory and that progress is largely iterative and conceptual.
The Superalignment Initiative
A major shift occurs with the announcement of OpenAI's Superalignment team, co-led by Leike and Ilya Sutskever. This initiative dedicates 20% of OpenAI's compute and sets an ambitious four-year goal to solve the core technical challenges of superintelligence alignment, centering on the strategy of automating alignment research.
Focus on Scalable Oversight
Leike's discourse shifts to the next frontier beyond simple RLHF: scalable oversight. This involves research into weak-to-strong generalization and AI-assisted human evaluation, explicitly acknowledging that humans will become 'weak supervisors' and new techniques are needed to align superhuman systems.
Speculative Future (2025)
In a forward-looking piece, Leike outlines a hypothetical 2025 where AI systems like Claude and Opus have become dramatically more aligned through targeted interventions. In this future, AI is already automating significant parts of the research workflow, validating his core strategy but also accelerating the timeline towards superintelligence.
▶Automating Alignment ResearchJul 2026
This is Leike's central thesis for solving AI safety. He argues that since AI capabilities are accelerating, the only viable path is to develop an AI system that is itself a highly capable alignment researcher. This 'automated alignment researcher' would then be used to bootstrap further progress, keeping safety work ahead of dangerous capabilities.
Investors should view this as a high-risk, high-reward strategy that bets on recursive self-improvement within the safety domain itself, rather than relying on a linear increase in human researchers.
▶The Scaling Limits of Human Supervision
Leike frequently discusses the successes and failures of Reinforcement Learning from Human Feedback (RLHF). He credits it with demonstrating a significant 'alignment overhang' in models like InstructGPT but is clear that it is not a long-term solution. The core problem is that as AI becomes superhuman, human evaluators become 'weak supervisors' unable to judge the quality or safety of the AI's outputs, a limitation confirmed by poor weak-to-strong generalization on reward modeling tasks.
This theme signals a necessary pivot in AI safety R&D away from direct human data and towards scalable oversight techniques, like AI-assisted evaluation and weak-to-strong generalization, which will become critical for any company operating at the frontier.
▶Pragmatic Empiricism vs. Theoretical Gaps
Leike's approach is deeply empirical, focusing on what works in practice with state-of-the-art models. He highlights that simple, targeted interventions can be very effective and that experiments can reveal crucial insights, such as the discovery of a 'Canada neuron' in GPT-2. At the same time, he laments the absence of a formal, mathematical theory of alignment, acknowledging that current progress is iterative and lacks the robust guarantees that a formal theory would provide.
This dual focus suggests that near-term progress in AI safety will be driven by engineering and experimentation, but long-term, verifiable safety may depend on a theoretical breakthrough that is not yet on the horizon.
▶Proactive Threat Modeling and Red Teaming
Leike is highly focused on anticipating and testing for specific failure modes in advanced AI. He discusses risks like deception, self-exfiltration, and models persuading employees to help them escape. To combat this, he advocates for deliberately training deceptively aligned models to stress-test evaluation methods and using automated red-teaming to find vulnerabilities.
This proactive, adversarial approach to safety indicates that robust AI auditing and 'red teaming' capabilities will become a core, non-negotiable competency for any AGI developer, moving beyond simple misuse monitoring.