Keep pulling the thread on Jan Leike.
It is currently unknown if it is possible to keep AI alignment perpetually ahead of AI capabilities or to find a "once-and-for-all" solution to the alignment problem.
A comprehensive, "once-and-for-all" solution to the AI alignment problem would require four components: a formal theory for alignment, a process to elicit human values, techniques for training aligned AI, and formal verification tools.
The field of AI alignment currently lacks a formal theory that is grounded in mathematics and can make precise, automatically checkable statements about a system's alignment.
The AI alignment field does not currently have a precise, agreed-upon definition of what it means for a system to be "fully aligned."
Jan Leike's favored approach to AI alignment research is to build a system that can perform alignment research better than humans.
The current focus of some AI alignment research is to build a better automated alignment researcher, not to solve the entire alignment problem.
Slowing down AI capabilities research enough for AI alignment research to keep pace is likely prohibitively difficult.
Cooperative inverse reinforcement learning is the closest existing work to a formal theory of AI alignment, but it fails to address key difficulties like tasks the principal cannot understand and multi-agent scenarios.
Current AI alignment techniques are developed iteratively and based on conceptual motivations, such as "evaluation is easier than generation," rather than a formal theory.
A formal proof of alignment for a GPT-3-sized model with 175 billion parameters would be at least 175GB in size.
State-of-the-art formal verification methods for neural networks are limited to verifying local adversarial robustness for small image classifiers on datasets like MNIST and CIFAR.
Critics who claim there is no meaningful progress on the AI alignment problem are mostly pointing to a lack of progress on creating a formal theory of alignment.