Keep pulling the thread on Jan Leike.
Failure to accurately estimate an untrusted policy's subversion capabilities means the safety margin gained from a control technique could be minimal.
It seems difficult to successfully use AI control as the primary method for safely deploying models that are not sufficiently aligned.
The AI safety field should focus on solving the alignment problem rather than just trying to control misaligned AI.
It is currently unclear how well AI control techniques will scale to superhuman-level models.
Solving the AI alignment problem is a more holistic and sustainable solution to misalignment risk than implementing control techniques.
A misaligned AI model could deliberately introduce subtle mistakes or bugs into its work, or attempt to self-exfiltrate.
It is difficult to be highly confident in AI control techniques without a better understanding of capability elicitation.
It is currently unclear how to determine if an AI control technique has been sufficiently red-teamed.
A monitoring model trained on human labels might fail to detect misaligned behavior that humans would also miss.
Deploying 'suspicion' monitors as an AI control technique is cheap and has almost no downside.
AI control techniques plausibly add a margin of safety, but it is difficult to quantify how large that margin is.
Developing AI control techniques seems tractable with current technology, but does not add much value at present because there is not yet much misalignment risk.