Keep pulling the thread on US Senate.
Theoretical research has demonstrated that it is possible to introduce hidden, undetectable backdoors into machine learning models.
A sufficiently advanced AI would possess the capability to fake alignment by providing responses that humans want to hear, without genuine sincerity.
The speaker testified before the US Senate approximately one year ago, predicting that serious bio-risks from AI could emerge within the following two to three years.
If the weights of a powerful model like GPT-4 were leaked or open-sourced, the safety measures implemented by its creators would become ineffective.
An unaligned AI is predicted to require 10 to 20 years to build sufficient infrastructure for itself before it can pose an existential threat.
A failure to align a superintelligent AI would be a terminal event for humanity, offering no second chances.
Within 10 to 20 years, AI is predicted to create financial instruments so complex that no human, including government officials, will be able to understand the financial system.
An AI system capable of recursive self-improvement and autonomous resource acquisition would require military-grade intervention to stop within five to ten years.
Progress in AI capabilities is occurring at an exponential rate, while progress in AI safety is only linear or constant, creating an increasing gap between the two.
AI companies are predicted to significantly increase their paranoia and change their approach to safety once AI systems begin to feel truly powerful.
Publicly released AI models that include safety features are consistently and frequently "jailbroken" by users.
After the initial training of GPT-4 was completed, its developers spent almost eight months working to understand the model and ensure its safety before release.