Keep pulling the thread on Nathan Labenz.
Interpretability science indicates that AIs are developing increasingly sophisticated internal world models.
With the successful scaling of reinforcement learning, AIs are no longer merely imitating humans and will likely not be limited by human knowledge for much longer.
It is plausible that AI could enable the majority of human diseases to be cured within the next decade.
The METER benchmark for long-context AI is becoming saturated, with models performing so well that evaluators are struggling to create tasks that are sufficiently long and difficult.
A defense-in-depth strategy combining techniques like Goodfire's intentional design, Redwood Research's AI control, and formal verification for cybersecurity could be sufficient to manage AI risks.
AI systems are expected to soon become better than the vast majority of humans at nearly all forms of cognitive work.
The current paradigm of scaling reinforcement learning (RL) is likely sufficient to create transformative AI.
DeepSeek's R1 model, as detailed in a January 2025 paper, demonstrated emergent higher-order cognitive behaviors like having an "aha moment" during reasoning, a result of reinforcement learning.
OpenAI's latest models are outperforming the human doctors who helped train them at the specific task of evaluating AI-generated medical outputs.
The training objective for frontier AI models has shifted from simple next-token prediction to reinforcement learning based on achieving correct, verifiable answers.
Interpretability techniques like sparse autoencoders have demonstrated that LLMs possess internal world models by allowing researchers to identify and causally intervene on specific concepts, such as the Golden Gate Bridge in an experiment with a Claude model.
The most plausible bottleneck to achieving transformative AI in the next few years is a major disruption to semiconductor fabrication facilities, such as a Chinese invasion of Taiwan.