Musings on the Alignment Problem
Summary
Musings on the Alignment Problem covering AI Safety, AI Alignment, Large Language Models (LLMs), and Automated Alignment Research. Notable guests include Jan Leike. Episodes span from Sep 2022 to Jan 2026.
Significant progress was made in aligning large language models during 2025, with models like Anthropic's Opus 4.5 and OpenAI's GPT-5.2 showing marked improvements over predecessors. Simple, target...
The author contrasts two AI safety strategies: 'AI control' (managing a potentially misaligned AI's behavior externally) and 'AI alignment' (building an inherently trustworthy AI). While control te...
The current method of AI value alignment, where tech companies make unilateral decisions, is unsustainable and risks being driven by commercial incentives rather than societal well-being. The autho...