FA
FAR․AI
Summary
FAR․AI covering AI Safety, AI Alignment, Superhuman AI, and Scalable Oversight. Notable guests include Jan Leike (AI Safety Researcher, Anthropic / OpenAI), and Chris Olah (AI Interpretability Researcher). Episodes span from Sep 2023 to Sep 2024.
2episodes
34total claims
12topics covered
2 episodes
Jan Leike – Supervising AI on hard tasks
›Sep 10, 2024 · WITH Jan Leike
The core AI alignment problem is framed as supervising AI on tasks too complex for humans, with the goal of 'fully eliciting the model's capabilities'. Experiments in 'weak-to-strong generalization...
AI SafetyAI AlignmentSuperhuman AIScalable Oversight+11 more
Chris Olah - Looking Inside Neural Networks with Mechanistic Interpretability
›Sep 1, 2023 · WITH Chris Olah
Mechanistic interpretability aims to reverse engineer neural networks into a human-readable, source-code-like format to understand their internal logic and algorithms. Superposition, where a networ...
AI SafetyMechanistic InterpretabilitySuperpositionTransformers+11 more