“Chris Olah's team at Anthropic has pioneered mechanistic interpretability by reverse-engineering mechanisms in small models.”