Mechanistic interpretability research, led by figures like Chris Olah, is unlikely to sufficientl..., Sonic AI
“Mechanistic interpretability research, led by figures like Chris Olah, is unlikely to sufficiently reverse engineer models like GPT-7 in time to solve alignment if AGI arrives soon.”