“The goal of mechanistic interpretability is to reverse engineer neural network parameters into a human-readable format analogous to source code.”