The linear representation hypothesis in mechanistic interpretability posits that as a neuron or c..., Sonic AI
“The linear representation hypothesis in mechanistic interpretability posits that as a neuron or combination of neurons fires more, it represents a greater confidence in the presence of a specific feature.”