Current interpretability tools like sparse autoencoders are only able to observe a small fraction..., Sonic AI
“Current interpretability tools like sparse autoencoders are only able to observe a small fraction of the features within a neural network, leaving a large amount of "dark matter" unobserved.”