Keep pulling the thread on Dan Balsam & Tom McGrath.
Mechanistic interpretability startup Goodfire recently announced a $150 million Series B fundraise at a valuation of $1.25 billion.
Goodfire has announced a new research agenda called 'intentional design' which aims to control what models learn during training by understanding and shaping the loss landscape.
Tom McGrath of Goodfire asserts that a key principle for successful model control is to avoid 'fighting backpropagation' and instead shape the loss landscape so the model naturally learns the desired behavior.
Goodfire demonstrated that removing the identified 'memorization weights' from a model can improve its performance on some reasoning tasks.
To prevent a model from learning to evade a hallucination detection probe during RL training, Goodfire's technique runs the probe on a frozen, separate copy of the model, making it computationally easier for the student model to change its behavior than to evade the fixed detector.
Tom McGrath states that based on the current level of scientific development, intentional design techniques should not be used on frontier model training runs.
Goodfire's analysis of Prima Mente's Alzheimer's model found it was overwhelmingly depending on cell-free DNA fragment length for its predictions, a biomarker not previously emphasized for Alzheimer's in the scientific literature.
Using the insight about DNA fragment length, Goodfire and Prima Mente constructed a simple logistic regression proxy model that recapitulated much of the original model's performance and generalized better than literature baselines to an independent patient cohort.
Granola was the number two company adding the most new customers, according to a recent Ramp monthly report on the fastest-growing software vendors.
Tom McGrath of Goodfire highlights a shift in mechanistic interpretability from using sparse autoencoders to identify distinct concepts to understanding the geometric structures these concepts inhabit within a model's latent space.
Goodfire's first proof of concept for 'intentional design' is a technique that uses a probe trained to detect hallucinations to both steer a model at runtime and provide a reward signal for reinforcement learning.
Goodfire research showed it is possible to distinguish model weights used for memorizing facts from those used for general-purpose reasoning.