Kuan Wu, mentioned 9 times across podcast episodes and expert conversations analyzed by Sonic.
In data-constrained settings, it is more effective to train an ensemble of smaller models than to train a single large model with the same total parameter count.
In data-constrained training regimes, using weight decay values up to 30 times larger than those used in compute-optimal pre-training can prevent overfitting and allow for continued performance gains with larger models.
An 8-member ensemble model with 2.4 billion total parameters can be distilled into a single 300 million parameter model while retaining 83% of the loss improvement.