Keep pulling the thread on Epoch.
The primary driver of improvement in AI models has been the addition of more and better data, along with scaling the compute used to process that data.
The primary reason open-source models can quickly catch up to frontier models is that data, the main driver of progress, can be easily distilled from public APIs.
Based on the Chinchilla scaling laws, even an infinite increase in model parameters would only reduce the required training data by a factor of 10, which is insufficient to match human sample efficiency.
The strategy of AI labs for automating complex jobs is to first automate AI research, and then use those automated AI researchers to solve the sample efficiency problem.
There has not been significant progress in the training sample efficiency of AI models over the last few years.
Reinforcement Learning (RL) is the main method used to improve AI models by generating synthetic data from large amounts of compute.
The data labeling and Reinforcement Learning (RL) environment industry currently generates billions of dollars in annual revenue.
The data labeling and Reinforcement Learning (RL) environment industry is expected to grow to tens of billions of dollars in annual revenue.
According to a report by Epoch, open-source models lag behind state-of-the-art frontier models by approximately four months.
A human is exposed to approximately 200 million tokens of language by adulthood, whereas frontier AI models are trained on tens to hundreds of trillions of tokens.
A teenager can learn to drive a car with about 20 hours of practice, which is three to four orders of magnitude less data than Waymo and Tesla use to train their self-driving models.
The argument that evolution pre-trained the human brain is flawed because the human genome, at only 3 gigabytes, is not large enough to store the parameters for such a network.