Keep pulling the thread on Kyle Corbitt.
Kyle Corbitt believes that access to compute is the primary constraint preventing Chinese companies from catching up to American frontier AI models.
Kyle Corbitt believes the AI industry is already in a recursive self-improvement loop.
The performance ceiling for models trained with reinforcement learning (RL) is higher than for those trained with supervised fine-tuning (SFT), even when the SFT data is high-quality human data.
Using a "teacher" model as a judge in a reinforcement learning setup is a viable path to training a "student" model that surpasses the teacher's performance.
The threshold for AI recursive self-improvement to accelerate dramatically is relatively low, requiring a model to simply be better than the smartest humans at identifying and solving research bottlenecks.
Companies creating reinforcement learning environments for AI labs are scaling to tens or hundreds of millions of dollars in revenue within months of launching.
Kyle Corbitt predicts that if recursive self-improvement continues, automated labs for physical sciences could become a major economic factor within two to three years.
By using reinforcement learning on smaller, specialized models, customers can often exceed the performance of frontier models for their specific task while also achieving significantly lower per-token costs.
For frontier labs, the immense cost of training runs, potentially hundreds of millions of dollars, makes it impractical to fix reward hacking or misaligned rewards by retraining from scratch.
CoreWeave acquired the reinforcement learning and custom fine-tuning company OpenPipe in 2023.
Reinforcement Learning (RL) fine-tuning is less likely to cause catastrophic forgetting in models compared to Supervised Fine-Tuning (SFT).
Chinese labs are using distillation strategies, including using LLMs as judges in RL post-training, to fast-follow American frontier models.