Luke, mentioned 6 times across podcast episodes and expert conversations analyzed by Sonic.
In practice, running self-play with large language models for an extended period results in a performance plateau where the model ceases to improve.
Standard self-play algorithms for LLMs fail because rewarding the task-generating model (conjecturer) for difficulty incentivizes it to create messy, artificially complex problems rather than useful ones.
The amount of compute spent on reinforcement learning post-training for large language models is now approaching or surpassing the amount spent on pre-training.