Keep pulling the thread on Noam Brown.
Adding a 30-second search at the end of a poker game improved the AI's performance by an amount equivalent to scaling up the model's parameters by 100,000x.
The 2017 poker bot Libratus defeated four top poker professionals by a margin of 15 big blinds per 100 over 120,000 hands.
The raw neural network of AlphaGo and AlphaGo Zero, without search, has an Elo rating of around 3,000, which is below top human performance.
As of 2024, no neural network has achieved superhuman performance in Go without using a search algorithm at inference time.
Pre-training costs for large language models are exceeding $100 million.
By scaling up inference compute, OpenAI's O-1 model improved its accuracy on the AIME math test from 20% to over 80%.
The full O-1 model achieves over 80% accuracy on the AIME math benchmark when using consensus.
On a competition coding benchmark, GPT-4 achieves the 11th percentile while O-1 achieves the 89th percentile among human competitors.
On the GPQA Diamond benchmark, O-1 scores approximately 78%, surpassing the expert human score of 70%.
OpenAI's O-1 model uses large-scale reinforcement learning to generate an optimized chain of thought before answering.
The ability to scale inference compute means the anticipated wall for AI progress due to astronomical pre-training costs does not exist.
The 2015 poker bot Claudico lost to four top poker professionals by a margin of 9.1 big blinds per 100 in an 80,000-hand competition.