Keep pulling the thread on Noam Brown.
Existing AI safety frameworks, such as responsible scaling policies and preparedness frameworks, do not adequately account for the variable of test-time compute.
The capability of modern AI models is a direct function of the computational budget applied at test-time.
Modern AI models like "5.5" can continue to improve performance on some benchmarks for weeks of computation before plateauing.
The proper way to evaluate AI models is to either set a fixed budget for the benchmark or to plot performance as a function of test-time compute.
Noam Brown predicts that in 6 to 12 months, a future AI model will be able to create a state-of-the-art poker solver zero-shot, a task equivalent to his PhD thesis.
Nobody knows the ceiling of capabilities for currently released AI models because the model release cycle, at every 2-3 months, is faster than the time required to fully test a model's limits.
An internal OpenAI model was used to disprove the Erdős unit distance conjecture.
The AI model "5.5" was capable of disproving the Erdős unit distance conjecture with sufficient scaffolding, at an estimated cost of $1,000 to $100,000.
The computational cost of solving complex problems like the Erdős unit distance conjecture drops by a factor of 10x to 100x with each new AI model release cycle.
OpenAI's strategic focus is on developing more capable models and releasing them safely, rather than using its internal models to solve all existing open problems in science.
Noam Brown does not believe an "overnight intelligence explosion" is likely because the reliance on large-scale test-time compute means that time itself is a bottleneck to rapid capability gains.
The performance of GPT-3 did not scale significantly with increased test-time compute.