Keep pulling the thread on Joel Becker, Chris Painter.
According to Meter's time horizon chart as of February 2024, Claude 3 Opus can complete tasks that take a human 11 hours and 59 minutes with a 50% success rate.
The doubling time for AI capabilities, as measured by Meter's time horizon metric, has recently accelerated to approximately four months.
Chinese AI models are estimated to be 9 to 12 months behind the capabilities of US models.
The exponential growth in R&D spending on compute by AI companies has occurred at essentially the same rate as the exponential progress in their models' time horizon capabilities.
Joel Becker believes it is implausible that AI capability progress will slow down in the next few years because significant compute and R&D investments are already baked in.
The organization Meter has become an industry standard benchmark for AI model performance.
Investment decisions are frequently being based on the performance charts published by Meter.
Meter is a research nonprofit focused on measuring when AI systems might pose catastrophic risks, specifically from threats related to AI autonomy.
Meter's benchmarks focus specifically on tasks typical for an engineer at a frontier AI lab, such as software engineering and fine-tuning AI models.
Joel Becker predicts that in eight months, AI models will achieve an 80% success rate on tasks where they currently have a 50% success rate.
Current AI models are less capable at ideation and self-awareness of their position in a problem compared to their raw software engineering skills.
The leaders of the major AI labs reportedly do not get along, which complicates efforts to coordinate on AI safety.