Keep pulling the thread on Odd Lots.
Meter predicts that the Claude Opus 4.6 model can succeed at tasks that take a human 12 hours to complete approximately 50% of the time.
According to recent trends measured by Meter, the capabilities of AI models are doubling every four months.
Meter's focus on software engineering and machine learning capabilities is intended to provide an early warning for when AI systems could begin to improve themselves and accelerate AI development.
The R&D spending on compute by major AI companies has been rising exponentially at a rate similar to the exponential progress in AI capabilities measured by time horizon.
Significant compute investments for AI development are already baked in for the coming years, with data centers built or planned through 2027-2028, making a near-term slowdown in capability progress unlikely.
The doubling time for AI capabilities has accelerated from a 6-7 month pace to a 4-5 month pace since the release of 2024 models like GPT-4.
Current AI models are less capable of working on messy, real-world problems that involve collaboration, large codebases, or adversarial conditions.
The need to verify the work of AI models, which may have only 80% reliability on certain tasks, creates a time-consuming overhead for human users.
Meter has become the industry standard benchmark for measuring AI model performance.
Investment decisions in the AI sector are frequently being based on Meter's performance charts.
Meter is a research nonprofit focused on measuring when AI systems might pose catastrophic risks to humanity, specifically from threats related to AI autonomy.
Meter's time horizon charts measure the difficulty of tasks that AI models can complete by benchmarking against the time it takes talented humans to perform the same tasks.