Keep pulling the thread on Opus 4.7.
On DataCurve's DeepSWE benchmark, GPT-5.5 achieved a score of 70%, placing it significantly ahead of competitors.
On the DeepSWE benchmark, GPT-5.5 was found to be more efficient than Opus 4.7, using approximately half the tokens, completing tasks in less than half the time, and costing about one-third as much.
According to The Information, Base10 is nearing the close of a billion-dollar fundraising round that would value the startup at $11 billion.
Sources report that Base10's annualized revenue tripled from $200 million to $600 million during the first quarter of the year.
OpenRouter raised a $113 million Series B led by CapitalG, valuing the company at $1.3 billion.
OpenAI is currently at a $30 billion revenue run rate, while Anthropic has surged to a $45 billion run rate.
The White House opposed Anthropic's plan to expand access to its powerful Mythos model, partly because the government wanted priority access to the model's tokens.
Research firm Epoch AI estimates that while global AI inference capacity is tripling annually, global demand for tokens is growing by approximately 10x per year.
DataCurve has released a new coding benchmark called DeepSWE, which features tasks built from scratch to reflect realistic engineering work and avoid issues like memorization.
On DataCurve's DeepSWE benchmark, GPT-5.4 scored 56%, while Opus 4.7 scored 54%.
On the DeepSWE benchmark, Chinese models performed poorly, with KIMI being the highest scoring at 24% and DeepSeek V4 scoring only 8%.
DataCurve is intentionally not uploading the solutions for its DeepSWE benchmark to GitHub to prevent them from being included in future model training data.