Keep pulling the thread on Claude Cowork.
In a new report, OpenAI audited the SWE-Bench Pro coding benchmark and found that 30% of its tasks were broken.
OpenAI has formally retracted its support for the SWE-Bench Pro benchmark, concluding it no longer reliably measures frontier coding capability.
OpenAI has published new national security principles stating it will not support the use of its technology for mass domestic surveillance or high-stakes decisions over the use of force without human judgment.
Anthropic has appointed former Federal Reserve Chair Ben Bernanke to the board of its long-term benefit trust.
By next year, Anthropic's long-term benefit trust will have majority control over the company's corporate board, though this can be overruled by a shareholder supermajority vote.
Meta is building a new $10 billion data center in Alberta, Canada, with plans for a total capacity of 1 gigawatt.
Meta's in-house AI chip program is set to enter production in the coming months, with the first chips scheduled for September.
Beginning next year, Meta plans to design a new custom AI chip every 6 months, a cadence significantly faster than the industry standard.
According to an internal memo, Meta plans to deploy 7 gigawatts of compute capacity this year and double that deployment pace in 2027.
OpenAI's new GPT-5.6 model family is split into three classes: a flagship model named Sol, a mid-sized version named Terra, and a small, cost-efficient version named Luna.
On the Coding Agent Index benchmark, GPT-5.6 is the new state-of-the-art model, outperforming Fable-5 by 3 points.
OpenAI has released ChatGPT Work, an agentic harness that extends the functionality of Codex to general knowledge work.