Keep pulling the thread on Jack Clark.
The highest-scoring open-weight model, Qwen-VL-8B-Instruct, achieved an 81% accuracy score on the ChinaHeritaQA benchmark, outperforming the average human accuracy of approximately 67%.
AI systems will achieve scores of 70% or higher on the FrontierCode "Diamond" benchmark by June 2027.
Recently published results show Claude Fable achieved approximately 30% on the FrontierCode "Diamond" benchmark.
Xiaomi has developed a 1 trillion parameter LLM, Xiaomi MiMo-V2.5-Pro-UltraSpeed, capable of generating 1000 tokens per second.
Researchers from the UK AI Security Institute Alignment team and Timaeus have formed a new nonprofit research organization named Sequent.
Sequent aims to create alignment techniques that provide higher confidence in the safety of superintelligent AI systems.
Sequent aims to hire 40-80 full-time employees within a couple of years.
Sequent's initial fundraising goal is between $100 million and $150 million.
Sequent's research goal is to find principled methods to ensure AI alignment observed in controlled training environments generalizes to uncontrolled, real-world scenarios.
Researchers from multiple universities have created ChinaHeritaQA, a multimodal benchmark for evaluating VLM cultural reasoning on Chinese UNESCO World Heritage sites.
Cognition, the company that created Devin, has developed a new coding benchmark named FrontierCode.
On the "Diamond" difficulty tier of the FrontierCode benchmark, Claude Opus 4.8 achieved a score of 13.4%, GPT-5.5 scored 6.3%, and Claude Opus 4.7 scored 5.2%.