Keep pulling the thread on Nathan Lambert.
Jensen Huang predicts that all language models will eventually become reasoning models because the cost of compute will fall enough to make them economically superior.
OpenAI's forthcoming open model is expected to be best-in-class for a specific size category and a subset of tasks.
Meta's recent strategy of offering extremely high compensation is driven by the calculation that the cost of top talent is dramatically lower than the cost of GPUs.
The rate of improvement from simply scaling up AI model size is slowing as the industry shifts focus towards AI agents.
AI2's Tulu project aims to compress complex, industrial-scale post-training recipes into a more tractable format that can be modified by individual researchers.
The post-training suite for AI2's Tulu model family consists of 10 to 15 tasks, whereas OpenAI's post-training process is estimated to involve hundreds of evaluations.
AI2's models, which were based on Llama, matched or exceeded the performance of Meta's Llama 3.1 on core evaluations at the time of their release.
For over a year, the Ultra Feedback dataset has remained the state-of-the-art dataset for open preference tuning in the academic community.
The concept for RLVR (Reinforcement Learning from Verifiable Rewards) was inspired by a statement from John Schulman of OpenAI that major labs perform reinforcement learning on model outputs.
OpenAI's O3 model reportedly uses feedback from Bing searches to determine its subsequent actions during multi-hop tool use.
OpenAI's Deep Research product is likely a fine-tuned version of the O3 model, with progress stemming from RL applied to sub-tasks like information retrieval and search rather than on the final report outcome.
Frontier Labs continue to use human preference data in their model training pipelines.