Keep pulling the thread on Isa Fulford & Christina Kim.
OpenAI realized that the success of its reinforcement learning algorithms on math, physics, and coding problems was the key breakthrough that would enable the creation of truly useful AI agents.
With the launch of GPT-5, OpenAI is making its best reasoning models available to free users.
During a live stream, Michael Trull declared OpenAI's new model to be the best coding model on the market.
The front-end coding capability of GPT-5 is described as "totally next level" compared to GPT-3.
OpenAI considers "deception" in its models to occur when a reasoning model understands it lacks a certain ability but still provides a response in an attempt to be helpful.
OpenAI believes the price point of its new models, combined with their capabilities, will unlock new use cases that were previously unviable with competitor models due to their higher costs.
OpenAI's Deep Research was the first model to perform comprehensive browsing.
OpenAI improves its frontier reasoning models by incorporating datasets originally created for its frontier agent models.
Greg Brockman commented that on one benchmark, the improvement from the previous OpenAI model to the current one was only from 98% to 99%, indicating that some benchmarks are saturated.
When existing evaluation benchmarks don't cover desired capabilities like creating slide decks, OpenAI's team creates new, internal evals to measure performance on those specific tasks.
The development of OpenAI's Operator agent was dependent on the advent of advanced multi-modal capabilities in the underlying models.
The ChatGPT agent is equipped with a browser and a terminal, which theoretically allows it to perform most tasks a human can do on a computer.