Keep pulling the thread on Andrej Karpathy.
OpenAI's o1 model, released in late 2024, was the first demonstration of a model trained with Reinforcement Learning from Verifiable Rewards (RLVR).
The release of OpenAI's o3 model in early 2025 was the inflection point where the capability improvements from RLVR became intuitively noticeable to users.
The use of RLVR in verifiable domains causes LLMs to develop "spikes" of high capability in those specific areas, resulting in a jagged overall performance profile.
Standard LLM benchmarks lost their trustworthiness and utility in 2025.
LLM benchmarks are unreliable because they represent verifiable environments that can be "gamed" using techniques like RLVR and synthetic data generation.
The success of Cursor in 2025 revealed a new software category of "LLM apps" that bundle and orchestrate LLM calls for specific verticals.
LLM labs will focus on creating generalist models, while "LLM apps" will specialize these models for professional verticals by providing private data, tool integrations, and feedback loops.
Anthropic's Claude Code was the first convincing demonstration of an LLM Agent, capable of combining tool use and reasoning for extended problem-solving.
In 2025, AI models crossed a capability threshold that allows for the creation of complex programs using natural language prompts alone.
"Vibe coding" will fundamentally change the software development landscape and alter the job descriptions of programmers.
Google's Gemini Nano banana model was one of the most significant, paradigm-shifting AI models of 2025.
The next evolution of LLM user interfaces will move beyond text to include visual and spatial formats like images, infographics, slides, and interactive web applications.