Keep pulling the thread on Junyang Lin.
The next major development in AI models after reasoning-focused models will be "agentic thinking," which involves thinking in order to act and interact with an environment.
A fundamental challenge in merging "thinking" and "instruct" model modes is that their underlying data distributions and behavioral objectives are substantially different.
The AI industry is transitioning from an era focused on training models to an era focused on training agents.
Training agentic models requires a clean decoupling of training and inference systems to maintain high throughput and prevent pipeline collapse.
The most significant challenge in training agentic AI systems that have tool access is preventing reward hacking.
The next major research bottlenecks in AI will be in environment design, evaluator robustness, and anti-cheating protocols.
In the upcoming "agentic era" of AI, competitive advantage will be determined by the quality of training environments, tighter train-serve integration, and superior harness engineering.
OpenAI's o1 model demonstrated that "thinking" can be a first-class capability that is explicitly trained for and exposed to users.
DeepSeek-R1 proved that reasoning-style post-training for language models can be reproduced and scaled by organizations other than the original innovators.
The development of reasoning models was equally dependent on advancements in infrastructure as it was on advancements in modeling.
The development of reasoning models marked a significant transition in AI from focusing on scaling pre-training to scaling post-training.
Qwen3 was a notable public attempt to create a unified model with both "thinking" and "instruct" modes.