Keep pulling the thread on Chip Huyen.
Frontier AI labs are increasingly focusing their energy on post-training techniques rather than pre-training, as the industry may be approaching the limits of available text data on the internet.
The market for AI data labeling is lopsided, with a small number of frontier labs as customers for a massive number of data provider startups, giving the labs significant pricing power.
The single biggest performance improvement for Retrieval-Augmented Generation (RAG) systems comes from better data preparation, not from choosing a specific vector database.
In an experiment at one company, the highest-performing senior engineers saw the biggest productivity boost from using the AI coding assistant Cursor.
Some companies are restructuring their engineering organizations by shifting senior engineers into peer review and process design roles, while junior engineers and AI generate the bulk of the code.
Chip Huen predicts that the rate of improvement in base model capabilities will slow down, and we are unlikely to see the same magnitude of step-change improvements as seen between GPT-2, GPT-3, and GPT-4.
Future improvements in AI performance will increasingly come from the post-training and application-building phases, rather than from breakthroughs in pre-trained base models.
When offered a choice between expensive coding agent subscriptions for their team or an additional headcount, most engineering managers choose the extra headcount.
VP-level executives, when offered a choice between an AI assistant subscription or an additional headcount for a team, will typically choose the AI assistant.
Chip Huen believes that to improve AI applications, developers should focus on talking to users, building reliable platforms, preparing better data, optimizing workflows, and writing better prompts.
Chip Huen observes that many developers mistakenly focus on staying current with AI news, adopting new agentic frameworks, choosing vector databases, and fine-tuning models, which are less effective for improving AI applications.
Many open-source models are trained using distillation, where a smaller model is trained to emulate the outputs of a larger, more capable proprietary model.