Keep pulling the thread on Tri Dao.
The research paper Flash Attention, co-authored by Tree Dao, has been a key reason for the significant reduction in model inference costs.
Networking is a primary bottleneck for large-scale AI model training, an area where NVIDIA currently has a significant lead.
Tree Dao estimates that approximately 90% of AI workloads currently run on NVIDIA hardware.
Tree Dao predicts that within the next couple of years, some AI workloads will become multi-silicon, running on chips from various manufacturers instead of just NVIDIA.
Tree Dao believes true hardware portability is a myth, as even successive generations of NVIDIA chips have significant architectural changes requiring complete software rewrites.
The lack of high-quality, modern training data is a major bottleneck for training AI models to automatically write correct and performant GPU kernels.
Tree Dao estimates that AI model inference costs have decreased by approximately 100x in the last couple of years since ChatGPT's debut.
A recent OpenAI model with 120 billion parameters uses 4-bit quantization for most of its layers, allowing it to fit within 60 gigabytes of memory.
Tree Dao predicts that AI inference costs will decrease by another 10x within the next year.
Tree Dao predicts that agentic AI, where models can take actions and gather information independently, will be the next major application paradigm.
Tree Dao believes that real-time video generation will be a major consumer application that could change the consumer landscape as significantly as TikTok did.
Tree Dao predicts that open-source models will close the quality gap with closed-source models within a year, driven by better tooling for reinforcement learning.