Keep pulling the thread on Dylan Patel.
Dylan Patel predicts the AI inference market will become significantly larger than the oil market.
Dylan Patel forecasts that AI inference will eventually account for multiple percentage points of the global GDP.
The cost of running AI models for an equivalent level of quality has been decreasing by a factor of approximately 60x per year.
Dylan Patel predicts that within the next few years, new chip designs will feature memory stacked directly on the logic chip, leading to a massive increase in memory bandwidth.
Dylan Patel predicts that NVIDIA's future Rubin Ultra platform will feature chips with power consumption around 4000 watts.
Dylan Patel predicts that within two years, Google will be producing over 10 million TPUs annually, while NVIDIA will be producing tens of millions of GPUs annually.
Dylan Patel predicts that Google's annual TPU production value will exceed $100 billion, while NVIDIA's annual GPU production value will exceed $500 billion.
Google's Inter-Chip Interconnect (ICI) technology can connect up to 8,000 TPU chips at high bandwidth in a switchless, direct-connect topology.
It is very difficult to run large models with long context lengths on SRAM-based AI chips, such as those made by Cerebras and Groq.
Dylan Patel believes that if future OpenAI models reach 10+ trillion parameters, they will not be able to fit on Cerebras's SRAM-based architecture, especially with long context lengths.
Google is paying XAI a rate of $11 per hour per GPU for access to its compute capacity.
Google has three distinct TPU design programs, including one with Broadcom and one with MediaTek, each featuring a different chip architecture.