Keep pulling the thread on Amin Vahdat.
Google's eighth-generation TPUs, the 8T and 8I, are scheduled to be generally available later in 2026.
The design of Google's TPU 8I was influenced by DeepMind's need for lower latency, leading to a novel network topology called BoardFly that reduces the network diameter.
Citadel uses Google's TPUs to make its trading systems 2x to 4x more efficient and has reduced costs by 30%.
Google's infrastructure can deliver over 97% 'goodput' on a 10,000-chip system, a metric that measures actual forward progress accounting for failures and recovery time.
Google is introducing its eighth generation of TPUs, which for the first time includes two distinct, custom-built chips: the TPU 8T for training and the TPU 8I for inference.
Google's new TPU 8T pod offers nearly 3x the floating-point power compared to the previous generation.
The Google TPU 8T features 2x the scale-out network bandwidth and 4x the scale-up network bandwidth of its predecessor.
The Google TPU 8I pod size has increased by more than 4x to 1,152 chips, designed for coordinating large-scale models and agents.
The Google TPU 8I pod delivers 10x the floating-point exaflops and has 7x larger HBM capacity compared to the previous generation.
Silent data corruption, where a chip intermittently produces incorrect results without failing completely, is a critical reliability challenge in large-scale AI systems.
Amin Vahdat predicts that CPUs will make a comeback due to the general-purpose compute needs of orchestrating AI agents.
Amin Vahdat predicts that the trend of hardware specialization will continue, and the industry may see companies develop more than two specialized chips per year for different workloads.