Keep pulling the thread on Andrew Feldman.
When performing inference tasks, GPUs are typically only 5% to 7% utilized, meaning 93% to 95% of their potential is wasted.
The AI industry's dependence on the transformer architecture is predicted to decrease within the next 3 to 5 years.
The fundamental GPU architecture, which relies on off-chip memory, is not well-suited for AI inference workloads.
Cerebras's wafer-scale architecture enables the use of large amounts of fast SRAM, providing both the speed and capacity needed for large AI models.
The GPU architecture that was an advantage for NVIDIA in graphics has become a weakness for the company in modern AI workloads.
Since its inference product launch on August 26th, Cerebras's architecture has been the fastest across a range of models tested by third parties like Artificial Analysis.
Cerebras solved the historical challenge of wafer-scale yield by designing its processor with hundreds of thousands of identical, redundant tiles, a technique adapted from memory manufacturing.
The AI market is predicted to grow by more than 100 times in the next five years.
Within five years, almost all data used for training AI models will be synthetic.
There is no "CUDA lock-in" for AI inference workloads, as it is simple for users to switch between different hardware providers.
NVIDIA's primary competitive moat is its dominant market share, which makes it the default solution, rather than a technical lock-in like CUDA.
NVIDIA's market share in AI hardware is predicted to decrease from its current near-total dominance to between 50% and 60% within five years.