Keep pulling the thread on Jonathan Ross.
Groq's LPU architecture scales across multiple chips without using external memory by spreading the model across the chips, enabling the use of faster SRAM.
A hybrid system combining GPUs and Groq's LPUs can achieve the optimal cost-per-token and capacity across any desired inference speed.
Groq's LPU architecture can run Mixture of Experts (MoE) models within its static scheduling framework by using scatter and gather operations to fetch different experts as needed for a given query.
Groq's LPU architecture is particularly well-suited for Mixture of Experts (MoE) models because it can be economically efficient with a small batch size, such as 10.
Groq's LPUs can perform speech transcription at a rate hundreds of times faster than real-time.
NVIDIA announced the Vera Rubin supercomputer at GTC in March, a system dedicated to AI inference, with a special focus on agentic workloads.
In a hybrid GPU-LPU system for LLM inference, Groq's recommended architecture runs the projection layers on its LPUs and the attention mechanism on the GPUs.
The explosion in AI usage is driven by agentic workflows where AI models delegate subtasks to other AI models, creating exponential growth in compute demand.
The demand for compute is effectively limitless because it is tied to solving major civilizational problems like curing cancer and aging.
The cost per unit of AI intelligence is decreasing, which, due to Jevons' paradox, will lead to an increase in total spending on AI and compute.
Google's speech recognition team developed a model that achieved super-human transcription performance but could not deploy it widely due to a lack of compute capacity.
The initial deployment of Google's super-human speech transcription model was limited to the Nexus phone user base due to compute constraints.