Keep pulling the thread on Jonathan Ross.
NVIDIA acts as a monopsony (a single buyer) for High Bandwidth Memory (HBM) and CoWoS interposers, giving it a cornered resource advantage.
Groq's LPU architecture avoids using external memory by distributing and storing all model parameters directly within a large number of interconnected chips.
Groq's LPU architecture is approximately 3 times more energy-efficient per token than GPU-based systems.
Groq completed a deployment in Saudi Arabia, going from contract signing to serving production tokens in 51 days.
Groq's hardware architecture eliminates the need for external network switches by having the chips communicate directly with each other, acting as their own switch.
Groq's LPU-based inference is more than 5 times lower in cost compared to GPU-based solutions.
The cost of the memory component alone in the latest GPUs exceeds the total deployed capital expenditure per Groq chip.
The operational expenditure (OpEx) alone to run a GPU for inference is equivalent to the combined capital and operational expenditure (CapEx + OpEx) of a Groq LPU to produce the same number of tokens.
NVIDIA's profit margin is between 70% and 80%.
Groq's business model involves partners funding the capital expenditure for deployments, with Groq paying them back with a return (IRR) from revenue, after which a revenue-sharing agreement flips in Groq's favor.
Power availability will become a hard bottleneck for AI compute expansion in 3 to 4 years.
The current lead time for ordering data center generators is 90 months.