Keep pulling the thread on Jonathan Ross.
Groq's hardware can run a 180 billion parameter model at approximately 200 tokens per second, which is about 4x faster than the sub-50 tokens per second expected from NVIDIA's next-generation B200 GPU.
Groq's solution provides inference at approximately 1/10th the cost per token compared to a modern GPU.
NVIDIA's B200 GPU cost $10 billion to develop.
Every 100-millisecond reduction in latency leads to 8% more engagement on desktop and 34% more on mobile.
The winner for scaled inference hardware has not yet been determined, and it is unlikely to be NVIDIA.
According to NVIDIA's latest earnings, 40% of its revenue is already from inference.
The AI compute market will eventually consist of 90-95% inference and 5-10% training.
Meta plans to have the equivalent of 650,000 H100 GPUs by the end of the current year.
Groq will deploy 100,000 of its LPUs by the end of the current year and 1.5 million LPUs by the end of next year.
NVIDIA deployed a total of 500 H100s last year.
Groq predicts it will have approximately 50% of the world's inference compute capacity by the end of next year.
Groq has signed a deal with Aramco Digital to deploy a large amount of compute.