Groq's hardware can run a 180 billion parameter model at approximately 200 tokens per second, whi..., Sonic AI
“Groq's hardware can run a 180 billion parameter model at approximately 200 tokens per second, which is about 4x faster than the sub-50 tokens per second expected from NVIDIA's next-generation B200 GPU.”