“In an LLM's decoder layer, GPUs are more efficient at the 'attention' portion, while Groq's LPUs are more efficient at applying the model weights.”