The LLM decode stage will likely be run entirely on GPUs for cost-sensitive applications, on a GP..., Sonic AI
“The LLM decode stage will likely be run entirely on GPUs for cost-sensitive applications, on a GPU-LPU combination for professional users, and exclusively on LPUs for extreme performance-critical tasks.”