In a combined GPU-LPU architecture, compute-constrained parts of LLM processing are run on the GP..., Sonic AI
“In a combined GPU-LPU architecture, compute-constrained parts of LLM processing are run on the GPU, while memory throughput-constrained parts are run on the LPU.”