The Mamba architecture achieves efficiency by keeping its large state in fast on-chip GPU memory ..., Sonic AI
“The Mamba architecture achieves efficiency by keeping its large state in fast on-chip GPU memory (SRAM) rather than writing it to the slower High Bandwidth Memory (HBM), thus avoiding I/O bottlenecks.”