“The Kimi Linear architecture achieves up to 6 times the decoding throughput for a 1 million token context compared to full attention models.”