“The Kimi Linear architecture reduces KV cache usage by up to 75% compared to full attention models.”