“The KIMI-Linear architecture is the first of its kind to outperform full attention on short context tasks, long input tasks, and long output tasks.”