Keep pulling the thread on Yang Zhilin.
The correctness tests for FlashKDA verify an exact match against a PyTorch reference implementation.
The `flash_kda.fwd` kernel API currently requires the key (K) and value (V) dimensions to both be 128.
Moonshot AI has released FlashKDA, a set of high-performance Kimi Delta Attention kernels built on the CUTLASS library.
FlashKDA requires a GPU with SM90 architecture or higher.
Once installed, FlashKDA is automatically dispatched from the `chunk_kda` function within the `flash-linear-attention` library.
Using FlashKDA as a backend requires `flash-linear-attention` version 0.5.0.