“Combining Flash Attention with block-sparse attention enables scaling to the Path-256 task, which has a sequence length of 64,000.”