flash-attention · Issues· 1308 open
Open on GitHubLocally synced open issues (discussions stay on GitHub)
- #2897
[CuTe, SM80] Forward silently ignores `block_sparse_tensors` and returns dense attention
Updated Sep 18, 2026 - #2832
Windows build fails with CUDA 13.4: missing /Zc:preprocessor, /std:c++17 pinned (plus no sm_121 gencode)
Updated Sep 17, 2026 - #2456
[RFC] FA4 — head_dim=256 & head_dim=512
Updated Sep 14, 2026 - #2581
Support head_dim=512 on SM89 (Ada) for Gemma 4 global attention layers
Updated Sep 9, 2026 - #1786
RuntimeError when the number of tokens is too large
Updated Sep 9, 2026 - #2729
[FA4/B200] return_lse=True makes D64 forward about 19% faster
Updated Sep 6, 2026 - #2861
Cannot install on Windows - Aiter path too long
Updated Sep 5, 2026 - #2860
[FA4][SM120] flash_attn_varlen_func / flash_attn_func illegal memory access at >=4 varlen segments under real-workload memory layouts (b26/b29, GeForce Blackwell)
Updated Sep 5, 2026 - #2852
FA4 SM103 backward preprocess predicate shape mismatch for padded head dimensions
Updated Sep 4, 2026 - #2842
splitkv early-exit path ignores unpadded_lse, writing -inf to the wrong LSE slot
Updated Sep 2, 2026 - #2413
FA4 support for RTX 6000 Pro Blackwell
Updated Aug 31, 2026 - #1926
pip install flash-attn hangs on Google Colab A100 with latest environment
Updated Aug 28, 2026 - #2818
Using `torch.library.wrap_triton` introduces measurable CPU overhead in `rotary`
Updated Aug 27, 2026 - #2801
[FA4] max_seqlen_q tensor causes an argument-annotation mismatch in varlen forward
Updated Aug 23, 2026 - #2811
tests/test_flash_attn.py fails collection without einops
Updated Aug 20, 2026