FlashMLA · Issues· 128 open
Open on GitHubLocally synced open issues (discussions stay on GitHub)
- #220
[Portability][MSVC] GNU always_inline attributes on SM90 device lambdas fail to compile
Updated Sep 6, 2026 - #218
[Portability][MSVC] Replace compound-literal-style dependent-type initialization in dense MLA decode
Updated Aug 30, 2026 - #217
[Portability][MSVC] Replace __int128_t in shared-memory load/store paths 3
Updated Aug 30, 2026 - #212
Sparse MLA decode is not batch-invariant: a row's output depends on its batchmates (1-2 ULP), flipping argmax at near-tie tokens
Updated Aug 21, 2026 - #211
PyBind11 interface: How should function signature and CUDA error handling be handled for registered kernels?
Updated Aug 18, 2026 - #210
10 rounds of hand-written fused attention that never beat a 3-kernel cuBLAS stack — trajectory, where each round died, and a correctness trap
Updated Aug 7, 2026 - #206
DeepSeek-Coder-V2-Lite-Instruct 16B running locally on an RTX 2050 laptop
Updated Aug 3, 2026 - #204
flashmla dense decoding test lack fp8 kv cache case
Updated Aug 3, 2026 - #169
dual gemm
Updated Jul 13, 2026 - #172
Use `pyproject.toml`
Updated Jul 9, 2026 - #192
Sparse MLA decode (V3.2 / FP8 KV): throughput cost of the B200 accuracy fix (5aa668c) scales with topk
Updated Jun 26, 2026 - #190
flash mla kernal是否支持同一batch下不同query的动态token数?
Updated Jun 22, 2026 - #179
Do you open to add SM80 support
Updated Jun 15, 2026 - #66
Ampere architecture FlashMLA bring-up
Updated Jun 1, 2026 - #180
Make FlashMLA a libtorch and cpython stable extension
Updated Apr 27, 2026