turboquant_plus · Issues· 47 open
Open on GitHubLocally synced open issues (discussions stay on GitHub)
- #99
CUDA vec FA kernel: turbo4 V-cache branch consumes 4 of 8 elements per iteration — half of every head's V output is zero on batch-1 decode (patch attached)
Updated Sep 3, 2026 - #97
uv pip install turboquant also required
Updated Jul 26, 2026 - #92
Can it be used together with the MTP draft model?
Updated May 21, 2026 - #27
Upstream: TurboQuant discussion + contribution requirements for llama.cpp
type:portP2Updated May 4, 2026 - #86
[ROCm] Scale operation fails with "invalid device function" during Gemma 4 loading
Updated Apr 23, 2026 - #85
Feature request: prebuilt binary for Windows CPU only
Updated Apr 23, 2026 - #84
Diagnostic: [RTX 3090]
Updated Apr 20, 2026 - #82
Question about QJL abliation study
Updated Apr 16, 2026 - #80
time and performance overhead of quantization and dequantization
Updated Apr 13, 2026 - #79
Reproducibility of QJL resurrection issue
Updated Apr 13, 2026 - #74
it run
Updated Apr 9, 2026 - #72
End-to-end inference test fails with --cache-type turbo4 on CPU
Updated Apr 9, 2026 - #70
Error when building from source
Updated Apr 3, 2026 - #69
polar coordinate transformation operations
Updated Apr 3, 2026 - #68
llama-quantize crashes with ios_base::failbit when using --tensor-type-file or --tensor-type (regression from #20503)
Updated Apr 3, 2026