Baike.dev
All toolsAI codingTrendingOpen sourceNewsSubmit
Log in
Back to tool

turboquant_plus · Issues· 47 open

Open on GitHub

Locally synced open issues (discussions stay on GitHub)

  • #99

    CUDA vec FA kernel: turbo4 V-cache branch consumes 4 of 8 elements per iteration — half of every head's V output is zero on batch-1 decode (patch attached)

    Updated Sep 3, 2026
  • #97

    uv pip install turboquant also required

    Updated Jul 26, 2026
  • #92

    Can it be used together with the MTP draft model?

    Updated May 21, 2026
  • #27

    Upstream: TurboQuant discussion + contribution requirements for llama.cpp

    type:portP2Updated May 4, 2026
  • #86

    [ROCm] Scale operation fails with "invalid device function" during Gemma 4 loading

    Updated Apr 23, 2026
  • #85

    Feature request: prebuilt binary for Windows CPU only

    Updated Apr 23, 2026
  • #84

    Diagnostic: [RTX 3090]

    Updated Apr 20, 2026
  • #82

    Question about QJL abliation study

    Updated Apr 16, 2026
  • #80

    time and performance overhead of quantization and dequantization

    Updated Apr 13, 2026
  • #79

    Reproducibility of QJL resurrection issue

    Updated Apr 13, 2026
  • #74

    it run

    Updated Apr 9, 2026
  • #72

    End-to-end inference test fails with --cache-type turbo4 on CPU

    Updated Apr 9, 2026
  • #70

    Error when building from source

    Updated Apr 3, 2026
  • #69

    polar coordinate transformation operations

    Updated Apr 3, 2026
  • #68

    llama-quantize crashes with ios_base::failbit when using --tensor-type-file or --tensor-type (regression from #20503)

    Updated Apr 3, 2026