Baike.dev
All toolsAI codingTrendingOpen sourceNewsSubmit
Log in
Back to tool

vllm · Issues· 8061 open

Open on GitHub

Locally synced open issues (discussions stay on GitHub)

  • #57406

    [Feature]: GLM 5.3 Performance Optimization

    feature requestglmUpdated Sep 17, 2026
  • #57230

    [ROCm][AMD] GLM5.2/5.3 Performance Optimization on gfx950 / MI355X

    feature requestrocmquantizationUpdated Sep 17, 2026
  • #50587

    [Feature]: Kimi K3 Performance Optimization

    feature requestkimik3Updated Sep 17, 2026
  • #52167

    [RFC]: Extended online quantization roadmap

    RFCquantizationUpdated Sep 17, 2026
  • #43501

    [Feature]: Porting `ActivationQuantFusionPass` to manual fusion

    feature requestUpdated Sep 17, 2026
  • #48895

    [Bug]: moe_wna16_marlin_gemm applies wrong per-row topk weights (mul_topk_weights=True) at gpt-oss NVFP4 MoE shapes — corrupt output

    bugUpdated Sep 17, 2026
  • #43224

    [RFC]: Porting compiler fusions to manual fusion

    RFCstaleUpdated Sep 17, 2026
  • #57346

    [Feature]: [CPU][GLM5Next][KDA] Add CPU KDA backend for GLM-5.3-Flash

    feature requestkimiglmUpdated Sep 17, 2026
  • #57383

    [RFC]: Asymmetric P/D Deployment for DeepSeek-V4.1 Flash

    RFCdeepseekDSv4.1Updated Sep 17, 2026
  • #57200

    [RFC] Model console logging configuration as CLI/serve configuration

    Updated Sep 17, 2026
  • #38175

    [RFC]: Support ViT Full CUDA Graph (Tracker)

    help wantedRFCmulti-modalitykimiUpdated Sep 17, 2026
  • #54207

    [Feature]: Add explicit cache IDs for reusable multimodal and inference state

    feature requestmulti-modalityUpdated Sep 17, 2026
  • #55578

    [Feature]: [vllm_gguf_plugin] Add --gguf-dequant-on-load option to unlock cuBLAS tensor cores for prefill-heavy workloads`

    feature requestquantizationUpdated Sep 17, 2026
  • #57324

    [Bug]: Claude Code tool search: first request of every session rejected with 400 (tool_addition content blocks not accepted by /v1/messages)

    tool-callingUpdated Sep 17, 2026
  • #56851

    [RFC]: Request level text and derender output on `/inference/v1/generate`

    RFCUpdated Sep 17, 2026