vllm · Issues· 8061 open
Open on GitHubLocally synced open issues (discussions stay on GitHub)
- #57406
[Feature]: GLM 5.3 Performance Optimization
feature requestglmUpdated Sep 17, 2026 - #57230
[ROCm][AMD] GLM5.2/5.3 Performance Optimization on gfx950 / MI355X
feature requestrocmquantizationUpdated Sep 17, 2026 - #50587
[Feature]: Kimi K3 Performance Optimization
feature requestkimik3Updated Sep 17, 2026 - #52167
[RFC]: Extended online quantization roadmap
RFCquantizationUpdated Sep 17, 2026 - #43501
[Feature]: Porting `ActivationQuantFusionPass` to manual fusion
feature requestUpdated Sep 17, 2026 - #48895
[Bug]: moe_wna16_marlin_gemm applies wrong per-row topk weights (mul_topk_weights=True) at gpt-oss NVFP4 MoE shapes — corrupt output
bugUpdated Sep 17, 2026 - #43224
[RFC]: Porting compiler fusions to manual fusion
RFCstaleUpdated Sep 17, 2026 - #57346
[Feature]: [CPU][GLM5Next][KDA] Add CPU KDA backend for GLM-5.3-Flash
feature requestkimiglmUpdated Sep 17, 2026 - #57383
[RFC]: Asymmetric P/D Deployment for DeepSeek-V4.1 Flash
RFCdeepseekDSv4.1Updated Sep 17, 2026 - #57200
[RFC] Model console logging configuration as CLI/serve configuration
Updated Sep 17, 2026 - #38175
[RFC]: Support ViT Full CUDA Graph (Tracker)
help wantedRFCmulti-modalitykimiUpdated Sep 17, 2026 - #54207
[Feature]: Add explicit cache IDs for reusable multimodal and inference state
feature requestmulti-modalityUpdated Sep 17, 2026 - #55578
[Feature]: [vllm_gguf_plugin] Add --gguf-dequant-on-load option to unlock cuBLAS tensor cores for prefill-heavy workloads`
feature requestquantizationUpdated Sep 17, 2026 - #57324
[Bug]: Claude Code tool search: first request of every session rejected with 400 (tool_addition content blocks not accepted by /v1/messages)
tool-callingUpdated Sep 17, 2026 - #56851
[RFC]: Request level text and derender output on `/inference/v1/generate`
RFCUpdated Sep 17, 2026