inference · Issues· 36 open
Open on GitHubLocally synced open issues (discussions stay on GitHub)
- #4633
ENH: SDNQ Diffusers Quantizer
enhancementfeaturestaleUpdated Sep 18, 2026 - #5548
BUG: stop vLLM inference when requests are aborted or disconnected
buggpuUpdated Sep 17, 2026 - #5542
Qwen3 reranker computes full-sequence vocabulary logits, causing multi-GB VRAM spikes and OOMs
gpuUpdated Sep 16, 2026 - #5523
cutery errors
Updated Sep 11, 2026 - #5522
[Docker CD] Failure on v3.4.0
gpuUpdated Sep 11, 2026 - #5494
[Docker CD] Failure on main
gpuUpdated Sep 10, 2026 - #5484
V3.2.0版本,rerank接口调用无响应
gpuUpdated Sep 8, 2026