nano-vllm · Issues· 96 open
Open on GitHubLocally synced open issues (discussions stay on GitHub)
- #274
[Bug] scheduler: `assert scheduled_seqs` crashes engine when decode runs out of KV-cache blocks
Updated Sep 14, 2026 - #261
[Bug] nano-vllm cannot start on Windows
Updated Sep 1, 2026 - #187
Potential TP correctness issue: BlockManager capacity may rely on per-rank local KV-cache estimation
Updated Aug 26, 2026 - #246
Tensor parallel shared-memory RPC lacks worker acknowledgement and can corrupt commands
Updated Jun 25, 2026 - #245
Low inference speed in A100 PCIE GPU
Updated Jun 20, 2026 - #244
Lightweight Go load-balancing proxy for nano-vllm
Updated Jun 17, 2026 - #240
[Bug] may_append allocates new block one token too late, causing KV cache write to unallocated block
Updated Jun 9, 2026 - #230
Documentation Enhancement Suggestion
Updated May 12, 2026 - #228
[change] Int8 KV Cache + Async Pipeline + Head-major reordering for 22% Throughput Boost
Updated May 8, 2026 - #225
Introducing Quantized Inference to Nano-vLLM - A Lightweight FP8 Runtime for Qwen Models
Updated May 7, 2026 - #220
[Feature Request] Support weightless RMSNorm (for FlashNorm weight folding trick)
Updated May 6, 2026 - #221
[Discussion] As vLLM-Omni is also complex, I tried building a minimal version (~1k LOC) inspired by nano-vllm
Updated May 1, 2026 - #219
[Discussion] Duplicate blocks for same-step prefix sharing
Updated Apr 28, 2026 - #206
Implement support for Qwen3.5
Updated Apr 20, 2026 - #114
[BUG] Crashes when the prompt length exactly equals kvcache_block_size
Updated Apr 13, 2026