colibri · Issues· 143 open
Open on GitHubLocally synced open issues (discussions stay on GitHub)
- #1601
coli serve: RAM auto-detect returns 0.0 GB and silently falls back to a tiny expert cache (no warning)
Updated Sep 18, 2026 - #1600
coli run: SNAP env var never set for non-glm engines (olmoe confirmed) — always fails with 'started without a model'
Updated Sep 18, 2026 - #1464
[Bug]: RTX 3080 sm_86, WSL2 — "hybrid batched block failed in MoE" on every prompt
Updated Sep 18, 2026 - #1594
[Bug]: Qwen 3.8 Flash Next is too slow
Updated Sep 17, 2026 - #1306
Qwen3.8-Flash-Next: is a GPU backend (CUDA/HIP) planned for the qwen38 engine? (offer: gfx1151 / ROCm 7.2.4 validation + benchmarks)
Updated Sep 17, 2026 - #1593
[Bug]: doctor reports "no tensor fills these core roles: token embedding, output head" for a healthy DeepSeek V4 container
Updated Sep 17, 2026 - #1592
[Feature]: How I can run Maxmini 2.7
Updated Sep 17, 2026 - #1581
[Bug]: Linux, sibling engines: `--auto-tier` silently drops the VRAM tier that `coli plan`/`coli doctor` advertise (qwen36 CUDA build, 11.8 → 21 tok/s with `--gpu auto`)
Updated Sep 17, 2026 - #1576
glm53: KV prefix reuse lost after bare tool-call turns — template writes "\n<tool_call>", the model does not
Updated Sep 17, 2026 - #1570
Vulkan is never torn down: no engine calls coli_vk_shutdown (Metal is, in kimi_k3)
Updated Sep 17, 2026 - #1577
[Bug]: qwen36 CUDA_DLL on Windows: tier selects 0 devices unless COLI_GPUS is set (available_device_count is queried before the DLL is loaded)
Updated Sep 17, 2026 - #1370
Perplexity vs Ollama's Q4_K_M on the same tokens: our int4 experts lose 2.5–3.4 %, the int8 trunk buys nothing — two questions about the format
Updated Sep 16, 2026 - #1498
[Bug]: GPU doesn't engage
Updated Sep 16, 2026 - #1550
[Feature]:Add Dark mode switch button
Updated Sep 16, 2026 - #1538
DeepSeek V4: incoherent GPU output on GB10 (sm_121, unified memory) from the routed-expert VRAM mirror cache
Updated Sep 15, 2026