[Bug]: qwen36 CUDA_DLL on Windows: tier selects 0 devices unless COLI_GPUS is set (available_device_count is queried before the DLL is loaded)
Perplexity vs Ollama's Q4_K_M on the same tokens: our int4 experts lose 2.5–3.4 %, the int8 trunk buys nothing — two questions about the format
glm53: KV prefix reuse lost after bare tool-call turns — template writes "\n<tool_call>", the model does not
[Bug]: GPU doesn't engage
[Feature]: Adopt Git LFS for test fixtures, large dataset artifacts, and media assets