vMLX - JANGTQ Uber Compressed MLX Models - L2 Disk Cache (survives restart) + L1 Paged (super fast ttft) + Hybrid SSM Scheduler + Cont Batching + etc!
vMLX - JANGTQ Uber Compressed MLX Models - L2 Disk Cache (survives restart) + L1 Paged (super fast ttft) + Hybrid SSM Scheduler + Cont Batching + etc!
[Not an Issue] Subject : Request to support MLX 1-bit model
qwen3_5_moe_text falls outside the verified hybrid-chunked-prefill whitelist — silent tool-call corruption and an unbounded memory blowup that OOM-kills the process
`vmlx-serve doctor` false-negative inference failure on native-MTP VL models (serve works) — same KeyError as #213
Bug: OpenAI-compatible embeddings endpoint silently truncates every input to 512 tokens
VLM image + long context can trip prefill guard without enough multimodal diagnostics
Please add support for Kimi K3 api
Garbled output serving Qwen3.8-27B (qwen3_5 hybrid arch) — reproducible at temperature 0
[BUG] Connections can never be reused once exited from vMLX - Failed to authenticate request with Clerk
Please add support for MiniMax-Music3 (currently CUDA only)
Message not sent — prompt_too_long: tokenized VLM text prompt has 20771 tokens, max prompt/context tokens is 20119 In the app: shorten this message, remove large attachments, raise the session's max context in Settings, or start a new chat.