omlx · Issues· 1372 open
Open on GitHubLocally synced open issues (discussions stay on GitHub)
- #3729
Add mlx-lm architecture support for xing4_0 (XingChen-AGI/Xing4.0-29B-A4B) — oQ fails with "sensitivity measurement produced no scores"
Updated Sep 18, 2026 - #3727
Support Prism ML Hadamard packs (Ternary Bonsai 2 27B, MLX 2-bit): dispatch to the pack's bundled loader + register its MODEL_CONFIG format
Updated Sep 18, 2026 - #3697
Can mlx-lm/mlx-vlm be kept up-to-date?
Updated Sep 18, 2026 - #3725
The model file size displayed on the downloader page is incorrect
Updated Sep 18, 2026 - #3704
Broken Cache Management
Updated Sep 18, 2026 - #3723
Decode throughput decays ~3.2x with server uptime on qwen4_exp (M3 Ultra): 20.5 tok/s after ~10h vs 66.5 tok/s on a fresh process
Updated Sep 18, 2026 - #1970
Concurrent cross-model request handling breaks when memory fits only one model
Updated Sep 17, 2026 - #3716
Bonsai/t5 load patch is installed for every 2-bit checkpoint, and no test covers that path
Updated Sep 17, 2026 - #1835
macOS 27: long-context inference 10x slower than macOS 26.6 (even after HOST_VM_INFO64 fix)
Updated Sep 17, 2026 - #3706
macOS 27: periodic Metal cache-clear drain hits kIOGPUCommandBufferCallbackErrorTimeout, then the SubmissionsIgnored penalty box the engine cannot self-recover from (0.5.7, M3 Ultra)
Updated Sep 17, 2026 - #3709
PDF processing engine not working
Updated Sep 17, 2026 - #3699
SSD prefix cache never stores on qwen4_exp hybrid: available_boundaries=0 while boundary snapshots are enabled
Updated Sep 17, 2026 - #3690
TurboQuant KV cache (4-bit) silently corrupts detail retrieval for requests that hit the prefix cache
Updated Sep 17, 2026 - #3707
[Feature Request] Native macOS menu-bar observability for inference workloads
Updated Sep 16, 2026 - #2029
Server dying with "[METAL] Command buffer execution failed: Caused GPU Timeout Error "
Updated Sep 16, 2026