mlx-audio · Issues· 106 open
Open on GitHubLocally synced open issues (discussions stay on GitHub)
- #961
streaming input architecture
Updated Sep 16, 2026 - #950
Whisper checkpoints without a HF processor (e.g. mlx-community/whisper-large-v3-turbo-8bit, converted with mlx-audio-plus) fail late with a confusing "Processor not found" error instead of a clear load-time message
Updated Sep 7, 2026 - #944
nemotron_asr: expose live-input sessions through the existing realtime STT contract
Updated Sep 4, 2026 - #938
Feature: Dynamic UI forms from model metadata
Updated Sep 2, 2026 - #928
Qwen3-ASR mishandles invalid, short, and no-speech audio inputs
Updated Aug 29, 2026 - #927
Qwen3-ASR streaming timestamps depend on max_tokens instead of audio time
Updated Aug 29, 2026 - #921
ICL voice cloning randomly degenerates (no EOS, 24s garbage) — root cause: ref_text not covering full ref_audio; confirmed not a port bug via official PyTorch A/B
Updated Aug 28, 2026 - #917
Whisper non_speech_tokens crashes with IndexError: encode(" -")[0] on empty result (0.4.4 STT)
Updated Aug 28, 2026 - #916
Voxtral TTS: decorative single quotes can add hallucinated speech
Updated Aug 27, 2026 - #903
Model request + implementation plan: dots.tts (rednote-hilab) zero-shot voice cloning
Updated Aug 27, 2026 - #902
STT load_model silently returns a randomly initialized model when checkpoint keys don't match (strict=False default)
Updated Aug 25, 2026 - #898
Streaming `/v1/audio/speech` emits one complete container file per chunk, inserting audible silence at every seam (all container formats)
Updated Aug 21, 2026 - #875
POST /v1/audio/transcriptions: ndjson streaming reports inference-time failures as 200 with an empty body
Updated Aug 7, 2026 - #44
how to use sesame labs?
repo/mlx-audioUpdated Jul 23, 2026 - #16
Support Stable Audio Open v1.0
repo/mlx-audioUpdated Jul 23, 2026