STT can hallucinate confident, mislabeled transcriptions with no confidence signal, leading to nonsensical input reaching the LLM

Author: alexdimmock95Created Sep 15, 2026Updated Sep 16, 2026

Issue

When the STT backend encounters audio it doesn't reliably support (e.g. an unsupported or under-resourced language), it can silently produce a mislabeled language and garbled/hallucinated transcription, with no confidence indicator. This nonsensical text is then forwarded to the LLM as if it were normal, high-confidence input, producing a confused or irrelevant response with no signal to the user about where things actually broke down. This applies regardless of which STT backend is selected, since it's the pipeline layer, not any individual model, that would need to surface such a signal.

Open question: do the STT backends here expose an internal confidence score the pipeline could surface, or would this need to be estimated another way? Raising this as a discussion point rather than assuming a specific fix.

Steps to Reproduce

  1. Run speech-to-speech local with the default STT backend (Parakeet TDT).
  2. Speak or play audio in a language outside that backend's supported coverage (e.g. Basque, which isn't in Parakeet's 25 supported languages).
  3. Observe: the transcription is garbled/hallucinated, tagged with a specific (often incorrect) language code, and passed to the LLM with no indication of low confidence.

Evidence

- Example 1 — Turkish audio, default (auto-detect) language setting, hallucinated as unrelated English fragments:

USER: E say there. Language: en ... USER: How can some of us need that? Language: en ... ASSISTANT: Sometimes we all need support—just a listening ear or a kind word. You're not alone.

- Example 2 — Chinese audio, default language setting, hallucinated as English, LLM responds fluently to nonsense:

USER: Uh So Joe. Language: en ... ASSISTANT: Hi Joe! How can I help you today?

- Example 3 — Basque audio, language explicitly forced to French (--parakeet_tdt_language fr), hallucinated as unrelated English text; note the LLM's confused response here is itself a symptom of the garbled STT input, not a separate LLM issue:

USER: Tato down this been Idaho, uh California, Utah. Language: fr ... ASSISTANT: I'm not sure I understand. Could you clarify or rephrase your request?

Note: when a language is explicitly forced via --parakeet_tdt_language, the logged Language: field appears to simply echo the configured value rather than reflect actual detection confidence; not sure if this is intended.

Proposal

Surface some signal (visually, in logs, or forwarded to the LLM) when STT confidence is low or the detected/configured language doesn't match what was actually produced, rather than presenting all transcriptions identically regardless of certainty.

Source: huggingface/speech-to-speech