通过带有声调的拉丁语音译或 G2P 来改进旁遮普语的发音
作者: selsie创建于 2026年8月10日更新于 2026年9月10日
标签enhancementstale
- Is this request related to a challenge you're experiencing? Tell us your story. I am testing Fish Speech for Punjabi TTS using a native Punjabi speaker reference recording. During testing, I found a substantial difference in pronunciation depending on how the same Punjabi sentence is represented. I compared: 1. Punjabi in Gurmukhi 2. Plain Latin transliteration 3. Latin transliteration with explicit diacritics and consonant gemination The plain Latin transliteration loses several pronunciation distinctions because spellings such as
a,aa,d,dh, etc. can be ambiguous without additional information. Interestingly, a diacritic-aware Latin representation produces substantially more natural Punjabi pronunciation in my listening tests. For example, I tested the following sentence: text Ajj mausam bahut vadhīā hai, tē tuhāḍē nāl gall karke merā dil khuś hō giā. Compared with: text Ajj mausam bahut waddiya hai, te tuhade naal gal karke mera dil khush ho gya, the diacritic-aware version consistently sounded much closer to natural Punjabi pronunciation. I generated the diacritic-based version approximately ten times, and the pronunciation was consistently noticeably better. This is particularly relevant for Punjabi because Latin transliteration can easily lose information represented in Gurmukhi, including vowel distinctions, retroflex consonants, consonant gemination and nasalization. For example: text ਗੱਲ is represented more explicitly as: text gall rather than: text gal Similarly, a form such as: text tuhāḍē contains substantially more pronunciation information than: text tuhade The current situation makes it difficult to obtain consistently natural Punjabi pronunciation when users cannot or do not want to provide phonologically explicit input. 2. What is your suggested solution? I would suggest adding better Punjabi-specific text normalization / grapheme-to-phoneme support. Ideally, Fish Speech could support Punjabi Gurmukhi directly while preserving the phonological information represented by the script. A second useful option would be documented support for a
内容来源: fishaudio/fish-speech