#237·OmniVoice

Reducing Inconsistent Word Pronunciation Across Generations

Author: kawshikbuet17Created Jul 28, 2026Updated Sep 18, 2026
Labelshelp wantedStale

Checks

  • This template is only for usage issues encountered.
  • I have thoroughly reviewed the project documentation but couldn't find information to solve my problem.
  • I have searched for existing issues, including closed ones, and couldn't find a solution.
  • I am using English to submit this issue to facilitate community communication.

Environment Details

I have been using OmniVoice voice cloning as a Bengali TTS system for the past few weeks. After experimenting with inference steps, precision, and other generation parameters, I have been able to get fairly close to the desired output quality.

However, I am still facing one pronunciation consistency issue. Some words are pronounced correctly most of the time, but are occasionally mispronounced when generating the same or similar text.

For example, the Bengali word “বিকাশ” is pronounced approximately as “bee-kash”. The expected final sound is “sh”, as in the English word “ship”. However, in some generations, it is pronounced closer to “bee-kas”, where the final sound becomes “s”, as in “see”. Occasionally, the middle vowel is also shortened, making it sound closer to “bee-kosh.”

Since the model can already generate the correct pronunciation in many cases, it seems to be a consistency issue rather than an unsupported pronunciation.

I have tried increasing the inference steps, modifying the Bengali spelling, using Latin transliterations such as bikash, and experimenting with phoneme-based inputs. These approaches sometimes improve the result, but none of them completely eliminate the occasional error.

Could you please suggest the best approach to minimize this type of pronunciation inconsistency, especially for production use?

Steps to Reproduce

N/A

✔️ Expected Behavior

No response

❌ Actual Behavior

N/A