[Feature Request] Suggestions for Future Updates: Improved Dialect Accuracy, Emotional Expression, and VRAM Utilization
Checks
- This template is only for feature request.
- I have thoroughly reviewed the project documentation but couldn't find any relevant information that meets my needs.
- I have searched for existing issues, including closed ones, and found not discussion yet.
- I am using English to submit this issue to facilitate community communication.
1. Is this request related to a challenge you're experiencing? Tell us your story.
Hi OmniVoice team and @saganaki22 (thanks for the ComfyUI port!),
First of all, the local inference speed and quality for Sichuan dialect and Japanese cloning are truly impressive on my RTX 5060 Ti (16GB).
As you plan for future updates, I would love to see improvements in the following areas:
Better Sichuan dialect Cloning: Enhance overall audio quality, phrasing logic, and natural flow, especially for conversational, everyday speech rather than formal/written text.
Advanced Emotion Control: Move beyond simple punctuation/expression tags to implement more nuanced, natural emotional control (e.g., laughter, joy, sighing, whispering, excitement, and surprise) as heard in real-world dialogue.
Optimized VRAM Strategy: While the current ~2GB model performs remarkably well, consider designing future models to fully leverage 12–16GB VRAM. This extra headroom could unlock the model’s true potential for desktop users while offering a great engineering challenge.
That’s all for now. Thank you so much for your amazing work!
2. What is your suggested solution?
See above. :)
3. Additional context or comments
No response
4. Can you help us with this feature?
- I am interested in contributing to this feature.
Source: k2-fsa/OmniVoice