License clarification: using Fun-CosyVoice3 outputs to train and sell an independent real-time VC mode
Hello CosyVoice Team,
I am 漆原かまたり (Urushibara Kamatari), an independent Windows audio software developer operating as かまたり本舗 (Kamatari Honpo) in Japan.
I am evaluating a Windows real-time voice conversion project called Seiso Voice Neural. It is intended to convert input speech, including male speech, into a female-sounding voice for communication on VRChat, Discord, and similar services. This is a technical feasibility study and may not become a product. If successful, however, it may be sold as paid Windows software with trained VC model weights bundled in the application.
I have confirmed that FunAudioLLM/Fun-CosyVoice3-0.5B-2512 is published under Apache License 2.0, supports Japanese, zero-shot multilingual synthesis, and instruction-based control. I also understand that some examples in the official materials are sourced from the internet. I would not use any official demo audio or internet-sourced example audio as training data.
My primary proposed workflow is:
Fun-CosyVoice3 used locally -> Japanese female speech WAV files generated either without reference audio, if supported, or with reference audio for which I hold the necessary training and commercial rights -> WAV files used as synthetic training data for an independent real-time voice conversion model -> only the trained VC model bundled with a paid Windows application.
The original CosyVoice software, original weights, reference audio, and generated WAV dataset would not be distributed. The resulting VC weights may be encrypted or stored in a regular model format and would run locally on users' PCs.
Could you please clarify the following?
- May audio generated by Fun-CosyVoice3-0.5B-2512 be used as training data for a separate real-time voice conversion model, and may the resulting VC model be sold as part of paid Windows software?
- If only the independently trained VC weights are distributed, are attribution, NOTICE, Apache-2.0 license text, source identification, or other conditions required? Is encrypted distribution of the VC weights permitted?
- May several different instruction or voice conditions be used to generate synthetic datasets for multiple Japanese female VC target voices or presets bundled in the product?
- Do the conditions differ between using generated WAV files only and using internal features, speaker embeddings, acoustic tokens, or output distributions as teacher signals? No direct conversion or transplantation of the original CosyVoice weights is planned.
- For generation without reference audio, if supported by this checkpoint, are there any speaker or training-data-related restrictions beyond Apache-2.0? For generation with reference audio, is holding the necessary rights to that reference audio sufficient from the CosyVoice side?
- Are any separate usage policies, third-party component licenses, commercial agreements, or other permissions required for this specific downstream commercial VC use?
If possible, please indicate whether each use is permitted, conditionally permitted, or not permitted, and provide links to any applicable official terms.
Thank you for your time.
Source: QwenAudio/CosyVoice