Baike.dev
All toolsAI codingTrendingOpen sourceNewsSubmit
Log in
< Back to tools
V

Voice-Cloning-App

> 编程语言
Open source

A Python/Pytorch app for easily synthesising human voices

1.4K stars0 likes0 views
WebsiteGitHub

About

A Python/Pytorch app for easily synthesising human voices

Voice Cloning App

A Python/Pytorch app for easily synthesising human voices

Documentation

Discord Server

Video guide

Voice Sharing Hub

FAQ's

System Requirements

  • Windows 10 or Ubuntu 20.04+ operating system
  • 5GB+ Disk space
  • NVIDIA GPU with at least 4GB of memory & driver version 456.38+ (optional)

Key features

  • Automatic dataset generation (with support for subtitles and audiobooks)
  • Additional language support
  • Local & remote training
  • Easy train start/stop
  • Data importing/exporting
  • Multi GPU support

Manual Guides

  • Installation
  • Building the dataset
  • Training
  • Synthesis
  • Making changes

Future Improvements

  • Add support for Talknet
  • Add GTA alignment for Hifi-gan
  • Improved batch size estimation
  • AMD GPU support

Other resources

  • Remote training notebook
  • Try out existing voices at uberduck.ai and Vocodes
  • Youtube data fetching (created by Diskr33t#5880)
  • Synthesize in Colab (created by mega b#6696)
  • Generate youtube transcription (created by mega b#6696)
  • Wit.ai transcription

Acknowledgements

This project uses a reworked version of Tacotron2. All rights for belong to NVIDIA and follow the requirements of their BSD-3 licence.

Additionally, the project uses DSAlign, Silero, DeepSpeech & hifi-gan.

Thank you to Dr. John Bustard at Queen's University Belfast for his support throughout the project.

Supported by uberduck.ai, reach out to them for live model hosting.

Also a big thanks to the members of the VocalSynthesis subreddit for their feedback.

Finally thank you to everyone raising issues and contributing to the project.

GitHub Issues· 46 open

View all on GitHub
  • #188

    languages support

    Updated Dec 1, 2024
  • #183

    i dont know ;)

    Updated Dec 1, 2024
  • #189

    bug: Only iterate over directories

    Updated Dec 1, 2024
  • #185

    How to clone my own voice locally

    Updated Sep 15, 2023
  • #187

    UnpicklingError: invalid load key, '<'.

    Updated Sep 15, 2023
  • #182

    Error in Remote Training

    Updated Sep 15, 2023
  • #186

    Cloned voice has issue

    Updated Aug 27, 2023
  • #125

    Docker images creates import error

    bugUpdated Aug 6, 2023
  • #177

    Error before training

    Updated Aug 2, 2023
  • #169

    Cannot use German for my dataset

    Updated Aug 2, 2023

Highlights

  • •Windows 10 or Ubuntu 20.04+ operating system
  • •5GB+ Disk space
  • •NVIDIA GPU with at least 4GB of memory & driver version 456.38+ (optional)
  • •Automatic dataset generation (with support for subtitles and audiobooks)
  • •Additional language support
  • •Local & remote training
  • •Easy train start/stop
  • •Data importing/exporting
  • •Multi GPU support
  • •Installation

> Tags

Pythondeep-learningpythonpytorchtacotron2

No comments yet. Be the first to share.

> Details

PublishedAug 1, 2026
UpdatedSep 17, 2026
Category编程语言
PricingOpen source

> Related tools

T
TypeScript
JavaScript 的超集,为前端与全栈提供静态类型
P
Python
通用编程语言,广泛用于 Web、数据与 AI
G
Go
Google 推出的简洁高效系统语言