Baike.dev
All toolsTrendingOpen sourceNewsSubmit
Log in
< 返回工具列表
C

chatterbox

> 编程语言
开源

SoTA open-source TTS

25.8K stars0 点赞0 次浏览
访问官网GitHub

工具介绍

SoTA open-source TTS

Chatterbox TTS

Made with ♥️ by

Chatterbox is a family of state-of-the-art, open-source text-to-speech models by Resemble AI.

Latest Release: Chatterbox Multilingual V3

Chatterbox Multilingual V3 is the latest general-purpose multilingual TTS model in the Chatterbox family. It keeps the same 0.5B model size while improving speaker similarity, reducing hallucinations, and producing more natural, conversational speech across languages.

V3 is designed for broad language coverage like V2, but with stronger stability and more expressive generation. It is the recommended multilingual model for users who want one voice cloning model that works across many languages.

Alongside V3, we are releasing the Single Language Pack: dedicated finetunes for priority languages where tighter quality control, stronger language-specific behavior, and more specialized speech generation are valuable.

  • Broad Multilingual Coverage: Designed as the main general-purpose multilingual Chatterbox model, supporting wide language coverage similar to V2.
  • Single Language Pack: Dedicated single-language models provide stronger specialization and quality control where language- and regional-dialect-specific performance matters most.
  • More Consistent Speaker Similarity: Improves voice identity and accent preservation across languages, making cross-language voice cloning more stable and reliable.
  • Reduced Hallucination: V3 is optimized to reduce unwanted continuation, repetition, and off-prompt speech, especially in cases where earlier multilingual models were less stable.

For low-latency English voice agents, Chatterbox-Turbo is our most efficient model. Built on a streamlined 350M parameter architecture, Turbo delivers high-quality speech with less compute and VRAM than our previous models. We have also distilled the speech-token-to-mel decoder, previously a bottleneck, reducing generation from 10 steps to just one, while retaining high-fidelity audio output.

Paralinguistic tags are now native to the Turbo model, allowing you to use [cough], [laugh], [chuckle], and more to add distinct realism. While Turbo was built primarily for low-latency voice agents, it excels at narration and creative workflows.

For the most resource-constrained deployments, Chatterbox-Nano shares Turbo's architecture in an even smaller 110M parameter package. It targets on-device and CPU inference — running 3x faster than realtime on 8 CPU cores — while keeping the same single-step decoder and native paralinguistic tag support. Nano is the recommended model when memory and latency budgets are tightest.

If you like the model but need to scale or tune it for higher accuracy, check out our competitively priced TTS service (link). It delivers reliable performance with ultra-low latency of sub 200ms—ideal for production use in agents, applications, or interactive media.

⚡ Model Zoo

Choose the right model for your application.

Model Size Languages Key Features Best For 🤗 Examples Chatterbox-Turbo 350M English Paralinguistic Tags ([laugh]), Lower Compute and VRAM Zero-shot voice agents, Production Demo Listen Chatterbox-Nano 110M English Same architecture as Turbo, Paralinguistic Tags, Runs on CPU (3x realtime on 8 cores) On-device / CPU inference, tight latency & memory budgets Model — Chatterbox-Multilingual V3 (Language list) 500M 23+ Improved speaker similarity, reduced hallucinations, more natural multilingual speech Global applications, localization, cross-language voice cloning Demo Listen Single Language Pack (Models) 500M each 6 dedicated finetunes Language- and region-specific quality control Priority languages and dialect-sensitive applications Models Demos Chatterbox (Tips and Tricks) 500M English CFG & Exaggeration tuning General zero-shot TTS with creative controls Demo Listen

Installation

pip install chatterbox-tts

Alternatively, you can install from source:

# conda create -yn chatterbox python=3.11
# conda activate chatterbox

git clone https://github.com/resemble-ai/chatterbox.git
cd chatterbox
pip install -e .

We developed and tested Chatterbox on Python 3.11 on Debian 11 OS; the versions of the dependencies are pinned in pyproject.toml to ensure consistency. You can modify the code or dependencies in this installation mode.

Usage

Chatterbox-Turbo
import torchaudio as ta
import torch
from chatterbox.tts_turbo import ChatterboxTurboTTS

# Load the Turbo model
model = ChatterboxTurboTTS.from_pretrained(device="cuda")

# Generate with Paralinguistic Tags
text = "Hi there, Sarah here from MochaFone calling you back [chuckle], have you got one minute to chat about the billing issue?"

# Generate audio (requires a reference clip for voice cloning)
wav = model.generate(text, audio_prompt_path="your_10s_ref_clip.wav")

ta.save("test-turbo.wav", wav, model.sr)
Chatterbox-Nano

Nano shares Turbo's architecture and is loaded through the same ChatterboxTurboTTS class by passing nano=True:

import torchaudio as ta
import torch
from chatterbox.tts_turbo import ChatterboxTurboTTS

# Load the Nano model (also runs on CPU: device="cpu")
model = ChatterboxTurboTTS.from_pretrained(device="cuda", nano=True)

# Generate with Paralinguistic Tags
text = "Hi there, Sarah here from MochaFone calling you back [chuckle], have you got one minute to chat about the billing issue?"

# Generate audio (requires a reference clip for voice cloning)
wav = model.generate(text, audio_prompt_path="your_10s_ref_clip.wav")

ta.save("test-nano.wav", wav, model.sr)
Chatterbox and Chatterbox-Multilingual
…

See example_tts.py, example_tts_turbo.py, example_tts_nano.py, and example_vc.py for more examples.

Supported Languages

The general-purpose Chatterbox Multilingual model supports the following languages:

Arabic (ar) • Danish (da) • German (de) • Greek (el) • English (en) • Spanish (es) • Finnish (fi) • French (fr) • Hebrew (he) • Hindi (hi) • Italian (it) • Japanese (ja) • Korean (ko) • Malay (ms) • Dutch (nl) • Norwegian (no) • Polish (pl) • Portuguese (pt) • Russian (ru) • Swedish (sv) • Swahili (sw) • Turkish (tr) • Chinese (zh)

Single Language Pack

The Single Language Pack provides dedicated finetunes for priority languages and regional variants. Use these when you want stronger language-specific behavior, tighter quality control, or dialect-aware generation beyond the general multilingual model.

Language Model Card Demo Space Chinese ResembleAI/Chatterbox-Multilingual-zh-cmn Demo Latam Spanish ResembleAI/Chatterbox-Multilingual-es-mx-latam Demo Brazilian Portuguese ResembleAI/Chatterbox-Multilingual-pt-br Demo Spain Spanish ResembleAI/Chatterbox-Multilingual-es-es Demo Portugal Portuguese ResembleAI/Chatterbox-Multilingual-pt-pt Demo Hindi ResembleAI/Chatterbox-Multilingual-hi Demo

Original Chatterbox Tips

  • General Use (TTS and Voice Agents):

    • Ensure that the reference clip matches the specified language tag. Otherwise, language transfer outputs may inherit the accent of the reference clip’s language. To mitigate this, set cfg_weight to 0.
    • The default settings (exaggeration=0.5, cfg_weight=0.5) work well for most prompts across all languages.
    • If the reference speaker has a fast speaking style, lowering cfg_weight to around 0.3 can improve pacing.
  • Expressive or Dramatic Speech:

    • Try lower cfg_weight values (e.g. ~0.3) and increase exaggeration to around 0.7 or higher.
    • Higher exaggeration tends to speed up speech; reducing cfg_weight helps compensate with slower, more deliberate pacing.

Built-in PerTh Watermarking for Responsible AI

Every audio file generated by Chatterbox includes Resemble AI's Perth (Perceptual Threshold) Watermarker - imperceptible neural watermarks that survive MP3 compression, audio editing, and common manipulations while maintaining nearly 100% detection accuracy.

Watermark extraction

You can look for the watermark using the following script.

import perth
import librosa

AUDIO_PATH = "YOUR_FILE.wav"

# Load the watermarked audio
watermarked_audio, sr = librosa.load(AUDIO_PATH, sr=None)

# Initialize watermarker (same as used for embedding)
watermarker = perth.PerthImplicitWatermarker()

# Extract watermark
watermark = watermarker.get_watermark(watermarked_audio, sample_rate=sr)
print(f"Extracted watermark: {watermark}")
# Output: 0.0 (no watermark) or 1.0 (watermarked)

Official Discord

👋 Join us on Discord and let's build something awesome together!

Evaluation

Chatterbox Turbo was evaluated using Podonos, a platform for reproducible subjective speech evaluation.

We compared Chatterbox Turbo to competitive TTS systems using Podonos' standardized evaluation suite, focusing on overall preference, naturalness, and expressiveness.

Evaluation reports:

  • Chatterbox Turbo vs ElevenLabs Turbo v2.5
  • Chatterbox Turbo vs Cartesia Sonic 3
  • [Chatterbox Turbo vs VibeVoice 7B](https://podonos.com/resembl

核心特点

  • •Broad Multilingual Coverage: Designed as the main general-purpose multilingual Chatterbox model, supporting wide language coverage similar to V2.
  • •Single Language Pack: Dedicated single-language models provide stronger specialization and quality control where language- and regional-dialect-specific performance matters most.
  • •More Consistent Speaker Similarity: Improves voice identity and accent preservation across languages, making cross-language voice cloning more stable and reliable.
  • •Reduced Hallucination: V3 is optimized to reduce unwanted continuation, repetition, and off-prompt speech, especially in cases where earlier multilingual models were less stable.
  • •General Use (TTS and Voice Agents):
  • •The default settings (exaggeration=0.5, cfg_weight=0.5) work well for most prompts across all languages.
  • •If the reference speaker has a fast speaking style, lowering cfg_weight to around 0.3 can improve pacing.
  • •Expressive or Dramatic Speech:
  • •Try lower cfg_weight values (e.g. ~0.3) and increase exaggeration to around 0.7 or higher.
  • •Higher exaggeration tends to speed up speech; reducing cfg_weight helps compensate with slower, more deliberate pacing.

> 标签

Python

暂无评论,来聊聊你的看法吧

> 工具信息

发布日期2026年8月1日
最后更新2026年9月9日
分类编程语言
定价开源

> 相关工具

T
TypeScript
JavaScript 的超集,为前端与全栈提供静态类型
P
Python
通用编程语言,广泛用于 Web、数据与 AI
G
Go
Google 推出的简洁高效系统语言
Baike.dev

baike.dev helps you discover great languages, frameworks, databases, DevOps and cloud-native tools.

Quick links

  • Home
  • All tools
  • Trending
  • Open source

About

  • About us
  • Community
  • News

Contribute

Found a great developer tool? Share it with the community.

Submit a tool
© 2026 baike.dev Developer EncyclopediaUpdated daily · Discover great developer tools