25MB 下最先进的TTS型号
New: Free Kitten TTS API available at https://platform.kittenml.com
Kitten TTS is an open-source, lightweight text-to-speech library built on ONNX. With models ranging from 15M to 80M parameters (25-80 MB on disk), it delivers high-quality voice synthesis on CPU without requiring a GPU.
Status: Developer preview -- APIs may change between releases.
Commercial support is available. For integration assistance, custom voices, or enterprise licensing, contact us.
speed parameter| Model | Parameters | Size | Download |
|---|---|---|---|
| kitten-tts-mini | 80M | 80 MB | KittenML/kitten-tts-mini-0.8 |
| kitten-tts-micro | 40M | 41 MB | KittenML/kitten-tts-micro-0.8 |
| kitten-tts-nano | 15M | 56 MB | KittenML/kitten-tts-nano-0.8 |
| kitten-tts-nano (int8) | 15M | 25 MB | KittenML/kitten-tts-nano-0.8-int8 |
Note: Some users have reported issues with the
kitten-tts-nano-0.8-int8model. If you encounter problems, please open an issue.
https://github.com/user-attachments/assets/d80120f2-c751-407e-a166-068dd1dd9e8d
Try Kitten TTS directly in your browser on Hugging Face Spaces.
pip install https://github.com/KittenML/KittenTTS/releases/download/0.8.1/kittentts-0.8.1-py3-none-any.whlfrom kittentts import KittenTTS
model = KittenTTS("KittenML/kitten-tts-mini-0.8")
audio = model.generate("This high-quality TTS model runs without a GPU.", voice="Jasper")
import soundfile as sf
sf.write("output.wav", audio, 24000)# Adjust speech speed (default: 1.0)
audio = model.generate("Hello, world.", voice="Luna", speed=1.2)
# Save directly to a file
model.generate_to_file("Hello, world.", "output.wav", voice="Bruno", speed=0.9)
# List available voices
print(model.available_voices)
# ['Bella', 'Jasper', 'Luna', 'Bruno', 'Rosie', 'Hugo', 'Kiki', 'Leo']pip install -r requirements_gpu.txtm = KittenTTS("KittenML/kitten-tts-mini-0.8", backend="cuda")Check out example_cuda.py
KittenTTS(model_name, cache_dir=None)Load a model from Hugging Face Hub.
| Parameter | Type | Default | Description |
|---|---|---|---|
model_name |
str |
"KittenML/kitten-tts-nano-0.8" |
Hugging Face repository ID |
cache_dir |
str |
None |
Local directory for caching downloaded model files |
model.generate(text, voice, speed, clean_text)Synthesize speech from text, returning a NumPy array of audio samples at 24 kHz.
| Parameter | Type | Default | Description |
|---|---|---|---|
text |
str |
-- | Input text to synthesize |
voice |
str |
"expr-voice-5-m" |
Voice name (see available voices) |
speed |
float |
1.0 |
Speech speed multiplier |
clean_text |
bool |
False |
Preprocess text (expand numbers, currencies, etc.) |
model.generate_to_file(text, output_path, voice, speed, sample_rate, clean_text)Synthesize speech and write directly to an audio file.
| Parameter | Type | Default | Description |
|---|---|---|---|
text |
str |
-- | Input text to synthesize |
output_path |
str |
-- | Path to save the audio file |
voice |
str |
"expr-voice-5-m" |
Voice name |
speed |
float |
1.0 |
Speech speed multiplier |
sample_rate |
int |
24000 |
Audio sample rate in Hz |
clean_text |
bool |
True |
Preprocess text (expand numbers, currencies, etc.) |
normalize_text(text, locale="en-US", return_spans=False)Normalize text for TTS without generating audio.
from kittentts import normalize_text
normalized = normalize_text("Dr. Rivera paid $12.50 at 3:05 p.m.")
# "Doctor Rivera paid twelve dollars and fifty cents at three oh five p m."
result = normalize_text("Fig. 2", return_spans=True)
print(result.text)
print(result.spans)When return_spans=True, the result includes original-to-normalized character spans for changed segments such as abbreviations, dates, times, numbers, currency, URLs, and punctuation.
model.available_voicesReturns a list of available voice names: ['Bella', 'Jasper', 'Luna', 'Bruno', 'Rosie', 'Hugo', 'Kiki', 'Leo']
A virtual environment (conda, venv, or similar) is recommended to avoid dependency conflicts.
We offer commercial support for teams integrating Kitten TTS into their products. This includes integration assistance, custom voice development, and enterprise licensing.
Contact us or email [email protected] to discuss your requirements.
This project is licensed under the Apache License 2.0.
流式传输支持
谐波源: 会话非确定性 (图中 RNG) 和 float32 阶段累加 (fp16 / 长形式限制)
如何找到这种小型的语音转文字 LLM 模型,请不要低声细语
特性:FunASR/SenseVoice 用于训练数据标注
建筑学?
培训详情
将 Python 替换为 Rust
为什么 Misaki 包含在要求文件中?
特性: 启用系统原生的 espeak-ng。
无法使用 pip 安装