百科.dev
全部条目AI 编程趋势榜开源项目技术资讯提交条目
登录
< 返回工具列表
F

Fun-ASR

> AI 编程
开源

基于 LLM 的开源 ASR 模型系列,适用于中文、方言、口音和多语言语音,包含 FunASR、vLLM、流式处理和 LLaMA.cpp 运行时。

1.5K stars0 点赞0 次浏览
访问官网GitHub

工具介绍

基于 LLM 的开源 ASR 模型系列,适用于中文、方言、口音和多语言语音,包含 FunASR、vLLM、流式处理和 LLaMA.cpp 运行时。

Fun-ASR

「简体中文」|「English」|「日本語」|「한국어」

Fun-ASR is a family of end-to-end speech recognition models from Tongyi Lab. Checkpoint capabilities are distinct: Fun-ASR-Nano-2512 is trained on tens of millions of hours of speech and supports Chinese, English, Japanese, and Chinese dialects and accents; Fun-ASR-MLT-Nano-2512 is an 800M multilingual checkpoint trained on hundreds of thousands of hours and supports 31 languages. Both checkpoints integrate with FunASR for inference and serving.

Model Name Task Details Training Data Parameters
Fun-ASR-Nano
(⭐ HF / Transformers · HF / FunASR)
Speech recognition supports Chinese, English, and Japanese. Chinese includes support for 7 dialects (Wu, Cantonese, Min, Hakka, Gan, Xiang, Jin) and 26 regional accents (Henan, Shanxi, Hubei, Sichuan, Chongqing, Yunnan, Guizhou, Guangdong, Guangxi and more than 20 other regions). English and Japanese cover multiple regional accents. Additional features include lyric recognition and rap speech recognition. Tens of millions of hours 800M
Fun-ASR-MLT-Nano
(⭐ )
Speech recognition supports Chinese, English, Cantonese, Japanese, Korean, Vietnamese, Indonesian, Thai, Malay, Filipino, Arabic, Hindi, Bulgarian, Croatian, Czech, Danish, Dutch, Estonian, Finnish, Greek, Hungarian, Irish, Latvian, Lithuanian, Maltese, Polish, Portuguese, Romanian, Slovak, Slovenian, and Swedish: 31 languages in total. Hundreds of thousands of hours 800M

What's New

  • FunASR 1.4.15 is the current Python release for source installs, MOSS discovery, and realtime or industrial deployment. Install with python -m pip install -U "funasr==1.4.15". Release ->
  • MOSS-Transcribe-Diarize is a third-party OpenMOSS model for offline long-form transcription, timestamps, and anonymous speaker labels, with FunASR service, Docker, Kubernetes, vLLM, SGLang, LocalAI, and FunClip deployment paths. Deploy MOSS ->
  • Production deployment covers realtime WebSocket serving, native vLLM batch/streaming paths, and verified llama.cpp / GGUF packages for Linux, macOS, and Windows. Runtime v0.2.6 -> · vLLM guide ->

Native Transformers quickstart

Transcribe with the released Transformers 5.17.0 package. No toolkit installation or remote Python code is needed. Base Nano supports Chinese, English and Japanese; the 31-language MLT checkpoint is separate.

python -m pip install 'transformers==5.17.0' 'torch==2.10.0' 'torchaudio==2.10.0' 'librosa==0.11.0' 'soundfile==0.13.1'

Full Python recipe · Local audio, batches and keywords · Notebook · Space

…

Core Features

Fun-ASR focuses on high-precision speech recognition, checkpoint-specific multilingual support, and industry customization capabilities.

  • Far-field High-noise Recognition: Deeply optimized for far-distance sound pickup and high-noise scenarios (such as conference rooms, in-vehicle environments, industrial sites, etc.), improving recognition accuracy to 93%.
  • Chinese Dialects and Regional Accents:
    • Supports 7 major dialects: Wu, Cantonese, Min, Hakka, Gan, Xiang, Jin
    • Covers 26 regional accents: including Henan, Shaanxi, Hubei, Sichuan, Chongqing, Yunnan, Guizhou, Guangdong, Guangxi and more than 20 other regions
  • Checkpoint-specific language coverage: Fun-ASR-Nano supports Chinese, English, Japanese, and Chinese dialects and accents. Fun-ASR-MLT-Nano supports 31 languages, with emphasis on East and Southeast Asian languages.
  • Music Background Lyric Recognition: Enhanced speech recognition performance under music background interference, supporting accurate recognition of lyric content in songs.

Environment Setup

git clone https://github.com/QwenAudio/Fun-ASR.git
cd Fun-ASR
pip install -r requirements.txt

Capability boundaries

  • Checkpoint-native character timestamps (hub-specific checkpoint state)

    The current ModelScope FunAudioLLM/Fun-ASR-Nano-2512 checkpoint includes all 86 trained ctc_decoder.* / ctc.* tensors (model.pt SHA-256 81fec8616083c69377f3ceef36aba3655660ee0ca69a5d4a1e9810cd340ca499) and produces native CTC timestamps. The Hugging Face checkpoint at revision 272c57b82523ada6fd87095e955f8e29100979ab is still the older text-only artifact (model.pt SHA-256 55ae0d2fee369f0f11cce0795f6927934ad17cf11b278a7e56a51272074160bb) with no CTC tensors. When this repository's current model.py is used with funasr>=1.3.26, incomplete checkpoints fail closed: transcription remains available, but timestamps are omitted instead of returning random 60 ms alignments. Use hub="ms" for checkpoint-native timestamps until the Hugging Face artifact and remote code are synchronized. See issue #70 and FunASR #3496.

  • Checkpoint-native speaker diarization

    Fun-ASR-Nano and Fun-ASR-MLT-Nano do not emit speaker labels by themselves. Compose them in FunASR with the separate fsmn-vad and cam++ models, as shown below. For one-pass anonymous diarization with transcription and timestamps, use the third-party OpenMOSS MOSS-Transcribe-Diarize deployment guide. It is a separate model rather than a Fun-ASR-Nano checkpoint feature.

  • Model training

Usage ️

Inference

Run on CPU / edge — llama.cpp / GGUF (no GPU, no Python)

Run Fun-ASR-Nano as a single self-contained binary — like whisper.cpp but for FunASR, with strong Chinese accuracy. Built-in FSMN-VAD, no Python at runtime.

bash runtime/llama.cpp/download-funasr-model.sh nano ./gguf
llama-funasr-cli --enc ./gguf/funasr-encoder-f16.gguf -m ./gguf/qwen3-0.6b-q8_0.gguf -a audio.wav --vad ./gguf/fsmn-vad.gguf

fsmn-vad.gguf is hosted in the shared FunAudioLLM/fsmn-vad-GGUF repo, not inside the Nano GGUF repo. The nano downloader above fetches it automatically; to fetch only VAD from the Hugging Face UI/CLI, use:

hf download FunAudioLLM/fsmn-vad-GGUF --include "*.gguf" --local-dir ./gguf

Prebuilt binaries: Releases · Download & quickstart: funasr.com/llama-cpp · GGUF: Nano encoder/LLM · FSMN-VAD · Docs & benchmarks: runtime/llama.cpp/

Using funasr for inference

…

Faster batch transcription (no vLLM)

When transcribing long audio or many files on the funasr (PyTorch) path, pass batch_size_s to batch the VAD segments through the LLM decoder together. This greatly improves GPU utilization:

res = model.generate(
    input=[wav_path],
    cache={},
    language="中文",
    itn=True,
    batch_size_s=120,   # batch VAD segments up to ~120s of audio per LLM call
)

On Fun-ASR-Nano-2512 (184 Chinese files / 11,539 s, single H100) this is about 1.6x faster than the default per-segment decoding (RTFx 19.8 -> 31.8) with no loss in accuracy. For the highest throughput, use the vLLM path below.

Speaker Diarization

This example is a composed FunASR pipeline: FSMN-VAD segments the audio, Fun-ASR-Nano transcribes it, CAM++ assigns speaker labels, and CT-Punc restores punctuation. The start and end values are VAD segment boundaries, not reliable checkpoint-native character timestamps.

…

Direct Inference

from model import FunASRNano

def main():
    model_dir = "FunAudioLLM/Fun-ASR-Nano-2512"
    m, kwargs = FunASRNano.from_pretrained(model=model_dir, device="cuda:0")
    m.eval()

    wav_path = f"{kwargs['model_path']}/example/zh.mp3"
    res = m.inference(data_in=[wav_path], **kwargs)
    text = res[0][0]["text"]
    print(text)

if __name__ == "__main__":
    main()
Parameter Description (click to expand)
  • model_dir: Model name or local disk model path.
  • trust_remote_code: Whether to trust remote code for loading custom model implementations.
  • remote_code: Specify the location of specific model code (e.g., model.py in the current directory), supporting both absolute and relative paths.
  • device: Specify the device to use, such as "cuda:0" or "cpu".

vLLM High-Throughput Inference

Fun-ASR natively integrates the vLLM engine for high-throughput batch inference and production-grade real-time streaming service.

Full guide: docs/vllm_guide.md | API docs: modelscope.github.io/FunASR/vllm.html

Three Modes

Mode Use Case Entry
Offline Batch Large-scale transcription AutoModelVLLM
Streaming SDK Real-time subtitles FunASRNanoStreamingVLLM
WebSocket Service Production deployment serve_realtime_ws.py

Offline Batch Inference (3-5x faster)

from funasr.auto.auto_model_vllm import AutoModelVLLM

model = AutoModelVLLM(
    model="FunAudioLLM/Fun-ASR-Nano-2512",
    tensor_parallel_size=2,      # Multi-GPU
    gpu_memory_utilization=0.8,
)

results = model.generate(
    ["audio1.wav", "audio2.wav", "audio3.wav"],
    language="中文",
    hotwords=["张三", "北京"],
)
for r in results:
    print(f"[{r['key']}] {r['text']}")

Long audio: AutoModelVLLM decodes each input in a single pass, so a long recording (e.g. a multi-minute meeting) can be truncated — pre-segment it with VAD and pass the segments, or use the high-level AutoModel(model=..., vad_model="fsmn-vad"), which segments long audio automatically.

Real-time WebSocket Service

# Start server (with dynamic VAD + speaker diari

Issues· 7 开放

查看全部 Issues在 GitHub 打开

暂无开放 Issues,或尚未同步最近议题。

> 标签

C31-languagesasraudioaudio-language-model

暂无评论,来聊聊你的看法吧

> 工具信息

发布日期2026年8月1日
最后更新2026年9月17日
分类AI 编程
定价开源

> 相关工具

G
GitHub Copilot
GitHub 官方 AI 编程助手,覆盖补全、Chat 与 Agent 模式。
C
Cursor
AI 原生代码编辑器,对话改代码、多文件 Agent 与规则体系是其核心。
S
skills
Skills for Real Engineers. Straight from my .agents directory.