Baike.dev
All toolsTrendingOpen sourceNewsSubmit
Log in
< 返回工具列表
S

sherpa-onnx

> 编程语言
开源

Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet conne

13.9K stars0 点赞0 次浏览
访问官网GitHub

工具介绍

Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet conne

Supported functions

Speech recognition [Speech synthesis][tts-url] [Source separation][ss-url] ✔️ ✔️ ✔️ Speaker identification [Speaker diarization][sd-url] Speaker verification ✔️ ✔️ ✔️ [Spoken Language identification][slid-url] [Audio tagging][at-url] [Voice activity detection][vad-url] ✔️ ✔️ ✔️ [Keyword spotting][kws-url] [Add punctuation][punct-url] [Speech enhancement][se-url] ✔️ ✔️ ✔️

Supported platforms

Architecture Android iOS Windows macOS linux HarmonyOS x64 ✔️ ✔️ ✔️ ✔️ ✔️ x86 ✔️ ✔️ arm64 ✔️ ✔️ ✔️ ✔️ ✔️ ✔️ arm32 ✔️ ✔️ ✔️ riscv64 ✔️

Supported programming languages

1. C++ 2. C 3. Python 4. JavaScript ✔️ ✔️ ✔️ ✔️ 5. Java 6. C# 7. Kotlin 8. Swift ✔️ ✔️ ✔️ ✔️ 9. Go 10. Dart 11. Rust 12. Pascal ✔️ ✔️ ✔️ ✔️

It also supports WebAssembly.

Supported frameworks

Flutter

Architecture Android iOS Windows macOS Linux Web x64 ✔️ ✔️ ✔️ ✔️ ✔️ x86 ✔️ ✔️ arm64 ✔️ ✔️ ✔️ ✔️ ✔️ ✔️ arm32 ✔️ ✔️
  • Pre-built Flutter demo apps: flutter releases

Tauri

Architecture Android iOS Windows macOS Linux x64 ✔️ ✔️ ✔️ ✔️ x86 ✔️ arm64 ✔️ ✔️ ✔️ ✔️ ✔️ arm32 ✔️
  • Pre-built Tauri demo apps: tauri releases

Supported NPUs

[1. Rockchip NPU (RKNN)][rknpu-doc] [2. Qualcomm NPU (QNN)][qnn-doc] [3. Ascend NPU][ascend-doc] ✔️ ✔️ ✔️ [4. Axera NPU][axera-npu] 5. Intel NPU (OpenVINO) ✔️ ✔️

Join our discord

Introduction

This repository supports running the following functions locally

  • Speech-to-text (i.e., ASR); both streaming and non-streaming are supported
  • Text-to-speech (i.e., TTS)
  • Speaker diarization
  • Speaker identification
  • Speaker verification
  • Spoken language identification
  • Audio tagging
  • VAD (e.g., [silero-vad][silero-vad])
  • Speech enhancement (e.g., [gtcrn][gtcrn], DPDFNet)
  • Keyword spotting
  • Source separation (e.g., [spleeter][spleeter], [UVR][UVR])

on the following platforms and operating systems:

  • x86, x86_64, 32-bit ARM, 64-bit ARM (arm64, aarch64), RISC-V (riscv64), RK NPU, Ascend NPU
  • Linux, macOS, Windows, openKylin
  • Android, WearOS
  • iOS
  • HarmonyOS
  • NodeJS
  • WebAssembly
  • [NVIDIA Jetson Orin NX][NVIDIA Jetson Orin NX] (Support running on both CPU and GPU)
  • [NVIDIA Jetson Nano B01][NVIDIA Jetson Nano B01] (Support running on both CPU and GPU)
  • [Raspberry Pi][Raspberry Pi]
  • [RV1126][RV1126]
  • [LicheePi4A][LicheePi4A]
  • [VisionFive 2][VisionFive 2]
  • [旭日X3派][旭日X3派]
  • [爱芯派][爱芯派]
  • [RK3588][RK3588]
  • [SpacemiT-K1][SpacemiT-K1]
  • [SpacemiT-K3][SpacemiT-K3]
  • etc

with the following APIs

  • C++, C, Python, Go, C#
  • Java, Kotlin, JavaScript
  • Swift, Rust
  • Dart, Object Pascal

Links for Huggingface Spaces

You can visit the following Huggingface spaces to try sherpa-onnx without installing anything. All you need is a browser. Description URL 中国镜像 Speaker diarization [Click me][hf-space-speaker-diarization] [镜像][hf-space-speaker-diarization-cn] Speech recognition [Click me][hf-space-asr] [镜像][hf-space-asr-cn] Speech recognition with [Whisper][Whisper] [Click me][hf-space-asr-whisper] [镜像][hf-space-asr-whisper-cn] Speech synthesis [Click me][hf-space-tts] [镜像][hf-space-tts-cn] Generate subtitles [Click me][hf-space-subtitle] [镜像][hf-space-subtitle-cn] Audio tagging [Click me][hf-space-audio-tagging] [镜像][hf-space-audio-tagging-cn] Source separation [Click me][hf-space-source-separation] [镜像][hf-space-source-separation-cn] Spoken language identification with [Whisper][Whisper] [Click me][hf-space-slid-whisper] [镜像][hf-space-slid-whisper-cn]

We also have spaces built using WebAssembly. They are listed below:

Description Huggingface space ModelScope space Voice activity detection with [silero-vad][silero-vad] [Click me][wasm-hf-vad] [地址][wasm-ms-vad] Real-time speech recognition (Chinese + English) with Zipformer [Click me][wasm-hf-streaming-asr-zh-en-zipformer] [地址][wasm-hf-streaming-asr-zh-en-zipformer] Real-time speech recognition (Chinese + English) with Paraformer [Click me][wasm-hf-streaming-asr-zh-en-paraformer] [地址][wasm-ms-streaming-asr-zh-en-paraformer] Real-time speech recognition (Chinese + English + Cantonese) with [Paraformer-large][Paraformer-large] [Click me][wasm-hf-streaming-asr-zh-en-yue-paraformer] [地址][wasm-ms-streaming-asr-zh-en-yue-paraformer] Real-time speech recognition (English) [Click me][wasm-hf-streaming-asr-en-zipformer] [地址][wasm-ms-streaming-asr-en-zipformer] VAD + speech recognition (Chinese) with Zipformer CTC [Click me][wasm-hf-vad-asr-zh-zipformer-ctc-07-03] [地址][wasm-ms-vad-asr-zh-zipformer-ctc-07-03] VAD + speech recognition (Chinese + English + Korean + Japanese + Cantonese) with [SenseVoice][SenseVoice] [Click me][wasm-hf-vad-asr-zh-en-ko-ja-yue-sense-voice] [地址][wasm-ms-vad-asr-zh-en-ko-ja-yue-sense-voice] VAD + speech recognition (English) with [Whisper][Whisper] tiny.en [Click me][wasm-hf-vad-asr-en-whisper-tiny-en] [地址][wasm-ms-vad-asr-en-whisper-tiny-en] VAD + speech recognition (English) with [Moonshine tiny][Moonshine tiny] [Click me][wasm-hf-vad-asr-en-moonshine-tiny-en] [地址][wasm-ms-vad-asr-en-moonshine-tiny-en] VAD + speech recognition (English) with Zipformer trained with [GigaSpeech][GigaSpeech] [Click me][wasm-hf-vad-asr-en-zipformer-gigaspeech] [地址][wasm-ms-vad-asr-en-zipformer-gigaspeech] VAD + speech recognition (Chinese) with Zipformer trained with [WenetSpeech][WenetSpeech] [Click me][wasm-hf-vad-asr-zh-zipformer-wenetspeech] [地址][wasm-ms-vad-asr-zh-zipformer-wenetspeech] VAD + speech recognition (Japanese) with Zipformer trained with [ReazonSpeech][ReazonSpeech] [Click me][wasm-hf-vad-asr-ja-zipformer-reazonspeech] [地址][wasm-ms-vad-asr-ja-zipformer-reazonspeech] VAD + speech recognition (Thai) with Zipformer trained with [GigaSpeech2][GigaSpeech2] [Click me][wasm-hf-vad-asr-th-zipformer-gigaspeech2] [地址][wasm-ms-vad-asr-th-zipformer-gigaspeech2] VAD + speech recognition (Chinese 多种方言) with a [TeleSpeech-ASR][TeleSpeech-ASR] CTC model [Click me][wasm-hf-vad-asr-zh-telespeech] [地址][wasm-ms-vad-asr-zh-telespeech] VAD + speech recognition (English + Chinese, 及多种中文方言) with Paraformer-large [Click me][wasm-hf-vad-asr-zh-en-paraformer-large] [地址][wasm-ms-vad-asr-zh-en-paraformer-large] VAD + speech recognition (English + Chinese, 及多种中文方言) with Paraformer-small [Click me][wasm-hf-vad-asr-zh-en-paraformer-small] [地址][wasm-ms-vad-asr-zh-en-paraformer-small] VAD + speech recognition (多语种及多种中文方言) with [Dolphin][Dolphin]-base [Click me][wasm-hf-vad-asr-multi-lang-dolphin-base] [地址][wasm-ms-vad-asr-multi-lang-dolphin-base] Speech synthesis (Piper, English) [Click me][wasm-hf-tts-piper-en] [地址][wasm-ms-tts-piper-en] Speech synthesis (Piper, German) [Click me][wasm-hf-tts-piper-de] [地址][wasm-ms-tts-piper-de] Speech synthesis (Matcha, Chinese) [Click me][wasm-hf-tts-matcha-zh] [地址][wasm-ms-tts-matcha-zh] Speech synthesis (Matcha, English) [Click me][wasm-hf-tts-matcha-en] [地址][wasm-ms-tts-matcha-en] Speech synthesis (Matcha, Chinese+English) [Click me][wasm-hf-tts-matcha-zh-en] [地址][wasm-ms-tts-matcha-zh-en] Speaker diarization [Click me][wasm-hf-speaker-diarization] [地址][wasm-ms-speaker-diarization] Voice cloning with ZipVoice (Chinese+English) [Click me][wasm-hf-voice-cloning-zipvoice] [地址][wasm-ms-voice-cloning-zipvoice] Voice cloning with Pocket TTS (English) [Click me][wasm-hf-voice-cloning-pocket] [地址][wasm-ms-voice-cloning-pocket]

Links for pre-built Android APKs

You can find pre-built Android APKs for this repository in the following table Description URL 中国用户

核心特点

  • •Pre-built Flutter demo apps: [flutter releases][flutter-releases]
  • •Pre-built Tauri demo apps: [tauri releases][tauri-releases]
  • •Speech-to-text (i.e., ASR); both streaming and non-streaming are supported
  • •Text-to-speech (i.e., TTS)
  • •Speaker diarization
  • •Speaker identification
  • •Speaker verification
  • •Spoken language identification
  • •Audio tagging
  • •VAD (e.g., [silero-vad][silero-vad])

> 标签

C++aarch64androidarm32asr

暂无评论,来聊聊你的看法吧

> 工具信息

发布日期2026年8月1日
最后更新2026年9月9日
分类编程语言
定价开源

> 相关工具

T
TypeScript
JavaScript 的超集,为前端与全栈提供静态类型
P
Python
通用编程语言,广泛用于 Web、数据与 AI
G
Go
Google 推出的简洁高效系统语言