百科
.dev
全部条目
趋势榜
开源项目
技术资讯
提交条目
Language
中文
English
登录
< 返回工具列表
Language
中文
English
W
whisperX
> 编程语言
开源
WhisperX: 自动语音识别并带有Word-level时标( & Diarization)
23.4K stars
0 点赞
0 次浏览
访问官网
GitHub
点赞
收藏
分享体验
核心特点
•
⚡️ Batched inference for 70x realtime transcription using whisper large-v2
•
🪶 faster-whisper backend, requires <8GB gpu memory for large-v2 with beam_size=5
•
🎯 Accurate word-level timestamps using wav2vec2 alignment
•
👯♂️ Multispeaker ASR using speaker diarization from pyannote-audio (speaker ID labels)
•
🗣️ VAD preprocessing, reduces hallucination & batching with no WER degradation
•
1st place at Ego4d transcription challenge 🏆
•
_WhisperX_ accepted at INTERSPEECH 2023
•
v3 transcript segment-per-sentence: using nltk sent_tokenize for better subtitlting & better diarization
•
v3 released, 70x speed-up open-sourced. Using batched whisper with faster-whisper backend!
•
v2 released, code cleanup, imports whisper library VAD filtering is now turned on by default, as in the paper.
>
标签
Python
asr
speech
speech-recognition
speech-to-text
>
讨论区
最新
发表
暂无评论,来聊聊你的看法吧
> 工具信息
发布日期
2026年8月1日
最后更新
2026年9月9日
分类
编程语言
定价
开源
> 相关工具
T
TypeScript
JavaScript 的超集,为前端与全栈提供静态类型
P
Python
通用编程语言,广泛用于 Web、数据与 AI
G
Go
Google 推出的简洁高效系统语言
报告问题