一款实时的无声语音识别工具。
A visual speech recognition (VSR) tool that reads your lips in real-time and types whatever you silently mouth. Runs fully locally.
Relies on a model trained on the Lip Reading Sentences 3 dataset as part of the Auto-AVSR project.
Watch a demo of Chaplin here.
cd into it:git clone https://github.com/amanvirparhar/chaplin
cd chaplin
./setup.sh
...which will automatically download the required model files from Hugging Face Hub and place them in the appropriate directories:chaplin/
├── benchmarks/
├── LRS3/
├── language_models/
├── lm_en_subword/
├── models/
├── LRS3_V_WER19.1/
├── ...
ollama, and pull the qwen3:4b model.uv.uv run --with-requirements requirements.txt --python 3.12 main.py config_filename=./configs/LRS3_V_WER19.1.ini detector=mediapipe
option key (Mac) or the alt key (Windows/Linux), and start mouthing words.option key (Mac) or the alt key (Windows/Linux) again. The raw VSR output will get logged in your terminal, and the LLM-corrected version will be typed at your cursor.q.暂无开放 Issues,或尚未同步最近议题。