Multilingual Automatic Speech Recognition with word-level timestamps and confidence
Multilingual Automatic Speech Recognition with word-level timestamps and confidence
Multilingual Automatic Speech Recognition with word-level timestamps and confidence.
Whisper is a set of multi-lingual, robust speech recognition models trained by OpenAI that achieve state-of-the-art results in many languages. Whisper models were trained to predict approximate timestamps on speech segments (most of the time with 1-second accuracy), but they cannot originally predict word timestamps. This repository proposes an implementation to predict word timestamps and provide a more accurate estimation of speech segments when transcribing with Whisper models. Besides, a confidence score is assigned to each word and each segment.
The approach is based on Dynamic Time Warping (DTW) applied to cross-attention weights, as demonstrated by this notebook by Jong Wook Kim. There are some additions to this notebook:
whisper-timestamped is able to process long files with little additional memory compared to the regular use of the Whisper model.whisper-timestamped is an extension of the openai-whisper Python package and is meant to be compatible with any version of openai-whisper.
It provides more efficient/accurate word timestamps, along with those additional features:
Disclaimer: Please note that this extension is intended for experimental purposes and may significantly impact performance. We are not responsible for any issues or inefficiencies that arise from its use.
An alternative relevant approach to recovering word-level timestamps involves using wav2vec models that predict characters, as successfully implemented in whisperX. However, these approaches have several drawbacks that are not present in approaches based on cross-attention weights such as whisper_timestamped. These drawbacks include:
An alternative approach that does not require an additional model is to look at the probabilities of timestamp tokens estimated by the Whisper model after each (sub)word token is predicted. This was implemented, for instance, in whisper.cpp and stable-ts. However, this approach lacks robustness because Whisper models have not been trained to output meaningful timestamps after each word. Whisper models tend to predict timestamps only after a certain number of words have been predicted (typically at the end of a sentence), and the probability distribution of timestamps outside this condition may be inaccurate. In practice, these methods can produce results that are totally out-of-sync on some periods of time (we observed this especially when there is jingle music). Also, the timestamp precision of Whisper models tends to be rounded to 1 second (as in many video subtitles), which is too inaccurate for words, and reaching better accuracy is tricky.
Requirements:
python3 (version higher or equal to 3.7, at least 3.9 is recommended)ffmpeg (see instructions for installation on the whisper repository)You can install whisper-timestamped either by using pip:
pip3 install whisper-timestamped
or by cloning this repository and running installation:
git clone https://github.com/linto-ai/whisper-timestamped
cd whisper-timestamped/
python3 setup.py install
If you want to plot alignment between audio timestamps and words (as in this section), you also need matplotlib:
pip3 install matplotlib
If you want to use VAD option (Voice Activity Detection before running Whisper model), you also need torchaudio and onnxruntime:
pip3 install onnxruntime torchaudio
If you want to use finetuned Whisper models from the Hugging Face Hub, you also need transformers:
pip3 install transformers
A docker image of about 9GB can be built using:
git clone https://github.com/linto-ai/whisper-timestamped
cd whisper-timestamped/
docker build -t whisper_timestamped:latest .
If you don't have a GPU (or don't want to use it), then you don't need to install the CUDA dependencies. You should then just install a light version of torch before installing whisper-timestamped, for instance as follows:
pip3 install \
torch==1.13.1+cpu \
torchaudio==0.13.1+cpu \
-f https://download.pytorch.org/whl/torch_stable.html
A specific docker image of about 3.5GB can also be built using:
git clone https://github.com/linto-ai/whisper-timestamped
cd whisper-timestamped/
docker build -t whisper_timestamped_cpu:latest -f Dockerfile.cpu .
When using pip, the library can be updated to the latest version using:
pip3 install --upgrade --no-deps --force-reinstall git+https://github.com/linto-ai/whisper-timestamped
A specific version of openai-whisper can be used by running, for example:
pip3 install openai-whisper==20230124
In Python, you can use the function whisper_timestamped.transcribe(), which is similar to the function whisper.transcribe():
import whisper_timestamped
help(whisper_timestamped.transcribe)
The main difference with whisper.transcribe() is that the output will include a key "words" for all segments, with the word start and end position. Note that the word will include punctuation. See the example below.
Besides, the default decoding options are different to favour efficient decoding (greedy decoding instead of beam search, and no temperature sampling fallback). To have same default as in whisper, use beam_size=5, best_of=5, temperature=(0.0, 0.2, 0.4, 0.6, 0.8, 1.0).
There are also additional options related to word alignement.
In general, if you import whisper_timestamped instead of whisper in your Python script and use transcribe(model, ...) instead of model.transcribe(...), it should do the job:
import whisper_timestamped as whisper
audio = whisper.load_audio("AUDIO.wav")
model = whisper.load_model("tiny", device="cpu")
result = whisper.transcribe(model, audio, language="fr")
import json
print(json.dumps(result, indent = 2, ensure_ascii = False))
Note that you can use a finetuned Whisper model from HuggingFace or a local folder by using the load_model method of whisper_timestamped. For instance, if you want to use whisper-large-v2-nob, you can simply do the following:
import whisper_timestamped as whisper
model = whisper.load_model("NbAiLab/whisper-large-v2-nob", device="cpu")
# ...
You can also use whisper_timestamped on the command line, similarly to whisper. See help with:
whisper_timestamped --help
The main differences with whisper CLI are:
Output files:
Some default options are different:
--output_dir . for Whisper default.--verbose True for Whisper default.--accurate (which is an alias for --beam_size 5 --temperature_increment_on_fallback 0.2 --best_of 5).There are some additional specific options:
--compute_confidence to enable/disable the computation of confidence scores for each word.--punctuations_with_words to decide whether punctuation marks should be included or not with preceding words.An example command to process several files using the tiny model and output the results in the current folder, as would be done by default with whisper, is as follows:
whisper_timestamped audio1.flac audio2.mp3 audio3.wav --model tiny --output_dir .
Note that you can use a fine-tuned Whisper model from HuggingFace or a local folder. For instance, if you want to use the whisper-large-v2-nob model, you can simply do the following:
whisper_timestamped --model NbAiLab/whisper-large-v2-nob <...>
In addition to the main transcribe function, whisper-timestamped provides some utility functions:
remove_non_speechRemove non-speech segments from audio using Voice Activity Detection (VAD).
from whisper_timestamped import remove_non_speech
audio_speech, segments, convert_timestamps = remove_non_speech(audio, vad="silero")
load_modelLoad a Whisper model from a given name or path, including support for fine-tuned models from HuggingFace.
from whisper_timestamped import load_model
model = load_model("NbAiLab/whisper-large-v2-nob", device="cpu")
Note that you can use the `
No open issues yet, or sync has not completed.