openai-whisper-cpu
yt-whisper
openai-whisper-cpu | yt-whisper | |
---|---|---|
5 | 3 | |
221 | 1,316 | |
- | - | |
10.0 | 0.0 | |
over 1 year ago | 4 months ago | |
Jupyter Notebook | Python | |
MIT License | MIT License |
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.
openai-whisper-cpu
-
How to run Llama 13B with a 6GB graphics card
I feel the same.
For example some stats from Whisper [0] (audio transcoding) show the following for the medium model (see other models in the link):
---
GPU medium fp32 Linear 1.7s
CPU medium fp32 nn.Linear 60.7
CPU medium qint8 (quant) nn.Linear 23.1
---
So the same model runs 35.7 times faster on GPU, and compared to an CPU-optimized model still 13.6.
I was expecting around an order or magnitude of improvement. Then again, I do not know if in the case of this article the entire model was in the GPU, or just a fraction of it (22 layers), which might explain the result.
[0] https://github.com/MiscellaneousStuff/openai-whisper-cpu
-
Whispers AI Modular Future
According to https://github.com/MiscellaneousStuff/openai-whisper-cpu the medium model needs 1.7 seconds to transcribe 30 seconds of audio when run on a GPU.
-
[P] Transcribe any podcast episode in just 1 minute with optimized OpenAI/whisper
There is a very simple method built-in to PyTorch which can give you over 3x speed improvement for the large model, which you could also combine with the method proposed in this post. https://github.com/MiscellaneousStuff/openai-whisper-cpu
-
[D] How to get the fastest PyTorch inference and what is the "best" model serving framework?
For CPU inference, model quantization is a very easy to apply method with great average speedups which is already built-in to PyTorch. For example, I applied dynamic quantization to the OpenAI Whisper model (speech recognition) across a range of model sizes (ranging from tiny which had 39M params to large which had 1.5B params). Refer to the below table for performance increases:
-
[P] OpenAI Whisper - 3x CPU Inference Speedup
GitHub
yt-whisper
-
me in real life
#28 transcribe local files
-
YouTubeTranscript.com
Or even better, yt-whisper, which uses OpenAI's Whisper speech to text. I guess it'd be better to first check whether the video has captions first before Whispering, so maybe both your command and this one could be used together.
https://github.com/m1guelpf/yt-whisper
-
[P] Transcribe any podcast episode in just 1 minute with optimized OpenAI/whisper
With minimal changes to https://github.com/m1guelpf/yt-whisper i got a setup to transcribe subs from YouTube videos or local files bit it might take an hour or so running the large model on my CPU.
What are some alternatives?
llama-cpp-python - Python bindings for llama.cpp
subsai - 🎞️ Subtitles generation tool (Web-UI + CLI + Python package) powered by OpenAI's Whisper and its variants 🎞️
intel-extension-for-pytorch - A Python package for extending the official PyTorch that can easily obtain performance on Intel platform
ChatGPT-YouTube-summarizer - This Chrome extension lets you summarize YouTube videos using the ChatGPT.
whisperX - WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
tldwol - Web API that summarizes multimedia from various sources using modern AI tools.
FlexGen - Running large language models on a single GPU for throughput-oriented scenarios.
auto-subtitle - Automatically generate and overlay subtitles for any video.
buzz - Buzz transcribes and translates audio offline on your personal computer. Powered by OpenAI's Whisper.
gentle - gentle forced aligner
kernl - Kernl lets you run PyTorch transformer models several times faster on GPU with a single line of code, and is designed to be easily hackable.
malayalam_english_subtitle_generator - Malayalam to English Subtitle Generator for audio files using OpenAI's Whisper.