openWakeWord
whisper-turbo
openWakeWord | whisper-turbo | |
---|---|---|
5 | 11 | |
457 | 1,573 | |
- | - | |
8.4 | 8.9 | |
about 1 month ago | 3 months ago | |
Jupyter Notebook | TypeScript | |
Apache License 2.0 | Apache License 2.0 |
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.
openWakeWord
-
OpenAI releases Whisper v3, new generation open source ASR model
https://github.com/dscripka/openWakeWord
Balancing wake reliability vs false wake activation is a tricky balance. OWW is decent but could certainly be better.
It's used with Home Assistant now so I expect the training data and implementation overall to get significantly better fairly soon.
-
Distil-Whisper: distilled version of Whisper that is 6 times faster, 49% smaller
There's also OpenWakeWord[0]. The models are readily available in tflite and ONNX formats and are impressively "light" in terms of compute requirements and performance.
It should be possible.
[0] - https://github.com/dscripka/openWakeWord
-
Real-Time Noise Suppression for PipeWire writen in Rust
hey, quick question. do you mind if I use your stft function in the speech preprocessing library I've been working on? we've been trying to add support for doing mel spectrograms to build a runner for openwakeword, but progress is pretty slow because I've been soloing something I really don't have the right background for(I've never directly studied or worked with signal processing)
-
I'm new to Rust but want to contribute
potentially build another runner for open wakeword
-
I want to contribute in a big project
here's what's on the pipeline next: - finish mel-spectrogram implementation - publish initial version on crates - move python caching rust side - finish implementing in the precise rust port - potentially build another runner for (open wakeword)[https://github.com/dscripka/openWakeWord] - build an android app that supports user-defined wakewords and has some popular defaults to load. ps not a voice assistant, just the thing that activates the voice assistant.
whisper-turbo
- Whisper Turbo: speech recognition in the browser using WebGPU
-
Show HN: Shadeup – A language that makes WebGPU easier
Even just looking at the ability to accelerate llms in the browser on any device without an installation is awesome
For example: fleetwood.dev has a really cool project that does audio transcription in browser on the GPU: https://whisper-turbo.com/#
- Run Whisper on WebGPU with a few lines of JS
- Run LLMs on my own Mac fast and efficient Only 2 MBs
-
Distil-Whisper: distilled version of Whisper that is 6 times faster, 49% smaller
You'd be surprised how capable old GPUs are! I've had great success with people running Whisper-Turbo in the browser on really old hardware: https://whisper-turbo.com/
- Running Whisper on Rust and WebGPU
-
Workers AI: serverless GPU-powered inference on Cloudflare’s global network
Whisper large is only 1.5B params, why not run it client side with something like https://github.com/FL33TW00D/whisper-turbo
(Disclaimer: I am the author)
- Whisper Turbo – Run Whisper Directly in the Browser with Rust and WebGPU
- Whisper Turbo: transcribe 20x faster than realtime using Rust and WebGPU
What are some alternatives?
WhisperInput - Offline voice input panel & keyboard with punctuation for Android.
faster-whisper - Faster Whisper transcription with CTranslate2
mfcc-rust
project-2501 - Project 2501 is an open-source AI assistant, written in C++.
whisperX - WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
Clippy - A bunch of lints to catch common mistakes and improve your Rust code. Book: https://doc.rust-lang.org/clippy/
willow - Open source, local, and self-hosted Amazon Echo/Google Home competitive Voice Assistant alternative
whisper-dictation - Dictation app based on the OpenAI speed to text models
discourse-ai
TX-2-simulator - Simulator for the pioneering TX-2 computer