STT
STT-examples
Our great sponsors
STT | STT-examples | |
---|---|---|
11 | 5 | |
2,130 | 111 | |
2.7% | 2.7% | |
0.6 | 0.0 | |
about 2 months ago | over 1 year ago | |
C++ | Python | |
Mozilla Public License 2.0 | Mozilla Public License 2.0 |
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.
STT
-
Rest in Peas: The Unrecognized Death of Speech Recognition (2010)
What has happened since then? I know Common Voice has come and gone https://en.wikipedia.org/wiki/Common_Voice https://github.com/coqui-ai/STT
And I've seen some neural approaches too
No idea where to look for comparisons though.
-
Numen - FOSS voice control for handsfree computing
I basically just used coqui stt https://github.com/coqui-ai/STT
-
Are there any OCR and Speech-to-Text services that are privacy friendly?
This speech-to-text works well: https://github.com/coqui-ai/STT. openai's "whisper" is probably better but I haven't tried it: https://towardsdatascience.com/transcribe-audio-files-with-openais-whisper-e973ae348aa7
-
Introducing Whisper
I use two SST to live-translate audio that I listen to so I can look back (in paragraph form) to see things that I or the youtube has previously said: https://github.com/coqui-ai/STT https://github.com/ratwithacompiler/OBS-captions-plugin
-
You can now tether any prod Vector to Wire's Open Source Escape Pod • thedroidyouarelookingfor
I did have to install Coqui STT and go-asticoqui manually before i was able to run Chipper.
-
Currently working on a custom Virtual Assistant ('Randy') to help automate things in my shed (mainly CNC equipment) and also perform basic tasks. This morning I was able to get it to publish events on my google calendar.
What do you use as STT? I have heard good things about coqui (https://github.com/coqui-ai/STT) and will use it for my Assistant-build.
- Speech to Text Best Resource
-
I put together a tutorial and overview on how to use DeepSpeech to do Speech Recognition in Python
If anyone is looking for a maintained version of DeepSpeech, checkout Coqui's repositories for STT and TTS. Coqui is lead by the engineers that used to work on DeepSpeech at Mozilla.
-
CoquiTTS: 🐸💬 - Open Source Text-to-Speech framework.
Link: https://github.com/coqui-ai/STT
- Mozilla Common Voice Adds 16 New Languages and 4,600 New Hours of Speech
STT-examples
-
Web Speech API is not available in the Quest browser
You're welcome! Your post actually got me a little interested in seeing what's new with Coqui STT (where all the old Mozilla STT folks moved to) and it seems someone was working on a WebAssembly binding for it, so one could probably finagle something themselves for testing purposes (the bandwidth of loading the model for every user on every page load is unfeasible from a production cost standpoint though)
-
DeepSpeech 60x Smaller, 9x faster, and 2x accuracy
I will add https://github.com/coqui-ai/STT, which is a continuation of DeepSpeech. Also, I've been messing around with https://github.com/ideasman42/nerd-dictation, which works on a VOSK backend - accuracy is decent, especially with the bigger model.
-
Any privacy friendly automated transcript app?
I don't know of a complete app, but https://github.com/coqui-ai/STT, which grew out of the now-unmaintained Mozilla Deepspeech project, works well and is easy to use. It could be a good starting point if you're comfortable writing a little code.
-
[N] 🐸Coqui and OVHCloud are organizing an open-source Speech Recognition Hackaton
👉CoquiSTT - https://github.com/coqui-ai/STT
-
Coqui, a startup providing open speech tech for everyone
https://github.com/coqui-ai/STT-examples
If you have any more specific requirements then we can point you in the right direction. Or just join us on Matrix: https://app.element.io/#/room/#coqui-ai_STT:gitter.im :)
What are some alternatives?
DeepSpeech - DeepSpeech is an open source embedded (offline, on-device) speech-to-text engine which can run in real time on devices ranging from a Raspberry Pi 4 to high power GPU servers.
vosk-api - Offline speech recognition API for Android, iOS, Raspberry Pi and servers with Python, Java, C# and Node
TTS - 🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production
LocalSTT - Android Speech Recognition Service using Vosk/Kaldi and Mozilla DeepSpeech
NeMo - A scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI (Automatic Speech Recognition and Text-to-Speech)
nerd-dictation - Simple, hackable offline speech to text - using the VOSK-API.
speech-to-text-benchmark - speech to text benchmark framework
TTS - :robot: :speech_balloon: Deep learning for Text to Speech (Discussion forum: https://discourse.mozilla.org/c/tts)
leopard - On-device speech-to-text engine powered by deep learning
OBS-captions-plugin - Closed Captioning OBS plugin using Google Speech Recognition