kaldi-active-grammar vs whisper-writer

kaldi-active-grammar

Python Kaldi speech recognition with grammars that can be set active/inactive dynamically at decode-time (by daanzu)

Source Code

Suggest alternative

Edit details

whisper-writer

💬📝 A small dictation app using OpenAI's Whisper speech recognition model. (by savbell)

openai Whisper dictation speech-recognition speech-to-text typing-assistant

Source Code

Suggest alternative

Edit details

Our great sponsors

InfluxDB - Power Real-Time Data Analytics at Scale

WorkOS - The modern identity platform for B2B SaaS

SaaSHub - Software Alternatives and Reviews

Our great sponsors

kaldi-active-grammar		whisper-writer
	Project
10	Mentions	2
329	Stars	184
-	Growth	-
0.0	Activity	6.9
10 months ago	Latest Commit	about 1 month ago
Python	Language	Python
GNU Affero General Public License v3.0	License	GNU General Public License v3.0 only

The number of mentions indicates the total number of mentions that we've tracked plus the number of user suggested alternatives.
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.

kaldi-active-grammar

Posts with mentions or reviews of kaldi-active-grammar. We have used some of these posts to build our list of alternatives and similar projects. The last one was on 2023-11-21.

Ask HN: How do you get started with adding voice commands to a computer system?
2 projects | news.ycombinator.com | 21 Nov 2023

https://github.com/dictation-toolbox/dragonfly
https://github.com/daanzu/kaldi-active-grammar
AMD Screws Gamers: Sponsorships Likely Block DLSS
4 projects | /r/Amd | 4 Jul 2023
Software I’m Thankful For
16 projects | news.ycombinator.com | 23 Sep 2022
Why, in 2022, is there no high quality method for voice control of a PC?
7 projects | news.ycombinator.com | 28 Jan 2022

With an open system/engine, you can train your own personal speech model. For kaldi-active-grammar (https://github.com/daanzu/kaldi-active-grammar), you can do so without all that much difficulty, although the process/documentation could certainly use improvement.
I bootstrapped my personal speech model by retaining the commands from me using WSR. My voice is quite abnormal, and it took only 10 hours of speech data to train a model orders of magnitude more accurate than any generic model I've ever used. And of course, I retain much of my usage now with Kaldi, so my model improves more and more over time. A virtuous flywheel!
Ask HN: Anyone voice code? I had a stroke and can't use my left side
2 projects | news.ycombinator.com | 16 Jan 2022

I have been coding entirely by voice for approximately 10 years now (by hand long before that). Most of that time I have been using the Dragonfly (https://github.com/dictation-toolbox/dragonfly) library to construct my own customized voice coding system. The library is highly flexible and open source, allowing you to easily customize everything to suit what you need to be productive. It is perhaps the power user analogue to Dragon Naturally Speaking. With it, you can certainly be highly productive coding by voice. In fact, I develop kaldi-active-grammar (https://github.com/daanzu/kaldi-active-grammar), a free and open source speech recognition backend usable by Dragonfly, itself entirely by voice. There's also a community of voice coders using Dragonfly and other tools that build on top of it, such as Caster (https://github.com/dictation-toolbox/Caster).
Ask HN: Who Wants to Collaborate?
58 projects | news.ycombinator.com | 1 Jan 2022

- Demo: https://www.youtube.com/watch?v=Qk1mGbIJx3s / Software: https://github.com/daanzu/kaldi-active-grammar
Far field audio is usually harder for any speech system to get correct, so having a good quality mic and using it nearby will _usually_ help with the transcription quality. As a long time Linux user, I would love to see it get some more powerful voice tools - really hope that this opens up over the next few years. Feel free to drop me an email (on my profile) happy to help with setup on any of the above.
How can I make Mycroft recognize non verbal audio sounds to command it?
3 projects | /r/Mycroftai | 29 Jul 2021
Linux Voice recognition/dictation/voice assistant/ one handed operation?
4 projects | /r/linuxquestions | 27 Jul 2021
Disabled computer science student ISO advice about single-handed keyboards
5 projects | /r/ErgoMechKeyboards | 21 Apr 2021

kaldi repo: https://github.com/daanzu/kaldi-active-grammar

whisper-writer

Posts with mentions or reviews of whisper-writer. We have used some of these posts to build our list of alternatives and similar projects. The last one was on 2023-05-08.

Show HN: WhisperWriter – Speech-to-text using OpenAI's Whisper, coded by ChatGPT
2 projects | news.ycombinator.com | 8 May 2023
Using ChatGPT to generate a GPT project end-to-end
4 projects | news.ycombinator.com | 6 May 2023

I've also made six small apps completely coded by ChatGPT (with GitHub Copilot contributing a bit as well). Here are the two largest:
PlaylistGPT (https://github.com/savbell/playlist-gpt): A fun little web app that allows you to ask questions about your Spotify playlists and receive answers from Python code generated by OpenAI's models. I even added a feature where if the code written by GPT runs into errors, it can send the code and the error back to the model and ask it to fix it. It actually can debug itself quite often! One of the most impressive things for me was how it was able to model the UI after the Spotify app with little more than me asking it to do exactly that.
WhisperWriter (https://github.com/savbell/whisper-writer): A small speech-to-text app that uses OpenAI's Whisper API to auto-transcribe recordings from a user's microphone. It waits for a keyboard shortcut to be pressed, then records from the user's microphone until it detects a pause in their speech, and then types out the Whisper transcription to the active window. It only took me two hours to get a working prototype up and running, with additions such as graphic indicators taking a few more hours to implement.
I created the first for fun and the second to help me overcome a disability that impacts my ability to use a keyboard. I now use WhisperWriter literally every day (I'm even typing part of this comment with it), and I used it to prompt ChatGPT to write the code for a few additional personal projects that improve my quality-of-life in small ways. If people are interested, I may write up more about the prompting and pair programming process, since I definitely learned a lot as I worked through these, including some similar lessons to the article!
Personally, I am super excited about the possibilities these AI technologies open up for people like me, who may be facing small challenges that could be easily solved with a tiny app written in a few hours tailored specifically to their problem. I had been struggling to use my desktop computer because the Windows Dictation tool was very broken for me, but now I feel like I can use it to my full capacity again because I can type with WhisperWriter. Coding now takes a minimal amount of keyboard use thanks to these AI coding assistants -- and I am super grateful for that!

What are some alternatives?

When comparing kaldi-active-grammar and whisper-writer you can also consider the following projects:

silero-vad - Silero VAD: pre-trained enterprise-grade Voice Activity Detector

WhisperLive - A nearly-live implementation of OpenAI's Whisper.

nerd-dictation - Simple, hackable offline speech to text - using the VOSK-API.

AI-Waifu-Vtuber - AI Vtuber for Streaming on Youtube/Twitch

pocketsphinx-python - Python interface to CMU Sphinxbase and Pocketsphinx libraries

playlist-gpt - 🎶👩‍💻 A fun little web app that analyzes your Spotify playlists with help from OpenAI's language models.

mycroft-precise - A lightweight, simple-to-use, RNN wake word listener

easy-chat - A ChatGPT UI for young readers, written by ChatGPT

Caster - Dragonfly-Based Voice Programming and Accessibility Toolkit

whisper-openai-gradio-implementation - Whisper is an automatic speech recognition (ASR) system Gradio Web UI Implementation

dragonfly - Speech recognition framework allowing powerful Python-based scripting and extension of Dragon NaturallySpeaking (DNS), Windows Speech Recognition (WSR), Kaldi and CMU Pocket Sphinx

shorthanddictation - Dictation program, which uses the reading speed unit syllables per minute

kaldi-active-grammar vs silero-vad whisper-writer vs WhisperLive kaldi-active-grammar vs nerd-dictation whisper-writer vs AI-Waifu-Vtuber kaldi-active-grammar vs pocketsphinx-python whisper-writer vs playlist-gpt kaldi-active-grammar vs mycroft-precise whisper-writer vs easy-chat kaldi-active-grammar vs Caster whisper-writer vs whisper-openai-gradio-implementation kaldi-active-grammar vs dragonfly whisper-writer vs shorthanddictation

Compare kaldi-active-grammar vs whisper-writer and see what are their differences.

kaldi-active-grammar

whisper-writer

kaldi-active-grammar

whisper-writer

What are some alternatives?