kaldi-active-grammar vs nerd-dictation

kaldi-active-grammar

Python Kaldi speech recognition with grammars that can be set active/inactive dynamically at decode-time (by daanzu)

Source Code

Suggest alternative

Edit details

nerd-dictation

Simple, hackable offline speech to text - using the VOSK-API. (by ideasman42)

Suggest topics

Source Code

Suggest alternative

Edit details

Our great sponsors

InfluxDB - Power Real-Time Data Analytics at Scale

WorkOS - The modern identity platform for B2B SaaS

SaaSHub - Software Alternatives and Reviews

Our great sponsors

kaldi-active-grammar		nerd-dictation
	Project
10	Mentions	28
329	Stars	1,158
-	Growth	-
0.0	Activity	3.6
10 months ago	Latest Commit	30 days ago
Python	Language	Python
GNU Affero General Public License v3.0	License	GNU General Public License v3.0 only

The number of mentions indicates the total number of mentions that we've tracked plus the number of user suggested alternatives.
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.

kaldi-active-grammar

Posts with mentions or reviews of kaldi-active-grammar. We have used some of these posts to build our list of alternatives and similar projects. The last one was on 2023-11-21.

Ask HN: How do you get started with adding voice commands to a computer system?
2 projects | news.ycombinator.com | 21 Nov 2023

https://github.com/dictation-toolbox/dragonfly
https://github.com/daanzu/kaldi-active-grammar
AMD Screws Gamers: Sponsorships Likely Block DLSS
4 projects | /r/Amd | 4 Jul 2023
Software I’m Thankful For
16 projects | news.ycombinator.com | 23 Sep 2022
Why, in 2022, is there no high quality method for voice control of a PC?
7 projects | news.ycombinator.com | 28 Jan 2022

With an open system/engine, you can train your own personal speech model. For kaldi-active-grammar (https://github.com/daanzu/kaldi-active-grammar), you can do so without all that much difficulty, although the process/documentation could certainly use improvement.
I bootstrapped my personal speech model by retaining the commands from me using WSR. My voice is quite abnormal, and it took only 10 hours of speech data to train a model orders of magnitude more accurate than any generic model I've ever used. And of course, I retain much of my usage now with Kaldi, so my model improves more and more over time. A virtuous flywheel!
Ask HN: Anyone voice code? I had a stroke and can't use my left side
2 projects | news.ycombinator.com | 16 Jan 2022

I have been coding entirely by voice for approximately 10 years now (by hand long before that). Most of that time I have been using the Dragonfly (https://github.com/dictation-toolbox/dragonfly) library to construct my own customized voice coding system. The library is highly flexible and open source, allowing you to easily customize everything to suit what you need to be productive. It is perhaps the power user analogue to Dragon Naturally Speaking. With it, you can certainly be highly productive coding by voice. In fact, I develop kaldi-active-grammar (https://github.com/daanzu/kaldi-active-grammar), a free and open source speech recognition backend usable by Dragonfly, itself entirely by voice. There's also a community of voice coders using Dragonfly and other tools that build on top of it, such as Caster (https://github.com/dictation-toolbox/Caster).
Ask HN: Who Wants to Collaborate?
58 projects | news.ycombinator.com | 1 Jan 2022

- Demo: https://www.youtube.com/watch?v=Qk1mGbIJx3s / Software: https://github.com/daanzu/kaldi-active-grammar
Far field audio is usually harder for any speech system to get correct, so having a good quality mic and using it nearby will _usually_ help with the transcription quality. As a long time Linux user, I would love to see it get some more powerful voice tools - really hope that this opens up over the next few years. Feel free to drop me an email (on my profile) happy to help with setup on any of the above.
How can I make Mycroft recognize non verbal audio sounds to command it?
3 projects | /r/Mycroftai | 29 Jul 2021
Linux Voice recognition/dictation/voice assistant/ one handed operation?
4 projects | /r/linuxquestions | 27 Jul 2021
Disabled computer science student ISO advice about single-handed keyboards
5 projects | /r/ErgoMechKeyboards | 21 Apr 2021

kaldi repo: https://github.com/daanzu/kaldi-active-grammar

nerd-dictation

Posts with mentions or reviews of nerd-dictation. We have used some of these posts to build our list of alternatives and similar projects. The last one was on 2023-07-07.

why nerd-dictation support in NixOS is stuck ?
2 projects | /r/NixOS | 7 Jul 2023
Is anyone doing always-on voice to text with a local llama at home?
5 projects | /r/LocalLLaMA | 25 Jun 2023
Apollo dev posts backend code to Git to disprove Reddit’s claims of scrapping and inefficiency
4 projects | /r/webdev | 9 Jun 2023

nerd-dictation
How to use notion in gnome
1 project | /r/gnome | 1 May 2023

There's no built-in way of doing this in GNOME, but you might already get a bit further with tools like https://github.com/ideasman42/nerd-dictation
What voice transcriber do you use?
1 project | /r/foss | 20 Apr 2023
Disability accessibility tools for Linux such as eyetrackers and voice commands?
2 projects | /r/linux | 19 Feb 2023

I'm not familiar with Talon so I don't know if this is a suitable suggestion but nerd-dictation seemed to have been well received here when it was last promoted and it looks like it's still in active development.
Voice Control was supposed to be the Future. Is Linux lagging behind?
4 projects | /r/linux | 3 Dec 2022

TBF Microsoft dropped IE, windows phone... that is not uncommon. But the OP is right, maybe not much for voice control but for dictation certainly. The FLOSS community is always far behind and thus always struggle with new technologies. We should be prepared. Since you've mentioned small open source project here's a demo of NerdDitaction. FYI Linux do have mobile devices developing.
I've made voice input for Linux that I use instead of a keyboard and mouse
1 project | /r/RSI | 6 Nov 2022

Yeah you get me. I did have RSI which was amplified by my other issue, but it was that issue that progressed and why can't type now, not RSI. I'd be interested in hearing about using numen in combination with typing, but it's likely not ideal yet. Maybe just using speech to text for some things could help? It's not my project but there's: https://github.com/ideasman42/nerd-dictation that uses the same speech recognition as numen.
Voice to text for Linux
1 project | /r/privacy | 1 Nov 2022
nerd-dictation: Simple, hackable offline speech to text - using the VOSK-API.
1 project | /r/planetemacs | 28 Sep 2022

What are some alternatives?

When comparing kaldi-active-grammar and nerd-dictation you can also consider the following projects:

silero-vad - Silero VAD: pre-trained enterprise-grade Voice Activity Detector

vosk-api - Offline speech recognition API for Android, iOS, Raspberry Pi and servers with Python, Java, C# and Node

pocketsphinx-python - Python interface to CMU Sphinxbase and Pocketsphinx libraries

recasepunc - Model for recasing and repunctuating ASR transcripts

mycroft-precise - A lightweight, simple-to-use, RNN wake word listener

cursorless - Don't let the cursor slow you down

Caster - Dragonfly-Based Voice Programming and Accessibility Toolkit

tortoise-tts - A multi-voice TTS system trained with an emphasis on quality

dragonfly - Speech recognition framework allowing powerful Python-based scripting and extension of Dragon NaturallySpeaking (DNS), Windows Speech Recognition (WSR), Kaldi and CMU Pocket Sphinx

monkeytype - The most customizable typing website with a minimalistic design and a ton of features. Test yourself in various modes, track your progress and improve your speed.

Common-Voice - Audio Classification with machine learning

vosk-android-demo - Offline speech recognition IME for Android with Vosk library.

kaldi-active-grammar vs silero-vad nerd-dictation vs vosk-api kaldi-active-grammar vs pocketsphinx-python nerd-dictation vs recasepunc kaldi-active-grammar vs mycroft-precise nerd-dictation vs cursorless kaldi-active-grammar vs Caster nerd-dictation vs tortoise-tts kaldi-active-grammar vs dragonfly nerd-dictation vs monkeytype kaldi-active-grammar vs Common-Voice nerd-dictation vs vosk-android-demo

Compare kaldi-active-grammar vs nerd-dictation and see what are their differences.

kaldi-active-grammar

nerd-dictation

kaldi-active-grammar

nerd-dictation

What are some alternatives?