SpeechRecognition vs aeneas

SpeechRecognition

Speech recognition module for Python, supporting several engines and APIs, online and offline. (by Uberi)

pypi.python.org

aeneas

aeneas is a Python/C library and a set of tools to automagically synchronize audio and text (aka forced alignment) (by readbeyond)

Audio Speech Data Speech Alignment Tts Python Linux MacOS Windows NLP Espeak espeak-ng Festival CLI Dtw Ffmpeg forced-alignment Text Srt Smil text-to-speech

Source Code

readbeyond.it

Docs

Suggest alternative

Edit details

Our great sponsors

WorkOS - The modern identity platform for B2B SaaS

InfluxDB - Power Real-Time Data Analytics at Scale

SaaSHub - Software Alternatives and Reviews

Our great sponsors

SpeechRecognition		aeneas
	Project
16	Mentions	4
8,040	Stars	2,379
-	Growth	-
8.7	Activity	0.0
7 days ago	Latest Commit	over 1 year ago
Python	Language	Python
BSD 3-clause "New" or "Revised" License	License	GNU Affero General Public License v3.0

The number of mentions indicates the total number of mentions that we've tracked plus the number of user suggested alternatives.
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.

SpeechRecognition

Posts with mentions or reviews of SpeechRecognition. We have used some of these posts to build our list of alternatives and similar projects. The last one was on 2023-08-23.

help with script (beginner)
1 project | /r/learnpython | 7 Dec 2023

Start and Stop Listening Example
MacWhisper: Transcribe audio files on your Mac
8 projects | news.ycombinator.com | 23 Aug 2023

There is a great library that has support not only with OpenAIs whisper but many others that also work offline. https://github.com/Uberi/speech_recognition
Unpopular Opinion: a lot of Obsidian community make Obsidian sound like something cringey/productivity guru-y
1 project | /r/ObsidianMD | 14 May 2023

This is the library: https://github.com/Uberi/speech_recognition
Nvim-VoiceRec : Add Speech-To-Text To Neovim! (useful for gpt)
4 projects | /r/neovim | 28 Apr 2023

It is python remote plugin that is a tin wrapper around speech_recognition package.
Speech-to-text software
1 project | /r/opensource | 15 Feb 2023
Voice commands in Doom Eternal possible?
1 project | /r/linux_gaming | 23 Dec 2022

I am less familiar with speech recognition myself. I have implemented something similar many years ago, back when Google had a REST API that allowed you to upload audio and they would respond with the recognized words/sentence. I think they still have the same API available, though. They limited how much you could send, but for voice commands it was pretty solid. However, SpeechRecognition looks like a library worth trying out for this, as that seems like it could do offline processing depending on the underlying library. They also have some examples to look at.
Build Simple CLI-Based Voice Assistant with PyAudio, Speech Recognition, pyttsx3 and SerpApi
7 projects | dev.to | 28 Nov 2022

SpeechRecognition
Need help with speech recognition
1 project | /r/learnpython | 4 Jul 2022
Wiki for the podcast
1 project | /r/Cortex | 3 Apr 2022

I found this one here
How to use my speaker as input and my mic as output?
1 project | /r/Python | 1 Jan 2022

https://github.com/Uberi/speech_recognition/blob/master/reference/library-reference.rst this might help. I guess your best bet is to rtfm.

aeneas

Posts with mentions or reviews of aeneas. We have used some of these posts to build our list of alternatives and similar projects. The last one was on 2023-02-03.

Anyone know of a tool to align (existing) subtitles to audio along sentence boundaries?
4 projects | /r/LanguageTechnology | 3 Feb 2023

You could try aeneas. Syncabook apparently uses the afaligner library, which says that it was inspired by aeneas but uses FastDTW to find an approximation to the optimal warping path. This might make it slightly less accurate than aeneas.
WhisperSync alternative for Plex Audiobooks and already owned E-Books
1 project | /r/selfhosted | 12 Sep 2021

Check out https://github.com/readbeyond/aeneas
Speech Recognition Training Data Tools?
2 projects | /r/LanguageTechnology | 27 Apr 2021

In case you have let's say: a 20min entry from an audio book, and the sentences seperatly in a txt file and you want to cut the sentences out of the audio manually you can look at a tool like aeneas. If you still have to annotated all your data yourself i do not really know a tool for this :/
Show HN: A retrainable subtitle synchronizer you can now build your own
3 projects | news.ycombinator.com | 31 Jan 2021

here's another solution: https://github.com/readbeyond/aeneas

What are some alternatives?

When comparing SpeechRecognition and aeneas you can also consider the following projects:

pydub - Manipulate audio with a simple and easy high level interface

Prosodylab-Aligner - Python interface for forced audio alignment using HTK and SoX

pyAudioAnalysis - Python Audio Analysis Library: Feature Extraction, Classification, Segmentation and Applications

allosaurus - Allosaurus is a pretrained universal phone recognizer for more than 2000 languages

espeak-ng - eSpeak NG is an open source speech synthesizer that supports more than hundred languages and accents.

speech-to-text-websockets-python

Watson Developer Cloud Python SDK - :snake: Client library to use the IBM Watson services in Python and available in pip as watson-developer-cloud

speechpy - :speech_balloon: SpeechPy - A Library for Speech Processing and Recognition: http://speechpy.readthedocs.io/en/latest/

audioread - cross-library (GStreamer + Core Audio + MAD + FFmpeg) audio decoding for Python

SpeechRecognition vs pydub aeneas vs Prosodylab-Aligner SpeechRecognition vs pyAudioAnalysis aeneas vs pyAudioAnalysis SpeechRecognition vs allosaurus aeneas vs espeak-ng SpeechRecognition vs speech-to-text-websockets-python aeneas vs Watson Developer Cloud Python SDK SpeechRecognition vs speechpy aeneas vs speechpy SpeechRecognition vs Watson Developer Cloud Python SDK aeneas vs audioread

Compare SpeechRecognition vs aeneas and see what are their differences.

SpeechRecognition

aeneas

SpeechRecognition

aeneas

What are some alternatives?