SpeechRecognition vs buzz

SpeechRecognition

Speech recognition module for Python, supporting several engines and APIs, online and offline. (by Uberi)

pypi.python.org

buzz

Buzz transcribes and translates audio offline on your personal computer. Powered by OpenAI's Whisper. (by chidiwilliams)

Whisper

Source Code

chidiwilliams.github.io

Suggest alternative

Edit details

Our great sponsors

WorkOS - The modern identity platform for B2B SaaS

InfluxDB - Power Real-Time Data Analytics at Scale

SaaSHub - Software Alternatives and Reviews

Our great sponsors

SpeechRecognition		buzz
	Project
16	Mentions	21
8,040	Stars	9,869
-	Growth	-
8.7	Activity	8.5
11 days ago	Latest Commit	18 days ago
Python	Language	Python
BSD 3-clause "New" or "Revised" License	License	MIT License

The number of mentions indicates the total number of mentions that we've tracked plus the number of user suggested alternatives.
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.

SpeechRecognition

Posts with mentions or reviews of SpeechRecognition. We have used some of these posts to build our list of alternatives and similar projects. The last one was on 2023-08-23.

help with script (beginner)
1 project | /r/learnpython | 7 Dec 2023

Start and Stop Listening Example
MacWhisper: Transcribe audio files on your Mac
8 projects | news.ycombinator.com | 23 Aug 2023

There is a great library that has support not only with OpenAIs whisper but many others that also work offline. https://github.com/Uberi/speech_recognition
Unpopular Opinion: a lot of Obsidian community make Obsidian sound like something cringey/productivity guru-y
1 project | /r/ObsidianMD | 14 May 2023

This is the library: https://github.com/Uberi/speech_recognition
Nvim-VoiceRec : Add Speech-To-Text To Neovim! (useful for gpt)
4 projects | /r/neovim | 28 Apr 2023

It is python remote plugin that is a tin wrapper around speech_recognition package.
Speech-to-text software
1 project | /r/opensource | 15 Feb 2023
Voice commands in Doom Eternal possible?
1 project | /r/linux_gaming | 23 Dec 2022

I am less familiar with speech recognition myself. I have implemented something similar many years ago, back when Google had a REST API that allowed you to upload audio and they would respond with the recognized words/sentence. I think they still have the same API available, though. They limited how much you could send, but for voice commands it was pretty solid. However, SpeechRecognition looks like a library worth trying out for this, as that seems like it could do offline processing depending on the underlying library. They also have some examples to look at.
Build Simple CLI-Based Voice Assistant with PyAudio, Speech Recognition, pyttsx3 and SerpApi
7 projects | dev.to | 28 Nov 2022

SpeechRecognition
Need help with speech recognition
1 project | /r/learnpython | 4 Jul 2022
Wiki for the podcast
1 project | /r/Cortex | 3 Apr 2022

I found this one here
How to use my speaker as input and my mic as output?
1 project | /r/Python | 1 Jan 2022

https://github.com/Uberi/speech_recognition/blob/master/reference/library-reference.rst this might help. I guess your best bet is to rtfm.

buzz

Posts with mentions or reviews of buzz. We have used some of these posts to build our list of alternatives and similar projects. The last one was on 2023-08-23.

Buzz: Transcribe and translate audio offline on your personal computer
1 project | news.ycombinator.com | 21 Mar 2024
MacWhisper: Transcribe audio files on your Mac
8 projects | news.ycombinator.com | 23 Aug 2023
Build Personal ChatGPT Using Your Data
14 projects | news.ycombinator.com | 8 Jul 2023

Easiest 1-click way to install and use Stable Diffusion on your computer."
https://github.com/easydiffusion/easydiffusion
And while Whisper is OpenAI, it is trivial to use locally and extremely usefull
https://github.com/chidiwilliams/buzz
automated transcription software that is HIPAA compliant?
1 project | /r/AskAcademia | 5 May 2023
Question: Does anyone know of an AI or ChatGPT tool to create automatic SRT caption files by uploading a video?
1 project | /r/ChatGPTPro | 21 Apr 2023
Brauchbare Speech-to-Text Lösungen für Windows?
1 project | /r/de_EDV | 25 Mar 2023
I've pretty much had it with Premiere.
3 projects | /r/editors | 10 Mar 2023

Install this for Resolve
As a Foreign student, I record letures alot so I can review it anytime. Thanks to Obsidian Audio Player its feels so effortless to review those records.
4 projects | /r/ObsidianMD | 24 Feb 2023

Just use this one: https://github.com/chidiwilliams/buzz
Whispers AI Modular Future
14 projects | news.ycombinator.com | 20 Feb 2023

What utilities related to Whisper do you wish existed? What have you had to build yourself?
On the end user application side, I wish there was something that let me pick a podcast of my choosing, get it fully transcribed, and get an embeddings search plus answer q&a on top of that podcast or set of chosen podcasts. I've seen ones for specific podcasts, but I'd like one where I can choose the podcast. (Probably won't build it)
Also on the end user side, I wish there was an Otter alternative (still paid $30/mo, but unlimited minutes per month) that had longer transcription limits. (Started building this, not much interest from users though)
Things I've seen on the dev tool side:
Gladia (API call version of Whisper)
Whisper.cpp
Whisper webservice (https://github.com/ahmetoner/whisper-asr-webservice) - via this thread
Live microphone demo (not real time, it still does it in chunks) https://github.com/mallorbc/whisper_mic
Streamlit UI https://github.com/hayabhay/whisper-ui
Whisper playground https://github.com/saharmor/whisper-playground
Real time whisper https://github.com/shirayu/whispering
Whisper as a service https://github.com/schibsted/WAAS
Improved timestamps and speaker identification https://github.com/m-bain/whisperX
MacWhisper https://goodsnooze.gumroad.com/l/macwhisper
Crossplatform desktop Whisper that supports semi-realtime https://github.com/chidiwilliams/buzz
Any suggestions for easy ways to add subtitles to YouTube videos?
1 project | /r/VideoEditing | 19 Feb 2023

What are some alternatives?

When comparing SpeechRecognition and buzz you can also consider the following projects:

pydub - Manipulate audio with a simple and easy high level interface

whisper - Robust Speech Recognition via Large-Scale Weak Supervision

pyAudioAnalysis - Python Audio Analysis Library: Feature Extraction, Classification, Segmentation and Applications

openai-whisper-cpu - Improving transcription performance of OpenAI Whisper for CPU based deployment

allosaurus - Allosaurus is a pretrained universal phone recognizer for more than 2000 languages

StoryToolkitAI - An editing tool that uses AI to transcribe, understand content and search for anything in your footage, integrated with ChatGPT and other AI models

aeneas - aeneas is a Python/C library and a set of tools to automagically synchronize audio and text (aka forced alignment)

audapolis - an editor for spoken-word audio with automatic transcription

speech-to-text-websockets-python

whisper-diarization - Automatic Speech Recognition with Speaker Diarization based on OpenAI Whisper

speechpy - :speech_balloon: SpeechPy - A Library for Speech Processing and Recognition: http://speechpy.readthedocs.io/en/latest/

text-to-speech-ubuntu - 🙊 Setup "selectable" text to speech / TTS on Ubuntu Linux 24.04 22.04 22.10 23.04 23.10 . Ideal for speed reading, programming, editing and writing.