buzz
mimic3
buzz | mimic3 | |
---|---|---|
21 | 24 | |
10,088 | 980 | |
- | 2.7% | |
8.5 | 0.0 | |
7 days ago | 5 months ago | |
Python | Python | |
MIT License | GNU Affero General Public License v3.0 |
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.
buzz
- Buzz: Transcribe and translate audio offline on your personal computer
- MacWhisper: Transcribe audio files on your Mac
-
Build Personal ChatGPT Using Your Data
Easiest 1-click way to install and use Stable Diffusion on your computer."
https://github.com/easydiffusion/easydiffusion
And while Whisper is OpenAI, it is trivial to use locally and extremely usefull
https://github.com/chidiwilliams/buzz
- automated transcription software that is HIPAA compliant?
- Question: Does anyone know of an AI or ChatGPT tool to create automatic SRT caption files by uploading a video?
- Brauchbare Speech-to-Text Lösungen für Windows?
-
I've pretty much had it with Premiere.
Install this for Resolve
-
As a Foreign student, I record letures alot so I can review it anytime. Thanks to Obsidian Audio Player its feels so effortless to review those records.
Just use this one: https://github.com/chidiwilliams/buzz
-
Whispers AI Modular Future
What utilities related to Whisper do you wish existed? What have you had to build yourself?
On the end user application side, I wish there was something that let me pick a podcast of my choosing, get it fully transcribed, and get an embeddings search plus answer q&a on top of that podcast or set of chosen podcasts. I've seen ones for specific podcasts, but I'd like one where I can choose the podcast. (Probably won't build it)
Also on the end user side, I wish there was an Otter alternative (still paid $30/mo, but unlimited minutes per month) that had longer transcription limits. (Started building this, not much interest from users though)
Things I've seen on the dev tool side:
Gladia (API call version of Whisper)
Whisper.cpp
Whisper webservice (https://github.com/ahmetoner/whisper-asr-webservice) - via this thread
Live microphone demo (not real time, it still does it in chunks) https://github.com/mallorbc/whisper_mic
Streamlit UI https://github.com/hayabhay/whisper-ui
Whisper playground https://github.com/saharmor/whisper-playground
Real time whisper https://github.com/shirayu/whispering
Whisper as a service https://github.com/schibsted/WAAS
Improved timestamps and speaker identification https://github.com/m-bain/whisperX
MacWhisper https://goodsnooze.gumroad.com/l/macwhisper
Crossplatform desktop Whisper that supports semi-realtime https://github.com/chidiwilliams/buzz
- Any suggestions for easy ways to add subtitles to YouTube videos?
mimic3
-
ArXiv Papers as Audiobooks
I'd like to take advantage of high quality TTS models but I'd prefer it to be one I may host myself.
Haven't found the right way yet, I'm considering: https://github.com/MycroftAI/mimic3
-
Any recommendation for human like voice AI model for conversation AI?
Fast or good, choose one
Mozilla's TTS is a python package installable with pip and uses cpu or gpu resources to render a choice of voices, they mostly sound natural and this is the good. https://github.com/mozilla/TTS
Mycroft's mimic3 is the default voice renderer for the Mycroft project that runs on pi hardware and sounds ok-ish, that is the fast. https://github.com/MycroftAI/mimic3
There are many others but these are the two I use according to if it needs to run on limited hardware or if the cycles fall freely from the sky.
- I used mimic3 in a few projects. It's relatively lightweight for a neural tts and gives acceptable results
- Mimic 3 Privacy-Focused Neural Text-to-Speech
-
Serverless voice chat with Vicuna-13B
It took quite a bit of digging to find the repo link https://github.com/MycroftAI/mimic3#readme and it's AGPL-3 for those interested in such things
- AI text-to-speech for private, non-commercial use?
-
Ask HN: Cool and Useful Dockerized Apps?
Recently seeing this testing mailserver linked on HN: https://tweedegolf.nl/en/blog/86/introducing-mailcrab
I was reminded of another useful tool, mimic3 by MycroftAI that gives you very nice TTS: https://github.com/MycroftAI/mimic3
So I was wondering: what other useful apps have been containerized for easy setup and great usefulness?
Basically the point of this question is to let people share these, for the benefit of all.
-
Text to speech
You could also try the successor but they didn't get around implementing the harvard voice yet and we don't like any of the voices that come with it.
-
Google tts/ amazon polly alternative?
https://github.com/MycroftAI/mimic3 this? it can be easily hosted on docker as well
- [D] Best TTS for a GPT powered voice assistant
What are some alternatives?
whisper - Robust Speech Recognition via Large-Scale Weak Supervision
piper - A fast, local neural text to speech system
openai-whisper-cpu - Improving transcription performance of OpenAI Whisper for CPU based deployment
TTS - 🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production
StoryToolkitAI - An editing tool that uses AI to transcribe, understand content and search for anything in your footage, integrated with ChatGPT and other AI models
bark - 🔊 Text-Prompted Generative Audio Model
audapolis - an editor for spoken-word audio with automatic transcription
mimic-recording-studio - Mimic Recording Studio is a Docker-based application you can install to record voice samples, which can then be trained into a TTS voice with Mimic2
whisper-diarization - Automatic Speech Recognition with Speaker Diarization based on OpenAI Whisper
tortoise-tts - A multi-voice TTS system trained with an emphasis on quality
text-to-speech-ubuntu - 🙊 Setup "selectable" text to speech / TTS on Ubuntu Linux 24.04 22.04 22.10 23.04 23.10 . Ideal for speed reading, programming, editing and writing.
mimic3-voices - Voice models for Mimic 3 text to speech system