frogbase
transcribe-anything
frogbase | transcribe-anything | |
---|---|---|
14 | 11 | |
754 | 355 | |
- | - | |
4.3 | 9.3 | |
7 months ago | 12 days ago | |
Python | Python | |
MIT License | MIT License |
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.
frogbase
-
For people who tried whisper
If you’re looking to use a local deployment and are comfortable with python projects, this deployment was fairly easy to use: https://github.com/hayabhay/whisper-ui
-
I have a two-step process for taking notes on transcripts that I'd like to share, but am also looking for feedback for the final step
I recommend using chat GPT to learn some very basic python/programming tho. I like this one https://github.com/hayabhay/whisper-ui which uses streamlit for easy to use UI for bulk transcriptions.
-
(Preferably) Self Hosted Podcasts with searchable transcripts
This tool 2 out of the 4 items you mentioned: https://github.com/hayabhay/whisper-ui
-
Whispers AI Modular Future
What utilities related to Whisper do you wish existed? What have you had to build yourself?
On the end user application side, I wish there was something that let me pick a podcast of my choosing, get it fully transcribed, and get an embeddings search plus answer q&a on top of that podcast or set of chosen podcasts. I've seen ones for specific podcasts, but I'd like one where I can choose the podcast. (Probably won't build it)
Also on the end user side, I wish there was an Otter alternative (still paid $30/mo, but unlimited minutes per month) that had longer transcription limits. (Started building this, not much interest from users though)
Things I've seen on the dev tool side:
Gladia (API call version of Whisper)
Whisper.cpp
Whisper webservice (https://github.com/ahmetoner/whisper-asr-webservice) - via this thread
Live microphone demo (not real time, it still does it in chunks) https://github.com/mallorbc/whisper_mic
Streamlit UI https://github.com/hayabhay/whisper-ui
Whisper playground https://github.com/saharmor/whisper-playground
Real time whisper https://github.com/shirayu/whispering
Whisper as a service https://github.com/schibsted/WAAS
Improved timestamps and speaker identification https://github.com/m-bain/whisperX
MacWhisper https://goodsnooze.gumroad.com/l/macwhisper
Crossplatform desktop Whisper that supports semi-realtime https://github.com/chidiwilliams/buzz
-
[P] Whisper-UI Update: You can now bulk-transcribe, save & search transcriptions with Streamlit & SQLAlchemy 2.0 [details in the comments]
Github Repo: https://github.com/hayabhay/whisper-ui
-
Self-host Whisper As a Service with GUI and queueing. Schibsted created a transcription service for our journalists to transcribe audio interviews and podcasts really quick.
People may also like this tool which is a bit more about searching the contents. https://github.com/hayabhay/whisper-ui
- Show HN: Self-host Whisper As a Service with GUI and queueing
-
Audio equivalent of paperless?
There is whisper ui to create meta data (speech 2 text).
- Whisper-UI: You can now bulk-transcribe, save & search transcriptions from YouTube with OpenAI's Whisper, Streamlit & SQLAlchemy 2.0
- Whisper-UI Update: You can now bulk-transcribe, save & search transcriptions with Streamlit & SQLAlchemy 2.0
transcribe-anything
-
Summarize audio recordings in text
transcribe-anything
-
$620,000 stolen from YouTuber Ethan Klein and the H3 Podcast by MCN BroadbandTV and their CEO Shahrzad Rafati
OpenAI whisper. Here is a tool that has it, a video downloader, and some other things bundled in with it: https://github.com/zackees/transcribe-anything
- 32 Open Source Libraries for Python's 32nd Birthday
-
Show HN: Self-host Whisper As a Service with GUI and queueing
People interested in this might also be interested in transcribe-anything [1].
It automates video fetching and uses whisper to generate .srt, .vtt and .txt files.
[1] https://github.com/zackees/transcribe-anything
-
[P] Free Youtube Subtitles Generator
Nice looks great, link broken but it just needed a hyphen https://github.com/zackees/transcribe-anything
-
Gpu accelerated ML apps will soon get a lot easier to deploy - Pytorch-cuda moving to 100% pypi hosting.
Right now the cuda accelerated whls are hosted outside of pypi which can only be accessed by using `--extra-index-url`, when installing from a requirements file (pip install -r requirements.txt). However pip install doesn't allow --extra-index-url for security reasons, which means deploying cuda accelerated ML apps on python is a complicated affair, see this [script](https://github.com/zackees/transcribe-anything/blob/main/install_cuda.py) as an example of what needs to be done to uninstall conflicting cpu only version of pytorch and replace it with cuda acceleration.
- Bro, listen: Interact with OpenAI using voice
- Convert YouTube to Text with OpenAI Whisper
-
Draw an owl
Transcribe Anything
-
Transcribe Video/Audio on the web using `transcribe-anything`, a front end to WhisperAI
Code Repo: https://github.com/zackees/transcribe-anything (please give my repo a like)
What are some alternatives?
whisper.cpp - Port of OpenAI's Whisper model in C/C++
Hentai-Diffusion - The official place for the best A.I.
whisper - Robust Speech Recognition via Large-Scale Weak Supervision
whisperX - WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
FlexGen - Running large language models on a single GPU for throughput-oriented scenarios.
subtitle-generator - Generate subtitles for youtube videos for free with https://text-generator.io
whisper-playground - Build real time speech2text web apps using OpenAI's Whisper https://openai.com/blog/whisper/
static_ffmpeg - Installs FFMPEG v5 On Win32/Ubuntu/MacOS
nlp
ai-notes - notes for software engineers getting up to speed on new AI developments. Serves as datastore for https://latent.space writing, and product brainstorming, but has cleaned up canonical references under the /Resources folder.
audio-files-addon - Audio file support for Docspell
DeepSpeech-examples - Examples of how to use or integrate DeepSpeech