frogbase
ai-notes
frogbase | ai-notes | |
---|---|---|
14 | 15 | |
754 | 4,669 | |
- | - | |
4.3 | 9.8 | |
7 months ago | 2 days ago | |
Python | HTML | |
MIT License | MIT License |
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.
frogbase
-
For people who tried whisper
If you’re looking to use a local deployment and are comfortable with python projects, this deployment was fairly easy to use: https://github.com/hayabhay/whisper-ui
-
I have a two-step process for taking notes on transcripts that I'd like to share, but am also looking for feedback for the final step
I recommend using chat GPT to learn some very basic python/programming tho. I like this one https://github.com/hayabhay/whisper-ui which uses streamlit for easy to use UI for bulk transcriptions.
-
(Preferably) Self Hosted Podcasts with searchable transcripts
This tool 2 out of the 4 items you mentioned: https://github.com/hayabhay/whisper-ui
-
Whispers AI Modular Future
What utilities related to Whisper do you wish existed? What have you had to build yourself?
On the end user application side, I wish there was something that let me pick a podcast of my choosing, get it fully transcribed, and get an embeddings search plus answer q&a on top of that podcast or set of chosen podcasts. I've seen ones for specific podcasts, but I'd like one where I can choose the podcast. (Probably won't build it)
Also on the end user side, I wish there was an Otter alternative (still paid $30/mo, but unlimited minutes per month) that had longer transcription limits. (Started building this, not much interest from users though)
Things I've seen on the dev tool side:
Gladia (API call version of Whisper)
Whisper.cpp
Whisper webservice (https://github.com/ahmetoner/whisper-asr-webservice) - via this thread
Live microphone demo (not real time, it still does it in chunks) https://github.com/mallorbc/whisper_mic
Streamlit UI https://github.com/hayabhay/whisper-ui
Whisper playground https://github.com/saharmor/whisper-playground
Real time whisper https://github.com/shirayu/whispering
Whisper as a service https://github.com/schibsted/WAAS
Improved timestamps and speaker identification https://github.com/m-bain/whisperX
MacWhisper https://goodsnooze.gumroad.com/l/macwhisper
Crossplatform desktop Whisper that supports semi-realtime https://github.com/chidiwilliams/buzz
-
[P] Whisper-UI Update: You can now bulk-transcribe, save & search transcriptions with Streamlit & SQLAlchemy 2.0 [details in the comments]
Github Repo: https://github.com/hayabhay/whisper-ui
-
Self-host Whisper As a Service with GUI and queueing. Schibsted created a transcription service for our journalists to transcribe audio interviews and podcasts really quick.
People may also like this tool which is a bit more about searching the contents. https://github.com/hayabhay/whisper-ui
- Show HN: Self-host Whisper As a Service with GUI and queueing
-
Audio equivalent of paperless?
There is whisper ui to create meta data (speech 2 text).
- Whisper-UI: You can now bulk-transcribe, save & search transcriptions from YouTube with OpenAI's Whisper, Streamlit & SQLAlchemy 2.0
- Whisper-UI Update: You can now bulk-transcribe, save & search transcriptions with Streamlit & SQLAlchemy 2.0
ai-notes
-
Minimal implementation of Mamba, the new LLM architecture, in 1 file of PyTorch
the field just moves fast. I have curated a list of non-hypey writers and youtubers who explain these things for a typical SWE audience if you are interested. https://github.com/swyxio/ai-notes/blob/main/Resources/Good%...
- SDXL Turbo: A Real-Time Text-to-Image Generation Model
-
DeepEval – Unit Testing for LLMs
added to my notes! https://github.com/swyxio/ai-notes/
- ChatGPT Code Interpreter Capabilities
-
Google just released a 100% free learning path on Generative AI with 9 Courses
and here are mine, organized by beginner/intermediate/advanced
https://github.com/swyxio/ai-notes/blob/main/README.md#top-a...
and then you can go into the individual modality specific notes for more reading
- Show HN: Self-host Whisper As a Service with GUI and queueing
-
Show HN: YouTube Summaries Using GPT
there's https://learnprompting.org/
i've also been keeping a popular series of notes https://github.com/sw-yx/ai-notes/blob/main/TEXT_PROMPTS.md
-
Show HN: I reverse prompt engineered every Notion AI feature
Direct link to the source prompts are here: https://github.com/sw-yx/ai-notes/blob/main/Resources/Notion...
- GitHub - sw-yx/prompt-eng: notes for prompt engineering
- My hand-curated list of major distros and forks of Stable Diffusion. Please suggest anything I missed!
What are some alternatives?
whisper.cpp - Port of OpenAI's Whisper model in C/C++
text2image-gui - Somewhat modular text2image GUI, initially just for Stable Diffusion
whisper - Robust Speech Recognition via Large-Scale Weak Supervision
diffusionbee-stable-diffusion-ui - Diffusion Bee is the easiest way to run Stable Diffusion locally on your M1 Mac. Comes with a one-click installer. No dependencies or technical knowledge needed.
transcribe-anything - Input a local file or url and this service will transcribe it using Whisper AI. Completely private and Free 🤯🤯🤯
m1_huggingface_diffusers_demo - Demo of how to get HuggingFace Diffusers working on an M1 Mac
FlexGen - Running large language models on a single GPU for throughput-oriented scenarios.
stable-diffusion-ui - Easiest 1-click way to install and use Stable Diffusion on your computer. Provides a browser UI for generating images from text prompts and images. Just enter your text prompt, and see the generated image. [Moved to: https://github.com/easydiffusion/easydiffusion]
whisper-playground - Build real time speech2text web apps using OpenAI's Whisper https://openai.com/blog/whisper/
perceiver-pytorch - Implementation of Perceiver, General Perception with Iterative Attention, in Pytorch
nlp
stable-diffusion - A latent text-to-image diffusion model