flashlight
TTS
Our great sponsors
flashlight | TTS | |
---|---|---|
16 | 229 | |
5,116 | 28,249 | |
1.0% | 6.9% | |
7.7 | 9.5 | |
9 days ago | 13 days ago | |
C++ | Python | |
MIT License | Mozilla Public License 2.0 |
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.
flashlight
-
MatX: Efficient C++17 GPU numerical computing library with Python-like syntax
I think a comparison to PyTorch, TensorFlow and/or JAX is more relevant than a comparison to CuPy/NumPy.
And then maybe also a comparison to Flashlight (https://github.com/flashlight/flashlight) or other C/C++ based ML/computing libraries?
Also, there is no mention of it, so I suppose this does not support automatic differentiation?
-
Project Resources
This Facebook ai project seems reasonably structured after looking at its CMakeLists.txt. CMake is a build generator for c++, it's how you make binaries to run your project: https://github.com/flashlight/flashlight
-
[D] Deep Learning Framework for C++.
I built and maintain Flashlight, a C++-first library for ML/DL. We built Flashlight to be:
-
What is the most used library for AI in C++ ?
I’ve never used it, but Facebook’s flashlight looks interesting
-
Mozilla Common Voice Adds 16 New Languages and 4,600 New Hours of Speech
I've had good results with https://github.com/flashlight/flashlight/blob/master/flashli.... Seems to work well with spoken english in a variety of accents. Biggest limitation is that the architecture they have pretrained models for doesn't really work well with clips longer than ~15 seconds, so you have to segment your input files.
TTS
-
Base TTS (Amazon): The largest text-to-speech model to-date
I've used coqui.ai's TTS models[0] and library[1] to great success. I was able to get cloned voice to be rendered in about 80% of the audio clip length, and I believe you can also stream the response. Do note the model license for XTTS, it is one they wrote themselves that has some restrictions.
- FLaNK Stack Weekly 12 February 2024
-
Coqui.ai Is Shutting Down
My only exposure to Coqui was their text to speech software. If I remember correctly the website was a commercialized service with TTS and probably some other related things. I hope the software work continues in the open.
-
Hello guys, any selfhosted alternative to eleven labs?
Coqui.ai TTS (https://github.com/coqui-ai/TTS)
-
Demo of Anagnorisis - completely local recommendation system powered by Llama 2. Radio mode. Work in progress.
"tts_models/multilingual/multi-dataset/xtts_v2" model from https://github.com/coqui-ai/TTS. It gives pretty good results and works with references, so it's pretty easy to change the voice. By the way the source code of the project is open: https://github.com/volotat/Anagnorisis but be ready, the code is pretty raw for now.
-
David Attenborough is now narrating my life
I converted a book to audiobook a couple of months ago using the tts code here: https://github.com/coqui-ai/TTS
Wasn't as good as eleven labs, but it had enough options that I found a voice that I liked to listen to.
-
Open Source Libraries
coqui-ai/TTS
-
How I created a text-adventure game using ChatGPT
For a realistic text to speech the program uses TTS by coqui ai. And I think it is amazing.
-
Looking to recreate a cool AI assistant project with free tools
- [Coqui TTS](https://github.com/coqui-ai/TTS) instead of Whisper for text-to-speech
What are some alternatives?
tortoise-tts - A multi-voice TTS system trained with an emphasis on quality
Real-Time-Voice-Cloning - Clone a voice in 5 seconds to generate arbitrary speech in real-time
silero-models - Silero Models: pre-trained speech-to-text, text-to-speech and text-enhancement models made embarrassingly simple
vosk-api - Offline speech recognition API for Android, iOS, Raspberry Pi and servers with Python, Java, C# and Node
text-generation-webui - A Gradio web UI for Large Language Models. Supports transformers, GPTQ, AWQ, EXL2, llama.cpp (GGUF), Llama models.
piper - A fast, local neural text to speech system
bark - 🔊 Text-Prompted Generative Audio Model
PaddleSpeech - Easy-to-use Speech Toolkit including Self-Supervised Learning model, SOTA/Streaming ASR with punctuation, Streaming TTS with text frontend, Speaker Verification System, End-to-End Speech Translation and Keyword Spotting. Won NAACL2022 Best Demo Award.
coqui-docker - Docker images for Coqui AI
gTTS - Python library and CLI tool to interface with Google Translate's text-to-speech API
DeepSpeech - DeepSpeech is an open source embedded (offline, on-device) speech-to-text engine which can run in real time on devices ranging from a Raspberry Pi 4 to high power GPU servers.
speech-and-text-unity-ios-android - Speed to text in Unity iOS use Native Speech Recognition