instructor-embedding
OpenLLM
instructor-embedding | OpenLLM | |
---|---|---|
4 | 25 | |
1,703 | 8,813 | |
3.1% | 3.6% | |
5.9 | 9.9 | |
10 days ago | 4 days ago | |
Python | Python | |
Apache License 2.0 | Apache License 2.0 |
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.
instructor-embedding
-
My experience on starting with fine tuning LLMs with custom data
If you li embeddings and vector DB, you should look into this: https://github.com/HKUNLP/instructor-embedding
-
Build Personal ChatGPT Using Your Data
If you look at a embeddings leaderboard [1], one of the top competitors called InstructorXL [2] is just a pip install away. It's neck and neck with Ada v2 except for a shorter input length and half the dimensions, with the added benefit that you'll always have the model available.
Most of the other options just work with the transformers library.
[1] https://huggingface.co/spaces/mteb/leaderboard
[2] https://github.com/HKUNLP/instructor-embedding
-
I've made a customisable SMS personal assistant which has infinite and persistent semantic memory.
Use instructor-embedding to to make it 100% local and even maybe quick relationship lookup (embed relationship info with sentiment analysis instruction)
-
Whisper Transcription Formatting
First.I believe having srt subtitles as whisper result would be better.Essentially you don't need just a list of words like YouTube does.You need something more structured.I don't remember what whisper outputs so I might be wrong.There is whisperx for that as example. And then maybe use gpt index over it.Or something like instructor model That can work.
OpenLLM
-
First 15 Open Source Advent projects
13. OpenLLM by BentoML | Github | tutorial
-
The Lost Arts of CLJS Frontend
for those who care: i am developing a cljs frontend for OpenLLM right now: https://github.com/bentoml/OpenLLM/pull/89
- Build Personal ChatGPT Using Your Data
- OpenLLM: An open platform for operating large language models (LLMs) in production.
-
GPT Weekly - 26the June Edition - 🎙️ Meta's Voicebox is Paused, 🖼️SDXL 0.9, 📜AI Compliance & EU Act and more
Run inference on any LLM using OpenLLM.
- OpenLLM: OSS to easily serve Open Source LLMs
What are some alternatives?
h2ogpt - Private chat with local GPT with document, images, video, etc. 100% private, Apache 2.0. Supports oLLaMa, Mixtral, llama.cpp, and more. Demo: https://gpt.h2o.ai/ https://codellama.h2o.ai/
FastChat - An open platform for training, serving, and evaluating large language models. Release repo for Vicuna and Chatbot Arena.
openai-cookbook - Examples and guides for using the OpenAI API
gpt4all - gpt4all: run open-source LLMs anywhere
Nuggt - An Autonomous LLM Agent that runs on Wizcoder-15B
text-generation-webui - A Gradio web UI for Large Language Models. Supports transformers, GPTQ, AWQ, EXL2, llama.cpp (GGUF), Llama models.
vlite - fast vector database made in numpy
alpaca-lora - Instruct-tune LLaMA on consumer hardware
easydiffusion - Easiest 1-click way to create beautiful artwork on your PC using AI, with no tech knowledge. Provides a browser UI for generating images from text prompts and images. Just enter your text prompt, and see the generated image.
lit-gpt - Hackable implementation of state-of-the-art open-source LLMs based on nanoGPT. Supports flash attention, 4-bit and 8-bit quantization, LoRA and LLaMA-Adapter fine-tuning, pre-training. Apache 2.0-licensed. [Moved to: https://github.com/Lightning-AI/litgpt]
unilm - Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalities