instructor-embedding
PdfGptIndexer
instructor-embedding | PdfGptIndexer | |
---|---|---|
4 | 4 | |
1,703 | 637 | |
3.1% | - | |
5.9 | 4.8 | |
10 days ago | 10 months ago | |
Python | Python | |
Apache License 2.0 | MIT License |
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.
instructor-embedding
-
My experience on starting with fine tuning LLMs with custom data
If you li embeddings and vector DB, you should look into this: https://github.com/HKUNLP/instructor-embedding
-
Build Personal ChatGPT Using Your Data
If you look at a embeddings leaderboard [1], one of the top competitors called InstructorXL [2] is just a pip install away. It's neck and neck with Ada v2 except for a shorter input length and half the dimensions, with the added benefit that you'll always have the model available.
Most of the other options just work with the transformers library.
[1] https://huggingface.co/spaces/mteb/leaderboard
[2] https://github.com/HKUNLP/instructor-embedding
-
I've made a customisable SMS personal assistant which has infinite and persistent semantic memory.
Use instructor-embedding to to make it 100% local and even maybe quick relationship lookup (embed relationship info with sentiment analysis instruction)
-
Whisper Transcription Formatting
First.I believe having srt subtitles as whisper result would be better.Essentially you don't need just a list of words like YouTube does.You need something more structured.I don't remember what whisper outputs so I might be wrong.There is whisperx for that as example. And then maybe use gpt index over it.Or something like instructor model That can work.
PdfGptIndexer
What are some alternatives?
h2ogpt - Private chat with local GPT with document, images, video, etc. 100% private, Apache 2.0. Supports oLLaMa, Mixtral, llama.cpp, and more. Demo: https://gpt.h2o.ai/ https://codellama.h2o.ai/
OpenLLM - Run any open-source LLMs, such as Llama 2, Mistral, as OpenAI compatible API endpoint in the cloud.
openai-cookbook - Examples and guides for using the OpenAI API
gpt-2 - Code for the paper "Language Models are Unsupervised Multitask Learners"
Nuggt - An Autonomous LLM Agent that runs on Wizcoder-15B
vlite - fast vector database made in numpy
gpt4all - gpt4all: run open-source LLMs anywhere
easydiffusion - Easiest 1-click way to create beautiful artwork on your PC using AI, with no tech knowledge. Provides a browser UI for generating images from text prompts and images. Just enter your text prompt, and see the generated image.
private-gpt - Interact with your documents using the power of GPT, 100% privately, no data leaks
lit-gpt - Hackable implementation of state-of-the-art open-source LLMs based on nanoGPT. Supports flash attention, 4-bit and 8-bit quantization, LoRA and LLaMA-Adapter fine-tuning, pre-training. Apache 2.0-licensed. [Moved to: https://github.com/Lightning-AI/litgpt]
buzz - Buzz transcribes and translates audio offline on your personal computer. Powered by OpenAI's Whisper.