minimal-llama vs langchain

minimal-llama

By zphang

langchain

⚡ Building applications with LLMs through composability ⚡ [Moved to: https://github.com/langchain-ai/langchain] (by hwchase17)

Suggest topics

DISCONTINUED

Suggest alternative

Edit details

InfluxDB - Power Real-Time Data Analytics at Scale

Get real-time insights from all types of time series data with InfluxDB. Ingest, query, and analyze billions of data points in real-time with unbounded cardinality.

www.influxdata.com

featured

SaaSHub - Software Alternatives and Reviews

SaaSHub helps you find the best software and product alternatives

www.saashub.com

featured

minimal-llama		langchain
	Project
4	Mentions	152
456	Stars	56,526
-	Growth	-
8.5	Activity	10.0
7 months ago	Latest Commit	9 months ago
Python	Language	Python
-	License	MIT License

The number of mentions indicates the total number of mentions that we've tracked plus the number of user suggested alternatives.
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.

minimal-llama

Posts with mentions or reviews of minimal-llama. We have used some of these posts to build our list of alternatives and similar projects. The last one was on 2023-03-21.

Show HN: Finetune LLaMA-7B on commodity GPUs using your own text
16 projects | news.ycombinator.com | 21 Mar 2023
Visual ChatGPT
8 projects | news.ycombinator.com | 9 Mar 2023

I can't edit my comment now, but it's 30B that needs 18GB of VRAM.
LLaMA-13B, GPT-3 175B level, only needs 10GB of VRAM with the GPTQ 4bit quantization.
>do you think there's anything left to trim? like weight pruning, or LoRA, or I dunno, some kind of Huffman coding scheme that lets you mix 4-bit, 2-bit and 1-bit quantizations?
Absolutely. The GPTQ paper claims negligible output quality loss with 3-bit quantization. The GPTQ-for-LLaMA repo supports 3-bit quantization and inference. So this extra 25% savings is already possible.
As of right GPTQ-for-LLaMA is using a VRAM hungry attention method. Flash attention will reduce the requirements for 7B to 4GB and possibly fit 30B with a 2048 context window into 16GB, all before stacking 3-bit.
Pruning is a possibility but I'm not aware of anyone working on it yet.
LoRa has already been implemented. See https://github.com/zphang/minimal-llama#peft-fine-tuning-wit...

langchain

Posts with mentions or reviews of langchain. We have used some of these posts to build our list of alternatives and similar projects. The last one was on 2023-07-20.

🗣️🤖 Ask to your Neo4J knowledge base in NLP & get KPIs
2 projects | dev.to | 20 Jul 2023

Langchain and the implementation of Custom Tools also is a great (and very efficient) way to setup a dedicated Q&A (for example for chat purpose) agent.
LangChain – Some quick, high level thoughts on improvements/changes
1 project | news.ycombinator.com | 18 Jul 2023
Claude 2 Internal API Client and CLI
4 projects | news.ycombinator.com | 14 Jul 2023

We're using it via langchain talking to Amazon Bedrock which is hosting Claude 1.x. It's comparable to GPT3.x, not bad. The integration doesn't seem to be fully there though, I think langchain is expecting "Human:" and "AI:", but Claude uses "Assistant:".
https://github.com/hwchase17/langchain/issues/2638
Any better alternatives to fine-tuning GPT-3 yet to create a custom chatbot persona based on provided knowledge for others to use?
4 projects | /r/ChatGPTPro | 12 Jul 2023

Depending on how much work you want to put into it, you can get started at HuggingFace with their models and datasets, but you'd need compute power, multiple MLOps, etc. I was introduced to the concept in this video, since Google has their Vertex AI tools on Google Cloud, and there's always LangChain but I'm not sure about anything recent.
langchain VS griptape - a user suggested alternative
2 projects | 11 Jul 2023

2 projects | 9 Jul 2023
Vector storage is coming to Meilisearch to empower search through AI
6 projects | dev.to | 11 Jul 2023

a documentation chatbot proof of concept using GPT3.5 and LangChain
ChatPDF: What ChatGPT Can't Do, This Can!
2 projects | /r/LangChain | 10 Jul 2023

I encourage everyone to pay attention to the Langchain open-source project and leverage it to achieve tasks that ChatGPT cannot handle.
LangChain Arbitrary Command Execution - CVE-2023-34541
2 projects | dev.to | 9 Jul 2023
Langchain Is Pointless
16 projects | news.ycombinator.com | 8 Jul 2023

Yeah I never know where memory goes exactly in langchain, it's not exactly clear all the time. But sure, the main insight I remember is this, take a look at their MULTI_PROMPT_ROUTER_TEMPLATE: https://github.com/hwchase17/langchain/blob/560c4dfc98287da1...
It's a lot of instructions for an LLM, they seem to forget an LLM is an auto-completion machine, and which data it is trained on. Using <<>> for sections is not a normal thing, it's not markdown, which probably the thing read way more often on the internet, instead of open json comments, why not type signatures, instead of so many rules, why not give it examples? It is an autocomplete machine!
They are relying too much on the LLM being smart because they probably only test stuff in GPT-4 and 3.5, but with GPT4All models this prompt was not working at all, so I had to rewrite it, for simple routing, we don't even need json, carying the `next_inputs` here is weird if you don't need it.
So this is my version of it: https://gist.github.com/rogeriochaves/b67676977eebb1936b9b5c...
It's so basic it's dumb, yet it is more powerful, as it does not rely on GPT-4 level intelligence, it's just what I needed

What are some alternatives?

When comparing minimal-llama and langchain you can also consider the following projects:

FlexGen - Running large language models on a single GPU for throughput-oriented scenarios.

semantic-kernel - Integrate cutting-edge LLM technology quickly and easily into your apps

visual-chatgpt - Official repo for the paper: Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models [Moved to: https://github.com/microsoft/TaskMatrix]

llama_index - LlamaIndex is a data framework for your LLM applications

whisper.cpp - Port of OpenAI's Whisper model in C/C++

llama - Inference code for Llama models

simple-llm-finetuner - Simple UI for LLM Model Finetuning

text-generation-webui - A Gradio web UI for Large Language Models. Supports transformers, GPTQ, AWQ, EXL2, llama.cpp (GGUF), Llama models.

alpaca-lora - Instruct-tune LLaMA on consumer hardware

gpt_index - LlamaIndex (GPT Index) is a project that provides a central interface to connect your LLM's with external data. [Moved to: https://github.com/jerryjliu/llama_index]

GPTQ-for-LLaMa - 4 bits quantization of LLaMA using GPTQ

AutoGPT - AutoGPT is the vision of accessible AI for everyone, to use and to build on. Our mission is to provide the tools, so that you can focus on what matters.

minimal-llama vs FlexGen langchain vs semantic-kernel minimal-llama vs visual-chatgpt langchain vs llama_index minimal-llama vs whisper.cpp langchain vs llama minimal-llama vs simple-llm-finetuner langchain vs text-generation-webui minimal-llama vs alpaca-lora langchain vs gpt_index minimal-llama vs GPTQ-for-LLaMa langchain vs AutoGPT

Compare minimal-llama vs langchain and see what are their differences.

minimal-llama

langchain

minimal-llama

langchain

What are some alternatives?