llama.cpp
StableLM
llama.cpp | StableLM | |
---|---|---|
773 | 43 | |
56,891 | 15,852 | |
- | 0.2% | |
10.0 | 5.0 | |
5 days ago | 23 days ago | |
C++ | Jupyter Notebook | |
MIT License | Apache License 2.0 |
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.
llama.cpp
-
Better and Faster Large Language Models via Multi-Token Prediction
For anyone interested in exploring this, llama.cpp has an example implementation here:
https://github.com/ggerganov/llama.cpp/tree/master/examples/...
- Llama.cpp Bfloat16 Support
-
Fine-tune your first large language model (LLM) with LoRA, llama.cpp, and KitOps in 5 easy steps
Getting started with LLMs can be intimidating. In this tutorial we will show you how to fine-tune a large language model using LoRA, facilitated by tools like llama.cpp and KitOps.
- GGML Flash Attention support merged into llama.cpp
-
Phi-3 Weights Released
well https://github.com/ggerganov/llama.cpp/issues/6849
- Lossless Acceleration of LLM via Adaptive N-Gram Parallel Decoding
- Llama.cpp Working on Support for Llama3
-
Embeddings are a good starting point for the AI curious app developer
Have just done this recently for local chat with pdf feature in https://recurse.chat. (It's a macOS app that has built-in llama.cpp server and local vector database)
Running an embedding server locally is pretty straightforward:
- Get llama.cpp release binary: https://github.com/ggerganov/llama.cpp/releases
- Mixtral 8x22B
- Llama.cpp: Improve CPU prompt eval speed
StableLM
-
The Era of 1-bit LLMs: ternary parameters for cost-effective computing
https://github.com/Stability-AI/StableLM?tab=readme-ov-file#...
-
Stable LM 3B: Bringing Sustainable, High-Performance LMs to Smart Devices
https://mistral.ai/news/announcing-mistral-7b/
looking at the 3b results (here https://github.com/Stability-AI/StableLM#stablelm-alpha-v2 ?), it looks like Mistral (which outperforms Llama-2 13b) is far more powerful
-
FreeWilly 1 and 2, two new open-access LLMs
Does this mean Stability gave up on StableLM?
I notice that the repo hasn’t been updated since April, and a question asking for an update has been ignored for at least a month: https://github.com/Stability-AI/StableLM/issues/83
-
In five years, there will be no programmers left, believes Stability AI CEO
I'm not "ignoring" StableLM, if anything it's the impetus for my post. The alpha models were so bad and unusable that it seems they may have simply abandoned the project. It's clear they basically didn't know what they were doing, which is silly for a company of their size and specialization.
-
Losing the plot
1) StableLM released a checkpoint at 800B for their 3B and 7B at 800B tokens with 4096 context size, but perform very poorly on different benchmarks and finetuning is discouraged with such a weak base model
-
UAE's Technology Innovation Institute Launches Open-Source "Falcon 40B" Large Language Model for Research & Commercial Utilization
It is the best open-source model currently available. Falcon-40B outperforms LLaMA, StableLM, RedPajama, MPT, etc. See the OpenLLM Leaderboard.
- Consulta API GPT
- Google "We Have No Moat, And Neither Does OpenAI"
-
New to StableLM--is it possible to use this locally to fine-tune on a small subset of documents yet?
Someone shared this link on another recent post
-
[N] Stability AI releases StableVicuna: the world's first open source chatbot trained via RLHF
Github: https://github.com/Stability-AI/StableLM
What are some alternatives?
ollama - Get up and running with Llama 3, Mistral, Gemma, and other large language models.
text-generation-webui - A Gradio web UI for Large Language Models. Supports transformers, GPTQ, AWQ, EXL2, llama.cpp (GGUF), Llama models.
gpt4all - gpt4all: run open-source LLMs anywhere
lm-evaluation-harness - A framework for few-shot evaluation of language models.
ggml - Tensor library for machine learning
GPTQ-for-LLaMa - 4 bits quantization of LLaMA using GPTQ
Open-Assistant - OpenAssistant is a chat-based assistant that understands tasks, can interact with third-party systems, and retrieve information dynamically to do so.
alpaca_lora_4bit
alpaca.cpp - Locally run an Instruction-Tuned Chat-Style LLM
llama - Inference code for Llama models