mamba | llm.f90 | |
---|---|---|
15 | 13 | |
9,506 | 48 | |
15.3% | - | |
8.1 | 8.4 | |
5 days ago | about 2 months ago | |
Python | Fortran | |
Apache License 2.0 | MIT License |
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.
mamba
-
Based: Simple linear attention language models
> how the recall can grow unbounded with no tradeoff
this? https://github.com/state-spaces/mamba/issues/175
-
Mamba: The Easy Way
If you want to learn this stuff as a computer engineer, you can read the code here [0]. I find the math quite helpful.
[0]: https://github.com/state-spaces/mamba
- FLaNK Stack 05 Feb 2024
- Introduction to State Space Models (SSM)
-
Fortran inference code for the Mamba state space language model
This model was discussed recently: https://news.ycombinator.com/item?id=38522428 It's a new kind of ML model architecture that can be used instead of a transformer in LLMs.
See also the original repo from the paper: https://github.com/state-spaces/mamba
-
Mamba outperforms transformers "everywhere we tried"
[2] - https://github.com/state-spaces/mamba
Out of curiosity, does anyone feel as though there's any benefit to linking to reddit when we can link to whatever the link is? I for one do not click the link and read discussion on reddit - if I wanted that sort of discussion, I would browse there, not HN.
- GitHub – State-Spaces/Mamba
-
Generate valid JSON with Mamba models
The library is compatible with any auto-regressive model, not transformers. To prove our point we integrated Mamba, a new state-space model architecture, to the library. Try it out!
-
[D] Thoughts on Mamba?
I ran the NanoGPT of Karparthy replacing Self-Attention with Mamba on his TinyShakespeare Dataset and within 5 minutes it started spitting out the following:
-
Mamba-Chat: A Chat LLM based on State Space Models
You might have come across the paper Mamba paper in the last days, which was the first attempt at scaling up state space models to 2.8B parameters to work on language data.
llm.f90
- llm.f90: LLM Inference in Fortran
-
karpathy/llm.c
I'd like to think he took the name from my llm.f90 project https://github.com/rbitr/llm.f90
It was originally based off of Karpathy's llama2.c but I renamed it when I added support for other architectures.
Probable a coincidence :)
-
Winteracter – The Fortran GUI Toolset
I'm a Fortran hobbyist. I'm working (unfortunately less frequently now) on a LLM framework in Fortan: https://github.com/rbitr/llm.f90
- Fortran implementation of phi-2 LLM
- Fortran implementation of phi-2 language model
-
TinyLlama: An Open-Source Small Language Model
Also, I should promote the code I wrote for running this. It runs models in ggml format, the one I made available is an older checkpoint though. It's easy to convert the newer one. And it's in Fortran but it should be easy to get gfortran if you don't have it installed.
https://github.com/rbitr/llm.f90/tree/optimize16/purefortran
- Mamba LLM Inference on CPU
-
Minimal implementation of Mamba, the new LLM architecture, in 1 file of PyTorch
The original mamba code has a lot of speed optimizations and other stuff that make it difficult to immediately get so this will help with learning.
I can't help but also plug my own Mamba inference implementation. https://github.com/rbitr/llm.f90/tree/master/ssm
- Mamba state-space LLM inference
-
Guide to the Mamba architecture that claims to be a replacement for Transformers
You may also be interested in https://github.com/rbitr/llm.f90/tree/master/ssm it's my inference only implementation of mamba which ends up being much simpler than the training code in the original repo
What are some alternatives?
miniforge - A conda-forge distribution.
rwkv.f90 - Port of the RWKV-LM model in Fortran (Back to the Future!)
pip - The Python package installer
neural-fortran - A parallel framework for deep learning
conda - A system-level, binary package and environment manager running on all major operating systems and platforms.
inference-engine - A deep learning library for use in high-performance computing applications in modern Fortran
mamba-chat - Mamba-Chat: A chat LLM based on the state-space model architecture 🐍
fastGPT - Fast GPT-2 inference written in Fortran
spack - A flexible package manager that supports multiple versions, configurations, platforms, and compilers.
mamba-minimal - Simple, minimal implementation of the Mamba SSM in one file of PyTorch.
pyenv - Simple Python version management
Fortran-code-on-GitHub - Directory of Fortran codes on GitHub, arranged by topic