SaaSHub helps you find the best software and product alternatives Learn more →
Top 12 Rust llm-inference Projects
-
plano
Plano is an AI-native proxy server and data plane for agentic apps. Smart LLM routing, observability, agent orchestration, and guardrails so you stay focused on your agents core logic.
Project mention: Show HN: Signals – finding the most informative agent traces without LLM judges | news.ycombinator.com | 2026-04-04Project where Signals are already implemented: https://github.com/katanemo/plano
Happy to answer questions on the taxonomy, implementation details, or where this breaks down.
-
AppSignal
Monitoring that respects your time & budget. APM, error tracking, and dashboards for modern web apps. Ten-minute setup, transparent flat pricing, and support from engineers who actually use the product.
-
shimmy
⚡ Pure-Rust WebGPU inference engine — OpenAI-API compatible, GGUF native, runs on any GPU. No Python. No llama.cpp. Single binary.
Project mention: Shimmy v1.7.0: Running 42B Moe Models on Consumer GPUs with 99.9% VRAM Reduction | news.ycombinator.com | 2025-10-08 -
RuVector
RuVector is a High Performance, Real-Time, Self-Learning Ai, Vector GNN, Memory DB built in Rust.
Project mention: High Performance, Real-Time, Self-Learning, Vector GNN and DB Built in Rust | news.ycombinator.com | 2026-02-26 -
-
-
Project mention: Qwen 3.6 27B is the sweet spot for local development | news.ycombinator.com | 2026-06-29
I usually doubt the 'small dataset tuned' variants. B/c ages ago (in the NN prehistory) I've done some NN training, and appreciate how hard it is to improve in general, and how easy it is to ruin a model in general while targeting a small dataset (LoRA-s are ok, that's different). That model/quant was the most recent one I was trying. But could not really use any of them, as the combo model + llama-server ground to a halt even at small context depth sizes on the amd gpu.
Yesterday I finally found a good combo! So writing this for the benefit for anyone that may have the same h/w. Got around to /GOAL search for something better for the h/w (amd 7900xtx), and pi agent found a new best that actually seems it will be useful for real. As the 40 tok/s speed starts dropping only at 260K context depth?? Served by hipfire from this repo https://github.com/Kaden-Schutt/hipfire, that worked the best got on llama-benchy:
| Context | Wall Time | PP (t/s) | TG (t/s) |
-
pmetal
PMetal: high-performance Apple Silicon framework for local LLM inference, LoRA/QLoRA fine-tuning, serving, quantization, and MLX/Metal acceleration.
Project mention: PMetal – (Powdered Metal) LLM Fine-Tuning Framework for Apple Silicon | news.ycombinator.com | 2026-03-16 -
Kargo
Stop Scripting Promotions. Start Shipping with Kargo. Kargo automates promotion across dev, staging, and prod with approval gates and verification. Open source, built by the team behind Argo CD. Download now.
-
-
Project mention: Atlas: An LLM inference engine written from scratch in Rust and CUDA | news.ycombinator.com | 2026-05-12
-
saient-quartz
Quartz — custom CUDA/Vulkan GGUF inference engine + SD1.5/SDXL image and video generation (Rust, no llama.cpp)
Project mention: Open-sourced a local WAN 2.2 mobile pipeline for on-device video generation | news.ycombinator.com | 2026-08-01 -
-
gguf-switchboard
Run GGUF through llama.cpp and SafeTensors through vLLM behind one OpenAI-compatible endpoint. Your coding tools select a model; the switchboard manages the local runtime, process, and resident-model change.
GGUF Switchboard is open source: GitHub — GGUF Switchboard
Rust llm-inference discussion
Rust llm-inference related posts
-
Atlas: An LLM inference engine written from scratch in Rust and CUDA
-
Atlas – Pure Rust Inference Engine
-
Show HN: Signals – finding the most informative agent traces without LLM judges
-
Show HN: Preference-aware routing for OpenClaw via Plano
-
Agent Safety is a [bounded] Box
-
Show HN: Plano – Edge and service proxy with orchestration for AI agents
-
Show HN: Claude Code router 2.0 – preference-aligned routing to multiple LLMs
-
A note from our sponsor - SaaSHub
www.saashub.com | 10 Sep 2026