BigMoeOnEdge Alternatives
Similar projects and alternatives to BigMoeOnEdge based on common topics and language
-
QuantumLeap---Llama.cpp-TurboQuant
๐ Run any LLM on any hardware. 130% faster MoE inference with ExpertFlow + TurboQuant KV compression. Ollama-compatible API. Built on llama.cpp.
-
Kargo
Stop Scripting Promotions. Start Shipping with Kargo. Kargo automates promotion across dev, staging, and prod with approval gates and verification. Open source, built by the team behind Argo CD. Download now.
-
xyntetik-runner
Single-binary GGUF model runtime in C โ CPU/CUDA/Metal, OpenAI-compatible. Serves, scores, and trains LoRA directly through the quantized weights it deploys, with byte-reproducible adapters. Tool calls survive the token limit; sparse MoE and schema-constrained decoding included.
-
fitllm-engine
Accurate LLM memory & speed calculator โ models sliding-window/linear/MoE attention where naive VRAM/RAM calculators are 4โ11ร off. Apple Silicon + NVIDIA RTX ยท GGUF ยท one MIT file. Powers fitllm.run.
-
TensorSharp
A C# inference engine for running large language models (LLMs) locally using GGUF model files. TensorSharp provides a console application, a web-based chatbot interface, and Ollama/OpenAI-compatible HTTP APIs for programmatic access. It supports Windows/MacOS/Linux with full GPU capability
-
-
Qwen-3.5-16G-Vram-Local
Discontinued Configs, launchers, benchmarks, and tooling for running Qwen3.5 GGUF models locally with llama.cpp on a 16GB NVIDIA GPU [GET https://api.github.com/repos/willbnu/Qwen-3.5-16G-Vram-Local: 404 - Not Found // See: https://docs.github.com/rest]
-
mbolt
Profile-guided layout optimization for llama.cpp GGUF MoE models โ BOLT/PGO for LLM weights. Trace expert routing, rewrite the file so co-activated experts sit together, prefetch as merged reads.
-
AppSignal
AppSignal knows why the f*#k it crashed. Stop vibe-debugging. Every exception, every backtrace, grouped so you see patterns, not noise.
BigMoeOnEdge discussion
BigMoeOnEdge reviews and mentions
Stats
Helldez/BigMoeOnEdge is an open source project licensed under Apache License 2.0 which is an OSI approved license.
The primary programming language of BigMoeOnEdge is C++.