BigMoeOnEdge

Run MoE models bigger than your RAM. Frontier-size MoE on a 12 GB phone, CPU only, lossless, on stock llama.cpp (by Helldez)

BigMoeOnEdge Alternatives

Similar projects and alternatives to BigMoeOnEdge based on common topics and language

  1. QuantumLeap---Llama.cpp-TurboQuant

    ๐Ÿš€ Run any LLM on any hardware. 130% faster MoE inference with ExpertFlow + TurboQuant KV compression. Ollama-compatible API. Built on llama.cpp.

  2. Kargo

    Stop Scripting Promotions. Start Shipping with Kargo. Kargo automates promotion across dev, staging, and prod with approval gates and verification. Open source, built by the team behind Argo CD. Download now.

    Kargo logo
  3. xyntetik-runner

    Single-binary GGUF model runtime in C โ€” CPU/CUDA/Metal, OpenAI-compatible. Serves, scores, and trains LoRA directly through the quantized weights it deploys, with byte-reproducible adapters. Tool calls survive the token limit; sparse MoE and schema-constrained decoding included.

  4. fitllm-engine

    Accurate LLM memory & speed calculator โ€” models sliding-window/linear/MoE attention where naive VRAM/RAM calculators are 4โ€“11ร— off. Apple Silicon + NVIDIA RTX ยท GGUF ยท one MIT file. Powers fitllm.run.

  5. TensorSharp

    A C# inference engine for running large language models (LLMs) locally using GGUF model files. TensorSharp provides a console application, a web-based chatbot interface, and Ollama/OpenAI-compatible HTTP APIs for programmatic access. It supports Windows/MacOS/Linux with full GPU capability

  6. runanywhere-sdks

    Production ready toolkit to run AI locally

  7. Qwen-3.5-16G-Vram-Local

    Discontinued Configs, launchers, benchmarks, and tooling for running Qwen3.5 GGUF models locally with llama.cpp on a 16GB NVIDIA GPU [GET https://api.github.com/repos/willbnu/Qwen-3.5-16G-Vram-Local: 404 - Not Found // See: https://docs.github.com/rest]

  8. mbolt

    Profile-guided layout optimization for llama.cpp GGUF MoE models โ€” BOLT/PGO for LLM weights. Trace expert routing, rewrite the file so co-activated experts sit together, prefetch as merged reads.

  9. AppSignal

    AppSignal knows why the f*#k it crashed. Stop vibe-debugging. Every exception, every backtrace, grouped so you see patterns, not noise.

    AppSignal logo
NOTE: The number of mentions on this list indicates mentions on common posts plus user suggested alternatives. Hence, a higher number means a better BigMoeOnEdge alternative or higher similarity.

BigMoeOnEdge discussion

Log in or Post with

BigMoeOnEdge reviews and mentions

Posts with mentions or reviews of BigMoeOnEdge. We have used some of these posts to build our list of alternatives and similar projects.

Stats

Basic BigMoeOnEdge repo stats
1
543
9.6
7 days ago

Helldez/BigMoeOnEdge is an open source project licensed under Apache License 2.0 which is an OSI approved license.

The primary programming language of BigMoeOnEdge is C++.


Sponsored
Stop Scripting Promotions. Start Shipping with Kargo
Kargo automates promotion across dev, staging, and prod with approval gates and verification. Open source, built by the team behind Argo CD. Download now.
akuity.io