ninfer

High-performance single-GPU inference for selected model checkpoints and GPUs. (by Neroued)

Ninfer Alternatives

Similar projects and alternatives to ninfer

  1. llama.cpp

    LLM inference in C/C++

  2. AppSignal

    Monitoring that respects your time & budget. APM, error tracking, and dashboards for modern web apps. Ten-minute setup, transparent flat pricing, and support from engineers who actually use the product.

    AppSignal logo
  3. caveman

    41 ninfer VS caveman

    🪨 why use many token when few token do trick — Claude Code skill that cuts 65% of tokens by talking like caveman

  4. ds4

    DeepSeek 4 Flash local inference engine for Metal

  5. warp

    Run the full 2.78-trillion-parameter Kimi K3 model or GLM-5.3-Flash beyond available RAM by streaming activated weights directly from NVMe. A dependency-free, embeddable C inference engine. (by sqliteai)

  6. models.dev

    7 ninfer VS models.dev

    An open-source database of AI models.

  7. turbo-fieldfare

    Gemma 4 26B-A4B inference in ~2 GB of RAM on any M-series MacBook

  8. Kargo

    Stop Scripting Promotions. Start Shipping with Kargo. Kargo automates promotion across dev, staging, and prod with approval gates and verification. Open source, built by the team behind Argo CD. Download now.

    Kargo logo
  9. MTPLX

    Native MTP Speculative Decoding On Apple Silicon | 2x - 2.5x decode TPS increase at temp 0.6 | MLX-native, OpenAI API/Anthropic-compatible serving, no external drafter.

  10. Biscotti

    Free private meeting transcription app for macOS

  11. FreeToken

    2 ninfer VS FreeToken

    FreeToken brings datacenter-scale model serving to your desktop. Run massive models locally, fast and efficiently.

  12. Swiftlet

    2 ninfer VS Swiftlet

    Swiftlet is a Swift and Metal runtime that runs large Qwen Mixture-of-Experts models locally on Apple devices by streaming expert weights from storage, enabling 35B and 80B models to run with low RAM, including on iPhone.

  13. memra

    2 ninfer VS memra

    Rust + CUDA LLM inference engine for Blackwell (Tuned specifically on RTX PRO 6000, RTX 5090, B200): OpenAI-compatible (+converse and ant) serving, per-model X hardware exactness gates. NVFP4/mixed (fp8 hybrid, 4o6, etc - correctness, performance, hardware specific adapted) main quant support.

  14. ninfer-3090

    Hyper optimised Qwen3.8-27B inference on one RTX 3090: ReplaySSM, MTP3, reasoning effort, C1-C8 batching, and native Windows and Linux builds.

  15. listing

    simple shared file manager for docker/Synology

  16. RusTODO

    1 ninfer VS RusTODO

    A tiny, fast todo list that lives on your desktop.

  17. SaaSHub

    SaaSHub - Software Alternatives and Reviews. SaaSHub helps you find the best software and product alternatives

    SaaSHub logo
NOTE: The number of mentions on this list indicates mentions on common posts plus user suggested alternatives. Hence, a higher number means a better ninfer alternative or higher similarity.

ninfer discussion

Log in or Post with

ninfer reviews and mentions

Posts with mentions or reviews of ninfer. We have used some of these posts to build our list of alternatives and similar projects. The last one was on 2026-09-03.

Stats

Basic ninfer repo stats
9
1,562
9.9
1 day ago

Neroued/ninfer is an open source project licensed under Apache License 2.0 which is an OSI approved license.

The primary programming language of ninfer is C++.


Sponsored
Monitoring that respects your time & budget
APM, error tracking, and dashboards for modern web apps. Ten-minute setup, transparent flat pricing, and support from engineers who actually use the product.
www.appsignal.com