vllm-mlx

High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal models, MCP tool calling, and Claude Code support. (by waybarrios)

Vllm-mlx Alternatives

Similar projects and alternatives to vllm-mlx

  1. ollama

    Get up and running with Kimi-K2.6, GLM-5.2, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.

  2. AppSignal

    Monitoring that respects your time & budget. APM, error tracking, and dashboards for modern web apps. Ten-minute setup, transparent flat pricing, and support from engineers who actually use the product.

    AppSignal logo
  3. n8n

    464 vllm-mlx VS n8n

    n8n is a workflow automation platform for building AI-powered workflows and agents, connecting any AI model to any business system with full control over data, security, and deployment. Build visually or in code while n8n handles infrastructure from prototype to production with fully auditable executions.

  4. autogen

    75 vllm-mlx VS autogen

    A programming framework for agentic AI

  5. mlx

    52 vllm-mlx VS mlx

    MLX: An array framework for Apple silicon

  6. llama.cpp

    Discontinued LLM inference in C/C++ [Moved to: https://github.com/ggml-org/llama.cpp] (by ggerganov)

  7. claude-code-router

    One local control plane for every AI agent: route across models, fuse new capabilities, orchestrate tools, and stay fully in control.

  8. crewAI

    32 vllm-mlx VS crewAI

    Framework for orchestrating role-playing, autonomous AI agents. By fostering collaborative intelligence, CrewAI empowers agents to work together seamlessly, tackling complex tasks.

  9. Kargo

    Stop Scripting Promotions. Start Shipping with Kargo. Kargo automates promotion across dev, staging, and prod with approval gates and verification. Open source, built by the team behind Argo CD. Download now.

    Kargo logo
  10. qmd

    15 vllm-mlx VS qmd

    mini cli search engine for your docs, knowledge bases, meeting notes, whatever. Tracking current sota approaches while being all local

  11. hyperlearn

    12 vllm-mlx VS hyperlearn

    2-2000x faster ML algos, 50% less memory usage, works on all hardware - new and old.

  12. vllm-metal

    Community maintained hardware plugin for vLLM on Apple Silicon

  13. security-ops-platform

    Discontinued Open-source security detection & response platform: 50+ tool integrations, an on-prem LLM investigation-agent fleet with MCP, self-healing Webex/Teams bots, and 80+ Flask SOC web apps. [GET https://api.github.com/repos/vinayvobbili/security-ops-platform: 404 - Not Found // See: https://docs.github.com/rest/repos/repos#get-a-repository]

  14. langchain-failover

    Primary/secondary failover wrapper for LangChain chat models, with tool-calling preserved across failover.

NOTE: The number of mentions on this list indicates mentions on common posts plus user suggested alternatives. Hence, a higher number means a better vllm-mlx alternative or higher similarity.

vllm-mlx discussion

Log in or Post with

vllm-mlx reviews and mentions

Posts with mentions or reviews of vllm-mlx. We have used some of these posts to build our list of alternatives and similar projects. The last one was on 2026-06-06.

Stats

Basic vllm-mlx repo stats
6
1,566
9.7
6 days ago

Sponsored
Monitoring that respects your time & budget
APM, error tracking, and dashboards for modern web apps. Ten-minute setup, transparent flat pricing, and support from engineers who actually use the product.
www.appsignal.com

Did you know that Python is
the 1st most popular programming language
based on number of references?