Rapid-MLX

Rapid-MLX

Rapid-MLX is an open-source LLM inference server designed for Apple Silicon, supporting OpenAI and Anthropic APIs.

MCPDevOpen source
Type
MCP
Transport
http
Open source
Yes
GitHub Stars
★ 3.9k
Source
mcp-github

Overview

Rapid-MLX is an open-source (Apache 2.0) LLM inference server built on MLX, specifically designed for Apple Silicon. It supports OpenAI and Anthropic-compatible APIs and focuses on providing reliable tool calling for coding agents. Compared to Apple's MLX (mlx-lm), Rapid-MLX achieves up to 4x faster performance with the same weights. Ideal for developers and researchers needing efficient inference and tool calling.

Capabilities

  • ▪OpenAI and Anthropic-compatible APIs
  • ▪Supports Apple Silicon
  • ▪Efficient tool call parsing
  • ▪Concurrent request batching
  • ▪In-memory Radix prefix caching

Use cases

Code generation and debuggingAutomated script writingRapid prototypingNatural language processing tasks

Setup

Requires: Python 3.10+Apple Silicon (M1, M2, M3, M4)
pip install rapid-mlx or install via Homebrew: brew install rapid-mlx

This information was compiled by AI from public sources and may contain inaccuracies — please refer to the source.

FAQ

Rapid-MLX supports which hardware?

Supports Apple Silicon (M1, M2, M3, M4)

Is Rapid-MLX open source?

Yes, Rapid-MLX is an open-source project under the Apache 2.0 license

Related skills