Rapid-MLX
Rapid-MLX is an open-source LLM inference server designed for Apple Silicon, supporting OpenAI and Anthropic APIs.
- Type
- MCP
- Transport
- http
- Open source
- Yes
- GitHub Stars
- ★ 3.9k
- Source
- mcp-github
- Repository
- github.com/raullenchai/Rapid-MLX
Overview
Rapid-MLX is an open-source (Apache 2.0) LLM inference server built on MLX, specifically designed for Apple Silicon. It supports OpenAI and Anthropic-compatible APIs and focuses on providing reliable tool calling for coding agents. Compared to Apple's MLX (mlx-lm), Rapid-MLX achieves up to 4x faster performance with the same weights. Ideal for developers and researchers needing efficient inference and tool calling.
Capabilities
- ▪OpenAI and Anthropic-compatible APIs
- ▪Supports Apple Silicon
- ▪Efficient tool call parsing
- ▪Concurrent request batching
- ▪In-memory Radix prefix caching
Use cases
Setup
pip install rapid-mlx or install via Homebrew: brew install rapid-mlx
This information was compiled by AI from public sources and may contain inaccuracies — please refer to the source.
FAQ
Rapid-MLX supports which hardware?
Supports Apple Silicon (M1, M2, M3, M4)
Is Rapid-MLX open source?
Yes, Rapid-MLX is an open-source project under the Apache 2.0 license
Related skills
answer-me-with-html
Let an AI Agent answer complex questions with a single HTML page.
sast-skills
Transform your AI coding assistant into a SAST scanner to automatically detect vulnerabilities in code.
Lightswind-UI-Library
AI-native CLI-first React component library with MCP Server support.
scientific-agent-skills
Transform any AI agent into a scientific assistant with 147 ready-to-use research skills.
susi_alexa_skill
A skill that enables question-and-answer interactions between Alexa and Susi AI.
claude-context
Provides code search capabilities for Claude Code, turning the entire codebase into context.