MTPLX
View on GitHubThe fastest way to run Qwen 3.8 Flash Next and Qwen 3.8 27B on a Mac: 125 tok/s in OpenCode on an M5 Max. Native MTP speculative decoding on Apple Silicon, exact at any temperature. OpenAI and Anthropic compatible local server.
MTPLX is an Apple Silicon LLM inference engine and Mac app that uses Qwen's native multi-token prediction heads for exact speculative decoding (1.6x-2.24x faster decode at any temperature). It serves local models over OpenAI- and Anthropic-compatible APIs for coding agents and chat.
Use Cases
Run Qwen models locally on Apple Silicon MacsServe an OpenAI-compatible local API endpointServe an Anthropic-compatible endpoint for Claude CodeAccelerate LLM decoding with native MTP speculative decodingPower local coding agents like OpenCode, Cline and CursorChat with local models via a native desktop appTune decoding depth to a specific Mac's hardwareBenchmark local model throughput (tok/s)Run large MoE models with SSD-streamed n-gram tablesCollect and manage local model libraries
Built With
- Language
- Python
- Frameworks
- MLX · mlx-lm · FastAPI · Hugging Face Hub · Transformers · setuptools · Vite · React
Tags
apple-silicon · mlx · metal · speculative-decoding · mtp · inference-engine · local-llm · openai-compatible · anthropic-compatible · quantization · macos · qwen · local-server · cli · desktop-app · turbo-decoding