omlx
View on GitHubLLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar
oMLX is an MLX-based LLM inference server for Apple Silicon with continuous batching, tiered hot-RAM/cold-SSD KV cache, multi-model serving, and VLMs/embeddings/rerankers behind an OpenAI-compatible API, managed from a macOS menu bar app and web dashboard.
Use Cases
Serve local LLMs on Apple Silicon via an OpenAI-compatible APIBack Claude Code / Codex / Copilot with a self-hosted model endpointRun vision-language and OCR models with continuous batchingReuse KV cache across requests with hot RAM + cold SSD tiersServe embeddings and rerankers alongside chat modelsDistributed multi-Mac inference over Ring/ThunderboltDownload and manage MLX models from HuggingFace in a dashboardBenchmark prefill/generation throughput from the admin UI
Built With
- Language
- Python
- Frameworks
- MLX · mlx-lm · mlx-vlm · mlx-embeddings · FastAPI · SwiftUI · JACCL
Tags
llm-inference · apple-silicon · mlx · openai-api · continuous-batching · kv-cache · local-llm · macos · menubar-app · vlm · embeddings · reranker · quantization · mcp · self-hosted · ssd-cache