omlx
★ 22KoMLX is an MLX-based LLM inference server for Apple Silicon with continuous batching, tiered hot-RAM/cold-SSD KV cache, multi-model serving, and VLMs/embeddings/rerankers behind an OpenAI-compatible API, managed from a macOS menu bar app and web dashboard.
AI Frameworks | Python · llm-inference · apple-silicon
View Project →