mlx-lm
View on GitHubRun LLMs with MLX
MLX LM is a Python package for running and fine-tuning LLMs on Apple silicon via MLX. It offers CLI and Python APIs for generation, chat, LoRA/full fine-tuning, quantization, GGUF conversion, prompt caching, and an OpenAI-compatible server.
Use Cases
Run LLMs locally on Apple siliconFine-tune models with LoRA or full fine-tuningQuantize and upload models to Hugging Face HubServe an OpenAI-compatible local inference serverInteractive chat REPL with a local modelCache long prompts for reuse across queriesDistributed inference and fine-tuning across MacsConvert HF models to MLX formatEvaluate perplexity and benchmark throughputStream token generation via a Python API
Built With
- Language
- Python
- Frameworks
- MLX · Hugging Face Hub · Transformers · safetensors · LoRA · GGUF · AWQ · GPTQ
Tags
llm-inference · apple-silicon · mlx · quantization · fine-tuning · lora · gguf · text-generation · prompt-caching · local-llm · openai-compatible · cli · chat-repl · model-conversion · distributed-inference · benchmarking