LocalAI
View on GitHubLocalAI is the open-source AI engine. Run any model - LLMs, vision, voice, image, video - on any hardware. No GPU required.
LocalAI is a self-hosted, OpenAI/Anthropic-compatible inference engine that runs LLMs, vision, voice, image and video models on CPU or GPU. Backends like llama.cpp, vLLM and whisper.cpp are pulled on demand, and it also ships agents, RAG and MCP support.
Use Cases
run LLMs locally with no GPUdrop-in replacement for OpenAI/Anthropic/ElevenLabs APIsprivate on-prem multi-modal inferenceself-hosted text-to-speech and ASRlocal image and video generationbuild RAG pipelines over private dataorchestrate tool-using AI agentsface recognition and voice biometricsdistributed multi-node inference clusterfine-tune and quantize models in the UIserve models as an MCP-enabled backendChatGPT-style local chat UI for teams
Built With
- Language
- Go
- Frameworks
- llama.cpp · vLLM · whisper.cpp · Stable Diffusion · MLX · Ollama · SGLang · parakeet.cpp · Piper TTS · MCP Go SDK · libp2p · NATS · Echo · Docker · WebRTC
Tags
llm-inference · local-inference · openai-compatible-api · self-hosted · multimodal · cpu-inference · gpu-acceleration · agents · mcp · rag · text-to-speech · speech-recognition · image-generation · video-generation · distributed-inference · docker