Vibe Coding Discover

AI Frameworks

LLM inference in C/C++

★ 129K23,544 forksC++MITggml-org

llama.cpp is a dependency-free C/C++ LLM and VLM inference engine built on ggml. It runs quantized GGUF models on CPU, CUDA, Metal, Vulkan and many other backends, and ships CLI, server, web UI and OpenAI-compatible API tools.

Use Cases

Run LLMs and VLMs locally on CPU, GPU, or NPUServe an OpenAI-compatible REST API from a local modelQuantize models to 1.5-8 bit for lower memory useConvert Hugging Face models to GGUF formatHybrid CPU+GPU inference for models larger than VRAMBuild lightweight C/C++ inference into apps and edge devicesRun models directly from Hugging Face with one CLI commandDistributed inference over RPC backend

Built With

Language
C++
Frameworks
ggml · CUDA · Metal · HIP/ROCm · Vulkan · SYCL · OpenVINO · CANN · WebGPU · OpenCL · MUSA · BLAS · PyTorch · Hugging Face Transformers

Tags

llm-inference · gguf · quantization · local-llm · cpu-inference · gpu-inference · cpp · openai-compatible-api · cuda · metal · vulkan · multimodal · edge-inference · no-dependencies · inference-engine · model-conversion