llama.cpp
★ 129Kllama.cpp is a dependency-free C/C++ LLM and VLM inference engine built on ggml. It runs quantized GGUF models on CPU, CUDA, Metal, Vulkan and many other backends, and ships CLI, server, web UI and OpenAI-compatible API tools.
AI Frameworks | C++ · llm-inference · gguf
View Project →