Vibe Coding Discover

AI Frameworks

DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm

★ 23K2,237 forksCMITantirez

DwarfStar is a model-specific native inference engine for running DeepSeek, GLM, and Qwen models on Apple Silicon, CUDA, and ROCm hardware. It includes a CLI, HTTP server, coding agent, vision support, and multi-GPU execution.

Use Cases

Run supported open-weight LLMs locallyServe models through an HTTP APIRun a native tool-using coding agentPerform multi-GPU inferenceRun models with SSD streaming on memory-constrained hardwareRun vision-capable models locally

Built With

Language
C
Frameworks
Metal · CUDA · ROCm · GGUF · llama.cpp · GGML

Tags

local inference · LLM inference · GPU acceleration · quantization · multi-GPU · tensor parallelism · vision · HTTP server · coding agent · SSD streaming