DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm
DwarfStar is a model-specific native inference engine for running DeepSeek, GLM, and Qwen models on Apple Silicon, CUDA, and ROCm hardware. It includes a CLI, HTTP server, coding agent, vision support, and multi-GPU execution.
Use Cases
Run supported open-weight LLMs locallyServe models through an HTTP APIRun a native tool-using coding agentPerform multi-GPU inferenceRun models with SSD streaming on memory-constrained hardwareRun vision-capable models locally
Built With
- Language
- C
- Frameworks
- Metal · CUDA · ROCm · GGUF · llama.cpp · GGML
Tags
local inference · LLM inference · GPU acceleration · quantization · multi-GPU · tensor parallelism · vision · HTTP server · coding agent · SSD streaming