shimmy
View on GitHub⚡ Pure-Rust WebGPU inference engine — OpenAI-API compatible, GGUF native, runs on any GPU. No Python. No llama.cpp. Single binary.
Shimmy is a Rust inference server for local GGUF language models, with WebGPU acceleration through Airframe and OpenAI-compatible chat, completion, and streaming APIs. It runs as a single binary without Python or llama.cpp.
Use Cases
Run local GGUF language modelsServe chat completions through an OpenAI-compatible APIAdd local LLM inference to existing SDK-based applicationsRun models with WebGPU acceleration across supported GPUsStream text-generation responses locally
Built With
- Language
- Rust
- Frameworks
- Airframe · WebGPU · WGSL · Axum · Tokio
Tags
LLM inference · local AI · inference server · WebGPU · GGUF · OpenAI-compatible API · Rust · GPU acceleration · streaming