shimmy
★ 5.9KShimmy is a Rust inference server for local GGUF language models, with WebGPU acceleration through Airframe and OpenAI-compatible chat, completion, and streaming APIs. It runs as a single binary without Python or llama.cpp.
AI Frameworks | Rust · LLM inference · local AI
View Project →