mesh-llm
View on GitHubDistributed AI/LLM for the people. Share compute privately or publicly to power your agents and chat.
Rust-based distributed LLM inference engine that pools compute across machines and exposes an OpenAI-compatible API. It routes requests across peers and can split large models into stages for multi-node serving.
Use Cases
Pool GPUs and memory across machines to serve LLMsExpose distributed models through an OpenAI-compatible APIRoute inference requests to peers that can serve a modelSplit models across nodes when they exceed single-machine capacityRun private or public inference meshesProvide local GGUF model serving
Built With
- Language
- Rust
- Frameworks
- Skippy · llama.cpp
Tags
distributed inference · LLM serving · GPU pooling · multi-node · OpenAI-compatible API · model routing · model splitting · GGUF · decentralized · private mesh