inference-gateway
View on GitHubAn open-source, cloud-native, high-performance gateway unifying multiple LLM providers, from local solutions like Ollama to major cloud providers such as OpenAI, Groq, Cohere, Anthropic, Cloudflare and DeepSeek.
A Go, cloud-native gateway that unifies many LLM providers (OpenAI, Anthropic, Groq, Ollama, DeepSeek, and more) behind one OpenAI-compatible API. Adds MCP tool discovery, streaming, vision, OIDC auth, and OpenTelemetry/Prometheus metrics for self-hosted LLM traffic.
Use Cases
Route requests to many LLM providers through one OpenAI-compatible endpointSelf-host a privacy-preserving LLM proxyExpose MCP server tools to LLMs automaticallySwitch or fail over between local and cloud models by model nameCollect Prometheus/OTLP metrics and traces for LLM trafficAdd OIDC auth, TLS and timeouts in front of LLM APIsStream tokens from any providerProxy image generation and vision/multimodal requestsDeploy horizontally scaled LLM gateway on KubernetesDynamically toggle MCP middleware per request or via env vars
Built With
- Language
- Go
- Frameworks
- Gin · OpenTelemetry · mcp-golang · Open Policy Agent · Prometheus · Kubernetes · oapi-codegen · Zap · OIDC · Docker
Tags
llm-gateway · proxy · openai-compatible · multi-provider · mcp · observability · kubernetes · self-hosted · go · opentelemetry · streaming · function-calling · multimodal · oidc · docker · provider-routing