opik
View on GitHubDebug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.
Apache-2.0, self-hostable LLM observability and evaluation platform from Comet. Provides deep agent/LLM tracing, datasets and experiments, LLM-as-a-judge metrics, CI/CD eval via PyTest, prompt playground, and production monitoring dashboards.
Use Cases
Trace multi-step agent and tool-call treesLog and inspect LLM calls in dev and prodRun offline experiments and datasets on LLM appsEvaluate RAG quality (answer relevance, context precision)Detect hallucinations and run moderation metricsLLM-as-a-judge scoring with feedback annotationsTest LLM pipelines in CI/CD via PyTestProduction monitoring dashboards for cost/latency/tokensOnline evaluation rules for live trafficPrompt playground for prompt/model experimentationOptimize prompts and agents with Agent OptimizerApply guardrails for safe/responsible AIQuery traces and run evals from coding agents via Opik MCP
Built With
- Language
- Python
- Frameworks
- LangChain · LlamaIndex · OpenAI SDK · Google ADK · AutoGen · Flowise AI · PyTest · MCP · TypeScript SDK · Pydantic
Tags
llm-observability · llm-evaluation · tracing · llmops · agent-tracing · prompt-management · llm-as-a-judge · rag-evaluation · monitoring · self-hosted · datasets · guardrails · dashboards · mcp-server · evaluation · apache-2.0