phoenix
View on GitHubAI Observability & Evaluation
Arize Phoenix is an open-source, OpenTelemetry-based AI observability platform for tracing, evaluating, and debugging LLM and agent applications, with datasets, experiments, prompt management, and a built-in MCP server. Self-hostable via pip or Docker.
Use Cases
Trace LLM and agent application runtime with OpenTelemetryRun LLM-as-judge response and retrieval evaluationsBenchmark prompts and compare models in the playgroundCreate versioned datasets for experiments and fine-tuningTrack prompt/LLM/retrieval changes across experimentsDebug agent traces and iterate prompts with PXI agentSelf-host an AI observability platform via pip or DockerConnect MCP clients like Cursor or Claude Code to query tracesMonitor RAG pipelines and detect hallucinationsInstrument apps automatically for OpenAI, Anthropic, Bedrock
Built With
- Language
- Python
- Frameworks
- LangChain · LangGraph · LlamaIndex · CrewAI · DSPy · OpenAI Agents SDK · Claude Agent SDK · Vercel AI SDK · Mastra · Pydantic AI · Google ADK · FastAPI · Strawberry GraphQL · OpenTelemetry
Tags
observability · llm-evaluation · tracing · opentelemetry · llmops · evals · datasets · prompt-management · experiments · playground · ai-monitoring · mcp-server · self-hosted · openinference · debugging · agents