deepeval
View on GitHubThe LLM Evaluation Framework
DeepEval is a Python framework for testing and evaluating LLM applications, agents, and RAG pipelines. It provides ready-made and custom metrics, synthetic test data generation, benchmarking, and pytest-based workflows for local or CI evaluation.
Use Cases
Test LLM applications in CIEvaluate agent task completion and tool useMeasure RAG answer quality and faithfulnessDetect hallucinations, bias, and toxicityGenerate synthetic evaluation datasetsBenchmark language modelsCompare prompts and model iterationsEvaluate MCP-based agents
Built With
- Language
- Python
- Frameworks
- LangChain · LangGraph · Pydantic AI · CrewAI · OpenAI Agents
Tags
LLM evaluation · evaluation metrics · LLM-as-a-judge · agent evaluation · RAG evaluation · testing · synthetic data · benchmarking · prompt optimization · MCP evaluation