Vibe Coding Discover

AI Tools

The LLM Evaluation Framework

★ 19K2,000 forksPythonApache-2.0confident-ai

DeepEval is a Python framework for testing and evaluating LLM applications, agents, and RAG pipelines. It provides ready-made and custom metrics, synthetic test data generation, benchmarking, and pytest-based workflows for local or CI evaluation.

Use Cases

Test LLM applications in CIEvaluate agent task completion and tool useMeasure RAG answer quality and faithfulnessDetect hallucinations, bias, and toxicityGenerate synthetic evaluation datasetsBenchmark language modelsCompare prompts and model iterationsEvaluate MCP-based agents

Built With

Language
Python
Frameworks
LangChain · LangGraph · Pydantic AI · CrewAI · OpenAI Agents

Tags

LLM evaluation · evaluation metrics · LLM-as-a-judge · agent evaluation · RAG evaluation · testing · synthetic data · benchmarking · prompt optimization · MCP evaluation