ragas
View on GitHubSupercharge Your LLM Application Evaluations 🚀
Python toolkit for evaluating RAG systems and other LLM applications with built-in and custom metrics. It can also generate test datasets and integrates with popular LLM frameworks and observability tools.
Use Cases
Evaluate RAG answer qualityScore LLM application outputsGenerate test datasetsTrack production feedback for AI applicationsCompare prompts and workflows
Built With
- Language
- Python
- Frameworks
- LangChain · LlamaIndex · Haystack · DSPy
Tags
LLM evaluation · RAG evaluation · LLMOps · evaluation metrics · test data generation · AI application testing