Vibe Coding Discover

AI Tools

awesome-evals

View on GitHub

A curated, non-BS library of the best resources for building and evaluating AI agents — papers, blogs, talks, tools, benchmarks. Maintained by BenchFlow.

★ 923112 forksCustombenchflow-ai

An annotated collection of papers, articles, talks, tools, and benchmarks for building and evaluating AI agents. Includes a practical playbook with runnable examples for graders, trajectory evaluation, error analysis, and CI gating.

Use Cases

Find papers and practical guides on evaluating AI agentsDesign LLM-as-judge evaluationsBuild benchmarks and evaluation datasetsAssess agent tool use and multi-turn behaviorAdd evaluation checks to CIAudit benchmark quality and safety

Tags

AI agent evaluation · LLM evaluation · benchmarks · evals · research resources · evaluation playbook