Vibe Coding Discover

AI Frameworks

Evaluate and improve models and agents using environments

★ 1.2K360 forksPythonApache-2.0NVIDIA-NeMo

Python library for building environments to evaluate and train LLMs and agents. It provides benchmark environments, agent harnesses, verifiers, rollout collection, and scalable execution for evaluation and reinforcement-learning workflows.

Use Cases

Evaluate agents in stateful task environmentsBenchmark models with reproducible verifiersCollect rollouts for reinforcement learningBuild custom training and evaluation environmentsRun large-scale concurrent agent evaluationsInspect and re-verify rollout results

Built With

Language
Python
Frameworks
vLLM · Ray · LangGraph · NeMo RL · VeRL · Unsloth

Tags

agent-evaluation · LLM-evaluation · reinforcement-learning · benchmarks · training-environments · rollouts · tool-calling · verifiers · agent-training · model-training · scalable-evaluation