AMA-Bench
★ 82AMA-Bench is an ICML 2026 evaluation framework for agentic memory: methods build memory from long agent trajectories, retrieve evidence, and answer QA scored by LLM-as-judge. Includes vLLM/API pipelines, cross-judge validation, and a HF leaderboard.
AI Agents | Python · agent-memory · long-horizon
View Project →