Vibe Coding Discover

AI Agents

[ICML 26] An evaluation framework assessing long-context retention and long-horizon memory performance for agentic applications (AMA-bench).

★ 8212 forksPythonMITAMA-Bench

AMA-Bench is an ICML 2026 evaluation framework for agentic memory: methods build memory from long agent trajectories, retrieve evidence, and answer QA scored by LLM-as-judge. Includes vLLM/API pipelines, cross-judge validation, and a HF leaderboard.

Use Cases

Benchmarking long-horizon memory in agentic applicationsEvaluating memory construction and retrieval from agent trajectoriesComparing memory methods (BM25, embedding, AMA-Agent) on open-ended QALLM-as-judge scoring with multi-model cross-validation and agreement statsTesting long-context retention of LLMsRunning local vLLM vs remote API evaluation pipelinesEvaluating tool-using coding agents (Codex, Claude Code) on trajectory QASubmitting results to a Hugging Face leaderboard

Built With

Language
Python
Frameworks
vLLM · PyTorch · Hugging Face Transformers · FAISS · Ray · Starlette · Gymnasium · MiniGrid · TextWorld · ALFWorld · Hugging Face Datasets · tiktoken

Tags

agent-memory · long-horizon · long-context · benchmark · evaluation · llm-as-judge · memory-retrieval · embeddings · bm25 · vllm · leaderboard · agent-trajectories · icml · python