kitaru
View on GitHubAgent traces you can run, not just read.
Kitaru records or imports production AI agent traces as sessions, then replays the real agent code against changes in model, prompt, or tools so you see what improved or regressed before shipping. Self-hosted, Apache 2.0, with Python/TypeScript SDKs, framework adapters, MCP server, and agent skills.
Use Cases
Replay production agent runs against a new model, prompt, or code changeImport existing traces from Langfuse, LangSmith, Braintrust, Logfire, or Arize PhoenixBuild a regression suite that gates agent changes in CIRecord live agent sessions with a single wrapper adapterCalibrate LLM evaluators from pinned human judgments on tracesCompare baseline replays vs forked replays to isolate a changeDrive the eval loop from Claude Code, Codex, Cursor, or Windsurf via MCP and skillsSelf-host an agent trace and eval server with no user code executing on it
Built With
- Language
- Python
- Frameworks
- PydanticAI · LangGraph · LangChain · Deep Agents · OpenAI Agents SDK · Claude Agent SDK · Mastra · Vercel AI SDK · FastAPI · Pydantic · Langfuse · LangSmith · Braintrust · Logfire · Arize Phoenix · Docker/Helm
Tags
agent-evals · observability · replay · tracing · llmops · regression-testing · self-hosted · mcp · durable-execution · checkpoints · ci-cd · session-recording · human-in-the-loop · python · typescript · open-source