headroom
View on GitHubCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
Python and TypeScript library, local proxy, and MCP server that compress tool outputs, logs, files, and RAG chunks before they reach an LLM. Includes coding-agent integrations and local retrieval of original content.
Use Cases
Compress tool outputs before sending them to an agentReduce tokens in logs and JSON payloadsCompress code and files in coding-agent workflowsReduce the size of retrieved RAG chunksRun a local proxy to compress LLM requestsRetrieve original content when an agent needs full context
Built With
- Language
- Python
- Frameworks
- MCP · LangChain · FastAPI
Tags
LLM context compression · token optimization · AI agents · MCP server · RAG · JSON compression · code compression · LLM proxy · cross-agent memory