LMCache
★ 12KLMCache is a vendor-neutral KV cache management layer for LLM inference. It stores and reuses KV cache across CPU RAM, disk and remote backends, cutting TTFT and boosting throughput for long-context, multi-turn agentic and RAG workloads on engines like vLLM.
AI Frameworks | Python · kv-cache · llm-inference
View Project →