CacheRoute is an innovative LLM scheduling scheme dedicated to enabling flexible KV cache reuse across LLM systems, improving task performance and system efficiency.
-
Updated
Aug 10, 2026 - Python
CacheRoute is an innovative LLM scheduling scheme dedicated to enabling flexible KV cache reuse across LLM systems, improving task performance and system efficiency.
Qwen3.8-27B on RTX 5090s — 262K ctx, 1.5M-token KV pool, ~300 t/s code decode. NVFP4 + vLLM + sm120 patches, reproducible.
Multimodal LLM inference gateway with KV-cache-aware routing and LMCache offload. OpenAI-compatible, benchmarked on GPUs.
Upstream-maintained home for vendor L2 adapter plugins for LMCache
KV-cache-aware LLM inference mesh on a single 6 GB GPU. 2× vLLM + LMCache shared KV + prefix-aware router on k3s, observed with Cilium/Hubble eBPF. Every number measured, every breakage documented.
Upstream-maintained home for vendor device plugins for LMCache
Benchmarking LMCache under simulated RTT
To associate your repository with the lmcache topic, visit your repo's landing page and select "manage topics."