Project
Tiered KV-cache for Ollama: GPU VRAM → Host RAM → Disk, enabling near-infinite context length for LLM inference on consumer GPUs
Cuda 42% Go 30% C 26% CMake 2% Makefile 1%
Write-up pending. This page is built from repository metadata alone. Adding docs/showcase.md to the repo fills in the sections below on the next sync.
Generated from databloom/ollama-kv-cache-tiering · synced 2026-08-19 00:13:54 UTC