Project

ollama-kv-cache-tiering

Tiered KV-cache for Ollama: GPU VRAM → Host RAM → Disk, enabling near-infinite context length for LLM inference on consumer GPUs

Maintenance Repository →

Cuda 42% Go 30% C 26% CMake 2% Makefile 1%

Write-up pending. This page is built from repository metadata alone. Adding docs/showcase.md to the repo fills in the sections below on the next sync.

Generated from databloom/ollama-kv-cache-tiering · synced 2026-08-19 00:13:54 UTC