Strata uses GPU-assisted I/O and cache-aware scheduling to cut the cost of loading cached KV states, raising long-context serving throughput by up to 5x at equal latency.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.DC 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Strata: Hierarchical Context Caching for Long Context Language Model Serving
Strata uses GPU-assisted I/O and cache-aware scheduling to cut the cost of loading cached KV states, raising long-context serving throughput by up to 5x at equal latency.