Predictive Speculative KV Replication for Bursty LLM Inference
Summary
The article's title suggests a research focus on predictive speculative key-value replication to support bursty LLM inference, aiming to reduce latency during spikes in demand. It likely describes mechanisms for prefetching or replicating cache state and discusses trade-offs, evaluation, and applicability to large-scale inference workloads.