42x faster prompt lookup drafting in llama.cpp
Summary
The article describes a series of performance optimizations for llama.cpp aimed at accelerating prompt lookup decoding. It reports significant speedups (up to 42x initially, with further improvements adding up to 140x when combined with parallel PR work) and memory reductions through architectural changes to n-gram caches, cache representations, and a precheck optimization. Benchmarking across WikiText-103 and multiple corpus sizes demonstrates the practical impact on drafting latency and static cache loading.