Benchmarking pocket-scale inference
Summary
The article reports on benchmarking pocket-scale inference on mobile devices using llama.cpp with quantization to 4-bit or smaller. It compares end-to-end generation time and peak memory across devices like iPhone 17 Pro and Galaxy S26 Ultra, focusing on models that fit within 8 GB of memory and real-device measurements. It notes methodology references and plans to expand coverage to more models, quantizations, and devices.