Speeding up gearhash on ARM64
Summary
A technical write-up detailing the ARM64 NEON backend optimization for the gearhash Rust crate, achieving roughly 2× speedups. The article walks through the motivation, the challenges of vectorizing a serial hash, and the successive optimizations from latency to throughput, including SIMD striping, load reductions, and combined boundary checks, with benchmark results across different mask densities. It concludes with results, observations on mask density impact, and future exploration of x86 backends.