Sorting, hashing, and sketches on 370,103 words
Summary
The Stochastic Blog post demonstrates sorting, hashing, and sketching on a real dataset of 370,103 English words. It benchmarks multiple algorithms and sketches, highlighting time/memory tradeoffs and showing that HyperLogLog with 4,096 registers estimates vocabulary size with about 2.71% error.