Full flattening of nested data parallelism
Summary
The post explains how Futhark's new flattening transformation converts nested data-parallel constructs into flat GPU kernels, using segmented arrays and careful handling of irregular shapes. It covers control-flow flattening, function lifting, and the challenges of performance and memory usage, including vectorization avoidance. It also references related work and outlines future directions.