The Zen of Parallel Programming: The Posture of a Kernel
Summary
The article explores GPU kernel design and parallel execution, focusing on how AMD architectures require different abstractions than NVIDIA. It highlights a specific eight-wave ping-pong schedule that overlaps computation and memory movement, using a philosophical lens to discuss portability and adaptation of abstractions across hardware.