Bringing PyTorch Monarch to AMD GPUs: Single-Controller Distributed Training on ROCm
Summary
The article details bringing PyTorch Monarch to AMD Instinct GPUs with ROCm, enabling single-controller, fault-tolerant distributed training across AMD hardware. It covers the architecture, porting work (HIP, RCCL, RDMA), and a fault-tolerant training case study on SLURM and Kubernetes, highlighting performance and ecosystem readiness.