Teaching Automatic1111 to Speak Metal on an M1
Summary
Derek Anderson details a tuned fork of Automatic1111 for Apple Silicon, achieving meaningful speedups on M-series Macs by selectively using Metal Flash Attention, memory-aware routing, and other GPU-focused optimizations. The write-up emphasizes staying with the familiar UI/workflow while integrating native-like performance, highlighting practical lessons about workload-driven optimization and the limits of chasing universal speedups. The project remains open-source and compatible with existing configurations.