ZML: Model to Metal
Summary
ZML introduces a production inference stack that decouples AI workloads from proprietary hardware, enabling models to run efficiently across NVIDIA, AMD, TPU, and Trainium with a single codebase. It emphasizes explicitness, composability, and predictability, avoids Python runtimes and abstraction overhead, and targets peak hardware performance. The article also notes the v2 release and points to GitHub and docs, highlighting open-source, self-hosted deployment options.