Kimi K3 (2.8T) at 1 token/s on a MacBook Pro, streamed from four SSDs
Summary
Deltafin documents running the Kimi K3 (2.8T) model on consumer hardware (Apple Silicon MacBook Pro) with streaming from four SSDs. The guide covers installation, building, and two model-loading approaches (full on-disk vs streaming), optional Qwen integration, and benchmarks to illustrate performance tradeoffs of local, unpruned large-model inference.