DeepSeek V4 Flash on a Single AMD MI300X
Summary
DeepSeek V4 Flash on AMD MI300X demonstrates a production-ready ROCm-based vLLM stack using a single MI300X. The repo provides patches, overlays, and tuning tables for FP8 KV, DSpark, and a 2,048-token scheduler with a 1,024-token long-prefill cap, and reports concrete benchmarks. This is valuable for SMBs evaluating on-prem AI inference on affordable AMD hardware.