vLLM Kunlun: Open-source plugin to run vLLM on Kunlun XPU
Summary
vLLM Kunlun is a community-maintained hardware plugin that lets you run vLLM on the Kunlun XPU. It emphasizes seamless integration, broad model support (Qwen, Llama, Gemma4, DeepSeek, etc.), and various quantization and optimization features (W8A8, AWQ, GPTQ, LoRA, FlashMLA, MTP). The repo provides prerequisites, architecture, and a Quick Start to launch an OpenAI-compatible API server.