DigiNews

Tech Watch by Johan Denoyer

← Back to articles

vLLM Kunlun: Open-source plugin to run vLLM on Kunlun XPU

Quality: 8/10 Relevance: 9/10

Summary

vLLM Kunlun is a community-maintained hardware plugin that lets you run vLLM on the Kunlun XPU. It emphasizes seamless integration, broad model support (Qwen, Llama, Gemma4, DeepSeek, etc.), and various quantization and optimization features (W8A8, AWQ, GPTQ, LoRA, FlashMLA, MTP). The repo provides prerequisites, architecture, and a Quick Start to launch an OpenAI-compatible API server.

🚀 Service construit par Johan Denoyer