kvcache-ai/ktransformers
Summary
KTransformers is an open-source research project focused on efficient inference and fine-tuning of large language models using CPU-GPU heterogeneous computing. The repository highlights two user-facing capabilities: high-performance kt-kernel inference and SFT (fine-tuning with LLaMA-Factory), along with ongoing updates and MoE optimization. It emphasizes quantization, cross-framework integration, and community contribution.