Qwen3.8 27B at 256K: 50 TPS on a 24 GB GPU
Summary
This article documents extensive local-inference benchmarking of Qwen3.8-27B at 256K context on a two-GPU RTX Pro 4000 setup, detailing NVFP4/Q5/Q6 quantization, Hermes iMatrix calibration, and MTP tuning. It concludes that throughput gains come from the interaction of quantization, drafter, CUDA kernels, and workload, with MTP n=8 identified as a sweet spot and full-context performance caveats highlighted.