Pushing the Speed-Cost Frontier for Qwen3-TTS
Summary
Nari Labs reports sub-50 ms p95 TTFA for Qwen3-TTS 1.7B CustomVoice and 10 RPS on a single NVIDIA H100 SXM, detailing a unified scheduler for Talker, Code Predictor, and Codec, plus caching and streaming optimizations. The post benchmarks several engines, describes performance tuning (removing leading silence and chunk sizing) and discusses cost implications of model serving.