DigiNews

Tech Watch by Johan Denoyer

← Back to articles

Pushing the Speed-Cost Frontier for Qwen3-TTS

Quality: 8/10 Relevance: 9/10

Summary

Nari Labs reports sub-50 ms p95 TTFA for Qwen3-TTS 1.7B CustomVoice and 10 RPS on a single NVIDIA H100 SXM, detailing a unified scheduler for Talker, Code Predictor, and Codec, plus caching and streaming optimizations. The post benchmarks several engines, describes performance tuning (removing leading silence and chunk sizing) and discusses cost implications of model serving.

🚀 Service construit par Johan Denoyer