Alibaba’s Qwen-Audio-3.0-TTS-Plus has taken the top spot in the Speech Arena, a benchmarking platform for text-to-speech (TTS) models, with a score of 1,236, surpassing Simba 3.2 by just 2 points.
A benchmark of speech quality
The Speech Arena is a widely used benchmark for evaluating the quality of TTS models. It assesses the naturalness, coherence, and overall speech quality of generated output, using human evaluators to grade the models against a set of predetermined criteria.
The latest results show Alibaba’s Qwen-Audio-3.0-TTS-Plus narrowly holding onto the top spot, ahead of Simba 3.2, Gemini 3.1 Flash TTS, and Sonic 3.5, which round out the top five.
Tiered offerings for different needs
The Qwen-Audio-3.0-TTS-Plus model comes in two flavors: Flash and Plus. Flash is optimized for real-time use, with latency of around 300 milliseconds, making it suitable for applications where speed is crucial. Plus, on the other hand, prioritizes higher-quality output, focusing on speech naturalness and timbre.
The dual-approach strategy from Alibaba acknowledges that different use cases require different trade-offs between speed and quality, reflecting the growing demand for TTS applications in areas like e-commerce, education, and customer service.
What this means
For developers and businesses, the Speech Arena results offer a clear indicator of a model’s strengths and weaknesses. The dominance of Alibaba’s Qwen-Audio-3.0-TTS-Plus model suggests it’s a reliable choice for applications where high-quality speech is paramount. However, for tasks that demand ultra-low latency, other models may still be more suitable.
The competition in the TTS arena is heating up, with multiple players pushing the boundaries of speech quality and speed. As the field advances, users can expect even more sophisticated and versatile solutions that cater to diverse needs.



