This week, Alibaba’s voice large model series has once again achieved outstanding results on the globally authoritative AI evaluation platform Artificial Analysis:

Fun-Realtime-ASR ranks first worldwide with a word error rate of only 1.8%;

Fun-Realtime-AudioChat has taken the top spot in both voice reasoning and conversational fluency;

Fun-Realtime-TTS-Preview maintains its position as China’s No. 1 speech synthesis technology, scoring 1190 Elo points.

Thus, in the three core areas of ASR (speech recognition), TTS (speech synthesis), and Chat (voice conversation), Alibaba’s voice large models have achieved an overall lead, securing a domestic “grand slam” achievement.

Voice intelligence deeply integrated into DingTalk and Wukong product experiences

As an important carrier for the practical application of Alibaba’s large model technologies, voice capabilities have been transformed into tangible user experiences through DingTalk and Wukong. Whether participating in DingTalk meetings, using the AI transcription feature, handling voice messages in instant communications, or pairing with hardware devices like the DingTalk A1 and Cleer smart headphones, users can enjoy a complete intelligent experience—spanning from online conferences to face-to-face discussions, and extending from text-based office work to voice interactions—easily overcoming language barriers posed by regional dialects and professional jargon.

The advanced models recently certified internationally will gradually be integrated into these scenarios, continuously enhancing productivity and communication quality.

Three key capabilities for voice interaction: hear clearly, speak well, converse naturally

A mature voice system must possess three core abilities: accurately recognizing spoken content (ASR), producing clear and natural audio output (TTS), and engaging in real-time conversations with comprehension and response capabilities (Chat). These three elements complement each other, forming seamless human-machine voice interactions.

As early as last year, DingTalk, together with the Tongyi Lab, launched the Fun-ASR large model, which has been widely adopted across multiple products, laying the foundation for voice intelligence.

Listen clearly: handle complex meetings and specialized expressions with ease

In DingTalk meetings, even when multiple speakers take turns, interruptions occur frequently, or industry-specific terms are abundant, Fun-ASR still captures every utterance with precision. After the meeting, it automatically generates near-zero-error transcripts, eliminating the need for manual organization.

Paired with the DingTalk A1 series of AI hardware and the AI transcription feature, whether conducting interviews, delivering lectures, or leading team brainstorming sessions, a single tap initiates high-quality, well-structured text recording. Combined with voice reasoning technology, the AI not only “hears clearly” but also “understands meaning,” automatically summarizing key points and flagging action items, greatly boosting information processing efficiency.

Real-time voice-to-text, fully supporting Cantonese and major dialects

In instant messaging, voice messages up to 60 seconds long can be instantly converted into text with exceptional accuracy, fully supporting various local dialects, accents, and industry-specific terms.

China’s linguistic diversity presents significant challenges for speech recognition. Leveraging Alibaba’s voice large model technology, DingTalk and Wukong now cover eight major dialect regions—Mandarin, Wu, Cantonese, Min, Hakka, Gan, Xiang, and Jin—and boast leading nationwide recognition capabilities for dozens of urban dialects, with mainstream dialect recognition accuracy exceeding 90%.

Whether you’re a finance professional communicating in Shanghainese, a local business owner negotiating deals in Cantonese, or a frontline worker speaking Sichuanese, Henanese, or Shaanxi dialects, DingTalk can accurately “understand” your expression.

According to the latest rankings from Artificial Analysis, the new-generation ASR model achieves a word error rate as low as 1.8%, with future enhancements aimed at further improving the precision of live captions and meeting transcriptions.

Speak well, converse naturally: smooth interactions like talking to a real person

A superior voice experience goes beyond simply “hearing clearly”; it also requires “speaking well” and “conversing smoothly.” On devices such as the Cleer smart headphones, powerful speech recognition and synthesis technologies enable users to receive messages, dictate replies, or issue commands while commuting, all with a more natural and fluid interaction.

Fun-Realtime-TTS-Preview has claimed the top spot among domestic speech synthesis technologies with an Elo score of 1190, making AI voices sound more lifelike and emotionally expressive, significantly elevating the quality of voice output across DingTalk and Wukong products.

Fun-Realtime-AudioChat has secured two global top honors, achieving 97.6% in voice reasoning and 97.8% in conversational fluency. This model not only comprehends semantic logic but also adeptly handles interruptions, follow-up questions, topic shifts, and other real-world dialogue scenarios, delivering interactions nearly indistinguishable from human conversation.

Deep customization of industry-specific language models—truly “understand jargon, grasp context”

General-purpose voice models often struggle with specialized terms like “LPR,” “put-call parity,” “Gleevec,” or “civil retrial,” leading to misinterpretations in real-world applications.

DingTalk, in collaboration with Alibaba’s voice team, has trained industry-specific language models tailored to major sectors—including Internet, artificial intelligence, finance, healthcare, law, automotive, manufacturing, education, and government—ensuring full vocabulary coverage and establishing dedicated domain-specific models.

In the financial sector, the model accurately recognizes terms like “reverse repurchase,” “MLF,” and “asset management regulations,” significantly improving the quality of meeting minutes and compliance documents.

In healthcare settings, proper nouns such as “methotrexate,” “EGFR-targeted therapy,” and “MDT multidisciplinary consultation” no longer pose recognition challenges.

In the legal field, everything from “unjust enrichment” to “bona fide acquisition” is faithfully reproduced in court statements and attorney interviews.

In manufacturing, technical terms like “tolerance fit,” “CNC machining,” and “PPAP” are recognized with remarkable accuracy, greatly accelerating production meeting processes.

This level of industry customization goes beyond simply adding words to a dictionary; it involves deep tuning at the model’s foundational level, enabling the system to truly master industry-specific contexts, distinguish homophones, decode abbreviations, and interpret colloquialisms, ensuring that every voice interaction is conducted with “industry jargon, understood intent.”

Voice is the most inherently human form of communication. From “hearing clearly” to “understanding meaning,” and from “being able to speak” to “capable of conversing,” Alibaba’s voice large model “grand slam” represents not only a technological breakthrough but also a comprehensive upgrade to product experiences. DingTalk and Wukong will continue to transform world-class voice AI technologies into everyday tools—making meetings more efficient, record-keeping easier, and communication unburdened by language barriers, truly enabling seamless collaboration across dialects and professions.

DomTech is DingTalk’s official designated service provider in Macau, specializing in providing DingTalk services to a wide range of customers. If you’d like to learn more about DingTalk platform applications, feel free to contact our online customer service, call +852 95970612, or email us at [email protected]. With an excellent development and operations team and extensive market service experience, we can offer you professional DingTalk solutions and services!

立即提升團隊協作效率

免費試用釘釘,改變你的工作方式。

免費開始