U2-ASR

Beyond hearing, truly understanding expression

100+ dialects to many languages · one-shot · real-time · long audio

U2-ASR

U2-ASR

Built on full-scenario capabilities, U2-ASR combines contextual understanding and custom vocabulary optimization to move beyond rigid word-by-word recognition into more natural semantic understanding. It covers 100+ dialects and multiple languages while balancing accuracy with clean, normalized text output. One-shot recognition and real-time transcription support instant, low-latency interactions—ready for personal, study, office, and everyday use.

90%

Noisy-Environment Accuracy

99.2%

AISHELL-1 Benchmark

98.4%

AISHELL-3 Benchmark

98.4%

Libri Clean Benchmark

Outstanding performance on the industrial dialect benchmark, comprehensively surpassing mainstream ASR models

U2-ASR Other Models

Xiang

90.1
U2-ASR
86.2
FunASR 1.5
83.4
Qwen3-ASR-1.7B
80.4
Seed-ASR

Hakka

83.4
U2-ASR
65
FunASR 1.5
42.6
Qwen3-ASR-1.7B
71.8
Seed-ASR

Min

84.9
U2-ASR
68.2
FunASR 1.5
73.1
Qwen3-ASR-1.7B
71.5
Seed-ASR

Gan

88.3
U2-ASR
64.6
FunASR 1.5
64
Qwen3-ASR-1.7B
74.1
Seed-ASR

Jin

89.6
U2-ASR
73.1
FunASR 1.5
73
Qwen3-ASR-1.7B
77.4
Seed-ASR

Wu (Suzhou)

82.2
U2-ASR
62
FunASR 1.5
53
Qwen3-ASR-1.7B
50.7
Seed-ASR

SW (Wuhan)

92.1
U2-ASR
89.2
FunASR 1.5
85.8
Qwen3-ASR-1.7B
85.6
Seed-ASR

Ji-Lu

96.2
U2-ASR
93
FunASR 1.5
93.7
Qwen3-ASR-1.7B
93
Seed-ASR

SW (Sichuan)

94.7
U2-ASR
89.2
FunASR 1.5
90.2
Qwen3-ASR-1.7B
90.1
Seed-ASR

Wu (Shanghai)

89
U2-ASR
74.6
FunASR 1.5
72.5
Qwen3-ASR-1.7B
79.8
Seed-ASR

Yue

93
U2-ASR
90.2
FunASR 1.5
89.9
Qwen3-ASR-1.7B
87.7
Seed-ASR

WS-Meeting

95.8
U2-ASR
95.2
Qwen3-ASR-1.7B
94.1
FunASR 1.5
93.7
Seed-ASR

Libri Clean

98.4
U2-ASR
97.9
Qwen3-ASR-1.7B
98.3
FunASR 1.5
98.4
Seed-ASR

Libri Other

96.7
U2-ASR
96.1
Qwen3-ASR-1.7B
96.6
FunASR 1.5
97.2
Seed-ASR

WS-Net

95.9
U2-ASR
95.3
Qwen3-ASR-1.7B
94.3
FunASR 1.5
95.3
Seed-ASR

Aishell-1

99.2
U2-ASR
98.6
Qwen3-ASR-1.7B
97.8
FunASR 1.5
98.8
Seed-ASR

Aishell-2

97.9
U2-ASR
97.5
Qwen3-ASR-1.7B
97.5
FunASR 1.5
97.2
Seed-ASR

Aishell-3

98.4
U2-ASR
98
Qwen3-ASR-1.7B
95.9
FunASR 1.5
98.4
Seed-ASR

Fluers(zh)

96.8
U2-ASR
96.8
Qwen3-ASR-1.7B
97.6
FunASR 1.5
97
Seed-ASR

Arabic

94.72
U2-ASR
89.23
FunASR 1.5
93.79
Seed-ASR
90.54
whisper-large-v3-turbo
74.06
cohere-transcribe-03-2026
91.31
Qwen3-ASR-1.7B

German

95.08
U2-ASR
88.88
FunASR 1.5
92.01
Seed-ASR
91.97
whisper-large-v3-turbo
91.94
cohere-transcribe-03-2026
93.21
Qwen3-ASR-1.7B

Spanish

95.4
U2-ASR
90.26
FunASR 1.5
93.61
Seed-ASR
91.9
whisper-large-v3-turbo
92.74
cohere-transcribe-03-2026
93.87
Qwen3-ASR-1.7B

French

94.31
U2-ASR
85.26
FunASR 1.5
90.79
Seed-ASR
91.22
whisper-large-v3-turbo
91.44
cohere-transcribe-03-2026
92.1
Qwen3-ASR-1.7B

Indonesian

96.59
U2-ASR
89.15
FunASR 1.5
94.08
Seed-ASR
89.91
whisper-large-v3-turbo
94.56
Qwen3-ASR-1.7B

Italian

95.19
U2-ASR
87.51
FunASR 1.5
90.67
Seed-ASR
91.91
whisper-large-v3-turbo
89.1
cohere-transcribe-03-2026
92.79
Qwen3-ASR-1.7B

Vietnamese

96.99
U2-ASR
89.51
FunASR 1.5
93.93
Seed-ASR
79.9
whisper-large-v3-turbo
93.49
cohere-transcribe-03-2026
93.54
Qwen3-ASR-1.7B

Russian

94.74
U2-ASR
88.55
FunASR 1.5
88.61
Seed-ASR
90.42
whisper-large-v3-turbo
91.07
Qwen3-ASR-1.7B

Portuguese

93.15
U2-ASR
85.87
FunASR 1.5
90.68
Seed-ASR
87.35
whisper-large-v3-turbo
87.78
cohere-transcribe-03-2026
90.54
Qwen3-ASR-1.7B

Korean

94.27
U2-ASR
87.22
FunASR 1.5
93.04
Seed-ASR
91.2
whisper-large-v3-turbo
80.68
cohere-transcribe-03-2026
93.3
Qwen3-ASR-1.7B

Core Capabilities

Full dialect and bilingual coverage, understanding every accent

Supports Chinese, English and 15 international languages including Arabic, German, Spanish, French, Indonesian, Japanese, Korean, Portuguese, Russian, Turkish, Vietnamese, Thai and Italian, while covering hundreds of dialects and regional accents such as Cantonese, Sichuanese, Shanghainese and Hokkien, spanning all seven major Chinese dialect families.

Fast single-sentence recognition, speech to text in seconds

Low-latency, high-efficiency response delivers text from a single utterance almost instantly, suited to voice commands, quick Q&A and other real-time interaction scenarios.

Context-aware understanding — beyond dictation to real meaning

Moves past mechanical word-by-word transcription, using conversational context and domain vocabularies to deeply understand meaning — robust against noisy environments, heavy accents and specialist jargon.

Full-scenario transcription with structured output for any audio length

Supports real-time streaming transcription and batch processing of long audio up to 5 hours, with speaker diarization, timestamps and smart punctuation producing clean, ready-to-use text.

Highlights

Strong noise resistance

Optimized for real recording environments, maintaining high recognition stability amid complex background noise, busy shopping malls and live meeting settings.

Broad dialect and multilingual coverage

Beyond Mandarin, the model supports over a hundred dialects and 15 foreign languages across Asia, Europe, the Middle East and Latin America — meeting cross-region and cross-border transcription needs through a single unified model.

Stronger domain semantic understanding

Context and hotword injection enable enhanced recognition of specialist terminology in healthcare, automotive, customer service and other professional domains.

Precise transcription with speaker separation

Supports speaker diarization, smart sentence segmentation, punctuation prediction and timestamp output for clear, organized transcripts.

Use Cases

Office Document Input

Generate work docs, emails, and draft plans quickly from speech to speed up content input.

Medical Record Entry

Recognize large volumes of medical terminology in real time and support physician dictation for faster EMR creation.

Communication and Translation

Accurately restores and clearly records different accents and ways of expression.

Meeting Audio Transcription

Parse long recordings in one click and output structured, well-formatted transcripts with less manual effort.

Capabilities

  • Supports audio-to-text for long-form scenarios such as meetings, lectures, customer service, and business recordings.

    Supports speaker diarization to distinguish segments from different speakers.

  • Supports smart sentence splitting and punctuation prediction to improve readability.

    Supports timestamp output for subtitles, search, and audio-video alignment.

  • Supports context and hotword enhancement to improve recognition of proper nouns and industry terms.

    Supports one-shot recognition with millisecond response for instant interaction.

  • Supports real-time streaming transcription with low-latency sync as you speak.

    Supports speaker diarization, distinguishing segments from different speakers.

  • Supports smart sentence segmentation and punctuation prediction for better readability.
  • Returns timestamp information, directly usable for subtitles, search and audio-video alignment.
  • Supports context and hotword enhancement, improving recognition of proper nouns and industry terms.
  • Supports fast single-sentence recognition with millisecond-level response for real-time interaction.
  • Supports real-time streaming transcription, generating text synchronously with low latency as speech happens.

Flexible pricing, custom solutions, private deployment

Flexible billing models and dedicated customization for speech recognition scenarios, with private deployment to ensure data security and compliance

Talk to an Expert