U2-ASR
Beyond hearing, truly understanding expression
100+ dialects to many languages · one-shot · real-time · long audio
U2-ASR
Built on full-scenario capabilities, U2-ASR combines contextual understanding and custom vocabulary optimization to move beyond rigid word-by-word recognition into more natural semantic understanding. It covers 100+ dialects and multiple languages while balancing accuracy with clean, normalized text output. One-shot recognition and real-time transcription support instant, low-latency interactions—ready for personal, study, office, and everyday use.
Noisy-Environment Accuracy
AISHELL-1 Benchmark
AISHELL-3 Benchmark
Libri Clean Benchmark
Xiang
Hakka
Min
Gan
Jin
Wu (Suzhou)
SW (Wuhan)
Ji-Lu
SW (Sichuan)
Wu (Shanghai)
Yue
WS-Meeting
Libri Clean
Libri Other
WS-Net
Aishell-1
Aishell-2
Aishell-3
Fluers(zh)
Arabic
German
Spanish
French
Indonesian
Italian
Vietnamese
Russian
Portuguese
Korean
Core Capabilities
Full dialect and bilingual coverage, understanding every accent
Supports Chinese, English and 15 international languages including Arabic, German, Spanish, French, Indonesian, Japanese, Korean, Portuguese, Russian, Turkish, Vietnamese, Thai and Italian, while covering hundreds of dialects and regional accents such as Cantonese, Sichuanese, Shanghainese and Hokkien, spanning all seven major Chinese dialect families.
Fast single-sentence recognition, speech to text in seconds
Low-latency, high-efficiency response delivers text from a single utterance almost instantly, suited to voice commands, quick Q&A and other real-time interaction scenarios.
Context-aware understanding — beyond dictation to real meaning
Moves past mechanical word-by-word transcription, using conversational context and domain vocabularies to deeply understand meaning — robust against noisy environments, heavy accents and specialist jargon.
Full-scenario transcription with structured output for any audio length
Supports real-time streaming transcription and batch processing of long audio up to 5 hours, with speaker diarization, timestamps and smart punctuation producing clean, ready-to-use text.
Highlights
Strong noise resistance
Optimized for real recording environments, maintaining high recognition stability amid complex background noise, busy shopping malls and live meeting settings.
Broad dialect and multilingual coverage
Beyond Mandarin, the model supports over a hundred dialects and 15 foreign languages across Asia, Europe, the Middle East and Latin America — meeting cross-region and cross-border transcription needs through a single unified model.
Stronger domain semantic understanding
Context and hotword injection enable enhanced recognition of specialist terminology in healthcare, automotive, customer service and other professional domains.
Precise transcription with speaker separation
Supports speaker diarization, smart sentence segmentation, punctuation prediction and timestamp output for clear, organized transcripts.
Use Cases
Office Document Input
Generate work docs, emails, and draft plans quickly from speech to speed up content input.
Medical Record Entry
Recognize large volumes of medical terminology in real time and support physician dictation for faster EMR creation.
Communication and Translation
Accurately restores and clearly records different accents and ways of expression.
Meeting Audio Transcription
Parse long recordings in one click and output structured, well-formatted transcripts with less manual effort.
Capabilities
-
Supports audio-to-text for long-form scenarios such as meetings, lectures, customer service, and business recordings.
Supports speaker diarization to distinguish segments from different speakers.
-
Supports smart sentence splitting and punctuation prediction to improve readability.
Supports timestamp output for subtitles, search, and audio-video alignment.
-
Supports context and hotword enhancement to improve recognition of proper nouns and industry terms.
Supports one-shot recognition with millisecond response for instant interaction.
-
Supports real-time streaming transcription with low-latency sync as you speak.
Supports speaker diarization, distinguishing segments from different speakers.
- Supports smart sentence segmentation and punctuation prediction for better readability.
- Returns timestamp information, directly usable for subtitles, search and audio-video alignment.
- Supports context and hotword enhancement, improving recognition of proper nouns and industry terms.
- Supports fast single-sentence recognition with millisecond-level response for real-time interaction.
- Supports real-time streaming transcription, generating text synchronously with low latency as speech happens.
Flexible pricing, custom solutions, private deployment
Flexible billing models and dedicated customization for speech recognition scenarios, with private deployment to ensure data security and compliance
Talk to an Expert