Unisound-USTC Joint Team Shines Again with Two Wins and a Third-Place Finish at the 10th ABAW International Challenge at CVPR 2026
The IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2026) recently announced the final results of the 10th Workshop and Competition on Affective Behavior Analysis in-the-Wild (10th ABAW). The joint team formed by Unisound and the group led by Associate Professor Jun Yu at the University of Science and Technology of China (USTC) stood out among teams from around the world, taking first place in both the Emotional Mimicry Intensity Estimation (EMI) and Action Unit Detection (AU) tracks and third place in the Expression Recognition (EXPR) track. Building on the technical innovations developed for the competition, the joint team also completed and publicly released three papers at the CVPR 2026 Workshop, highlighting its leading research capabilities in affective behavior analysis in the wild and multimodal understanding.

Jointly organized by IEEE and CVF, CVPR is one of the world's most influential top-tier academic conferences in computer vision and, together with ICCV and ECCV, is widely regarded as one of the field's three premier conferences. ABAW, a long-running CVPR workshop and international competition, focuses on human affective behavior analysis in the wild and has become one of the most influential competitions in affective computing worldwide.
Against strong competition from universities, research institutions, and industry teams worldwide, the Unisound-USTC joint team achieved outstanding results across multiple tracks through solid algorithmic innovation and engineering expertise.
Human affective behavior analysis integrates multimodal information such as vision, speech, and text to automatically understand and analyze human emotional states, behavioral patterns, and social interactions. It is a key research direction as AI advances toward more natural human-computer interaction. With the rise of multimodal large models and generative AI, affective behavior analysis is showing broad application potential in intelligent assistants, digital humans, mental health assessment, smart education, and other fields.
The 10th ABAW competition covered several challenging tasks, including Expression Recognition (EXPR), Action Unit Detection (AU), Emotional Mimicry Intensity Estimation (EMI), Violence Detection (VD), and Ambiguous Hesitation Recognition (AH). The Unisound-USTC joint team conducted in-depth research into key issues including multimodal fusion, long-range temporal modeling, and robust learning under missing modalities, producing a series of innovative results.
I. EMI Track: Robust Text-Anchored Multimodal Emotional Mimicry Intensity Estimation
Paper:
To address the vulnerability of visual and audio signals to noise and the frequent occurrence of missing modalities in Emotional Mimicry Intensity Estimation (EMI), the research team proposed the TAEMI (Text-Anchored Emotional Mimicry Intensity Estimation) framework.
The method innovatively uses text as a stable semantic anchor and introduces a Text-Anchored Dual Cross-Attention mechanism, allowing textual semantics to actively guide the alignment and fusion of visual and audio modalities. It also employs a Missing-Modality Token and Modality Dropout to strengthen robustness when modalities are missing, enabling stable prediction of continuous changes in emotional intensity in complex real-world environments.
Experimental results show that the method significantly outperforms the official baseline on the Hume-Vidmimic2 dataset, effectively improving the accuracy and stability of emotional mimicry intensity estimation.
II. AU Track: Multimodal Action Unit Detection Through Hierarchical Granularity Alignment and State-Space Modeling

Paper:
To address large variations in facial pose, long temporal dependencies, and complex audio-visual relationships in Action Unit (AU) detection in the wild, the research team proposed a multimodal framework based on Hierarchical Granularity Alignment and a State Space Model.
The method uses two foundation models, DINOv2 and WavLM, to extract high-quality visual and audio features, and employs a hierarchical granularity alignment module to establish fine-grained associations between local muscle movements and global facial semantics. It also introduces a Vision-Mamba architecture for ultra-long temporal modeling with linear complexity and designs an Audio-Guided State Space mechanism that uses audio cues to dynamically modulate the visual state update process.
The method achieved leading performance on the Aff-Wild2 dataset, demonstrating its application value for affective behavior analysis in real-world scenarios.
III. EXPR Track: Dual-Branch Transformer Framework for Expression Recognition Under Missing Modalities and Class Imbalance

Paper:
To address visual occlusion, missing modalities, and long-tailed class distributions frequently encountered in real-world expression recognition, the research team proposed a dual-branch Transformer framework for multimodal expression recognition.
The method uses a Dual-Branch Transformer to model visual and audio features separately and Cross-Attention to enable cross-modal information exchange. A Safe Attention mechanism addresses numerical stability when visual input is missing, allowing the model to automatically fall back to audio-driven decision-making. In addition, Focal Loss mitigates class imbalance, while a sliding-window soft-voting strategy supports stable emotion prediction in long videos. Experimental results show that the method can effectively improve the robustness of expression recognition systems in complex environments.
In recent years, Unisound and USTC have pursued extensive collaboration in multimodal AI, affective computing, and large-model technologies, advancing the deployment of related technologies in smart healthcare, intelligent customer service, human-computer interaction, and other fields. The results achieved at the 10th ABAW competition demonstrate the joint team's accumulated expertise and innovative capabilities in affective behavior analysis in the wild.
Looking ahead, the Unisound-USTC joint team will continue to deepen its work in affective computing, multimodal learning, and fundamental AI research; advance human-machine emotion understanding in complex real-world scenarios; and contribute to building more natural, trustworthy, and emotionally intelligent human-computer interaction systems.