Cross-modal audio-text attention for multimodal multitask speech emotion recognition in low-resource Urdu
Speech Emotion Recognition (SER) in low-resource languages remains challenging due to limited annotated data, speaker variability, and the multimodal nature of emotional expression. This paper repositions established components wav2vec 2.0, XLM-R, cross-modal attention, and multitask affective modeling into a framework jointly validated across speaker-independent, cross-lingual zero-shot, and attribution-faithfulness generalization for Urdu, a combination not jointly reported in prior Urdu SER work. The proposed multimodal multitask model achieves 91.3% emotion recognition accuracy on the Urdu Speech Emotion Corpus (UrSEC), outperforming strong audio-only and text-only baselines, with joint valence-arousal learning consistently improving over emotion-only training.
Speaker-independent evaluation shows a performance drop relative to random-split testing but confirms substantial robustness to speaker-specific bias. Cross-lingual zero-shot evaluation on English datasets yields 80.2 ± 1.3% (IEMOCAP) to 86.3 ± 0.9% (CREMA-D) accuracy without fine-tuning, indicating substantial cross-lingual transfer, though this alone does not establish full language-neutrality. All performance gains are statistically validated across multiple runs, and attention/Integrated Gradients analyses, supported by quantitative faithfulness testing, show the model relies on emotionally salient acoustic regions and Urdu tokens rather than spurious correlations.
The work was done with partial support from grants 20260626 (G.S.), and 20260496 (O.K.) by Secretary of Research and Posgraduate Studies (SIP) of Instituto Politécnico Nacional, Mexico. Centro de Investigación en Computación, Instituto Politécnico Nacional, Mexico City, 07320, Mexico Abdullah, Muhammad Ateeb Ather, Olga Kolesnikova & Grigori Sidorov Department of Computer Sciences, Bahria University, Lahore, 54600, Pakistan Faculty of Allied Health Sciences, Superior University, Lahore, 54000, Pakistan Correspondence to Olga Kolesnikova or Grigori Sidorov. The authors declare no competing interests.
This study exclusively utilized a publicly available and pre-annotated dataset (UrSEC). The dataset contains professionally recorded speech data from actors and does not include any sensitive, personal, or identifiable information. According to institutional and international research guidelines, studies involving publicly available, non-sensitive, and anonymized datasets do not require additional ethical approval.
All data collection and annotation procedures for the UrSEC dataset were conducted by the original dataset creators in accordance with applicable ethical standards. The UrSEC dataset consists of speech recordings performed by professional actors. Informed consent was obtained by the original dataset creators from all participants involved in the recordings.
The present study involves only secondary analysis of this publicly available dataset and does not involve direct interaction with human participants. Therefore, no additional informed consent was required for this research. Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Extract — continue reading at the source.