sözaltı news Science
Science
EN AZ

Speech intelligibility comparison of standalone and two-stage deep learning architectures for behind-the-ear-to-binaural enhancement

nature.com 18.09.2026 02:00 1 views

The relative perceptual performance of alternative deep learning based speech processing architectures remains uncertain under complex acoustic conditions. This study compared speech intelligibility, in listeners with normal hearing and with mild-to-moderate hearing loss, obtained with two processing models: a speech-target-oriented two-stage model (SED + U-Net), designed to estimate speech-target activity before denoising, and a standalone U-Net model. These two architectures were selected as representative instances of two contrasting design philosophies in deep-learning-based speech enhancement: unconstrained direct denoising versus denoising explicitly guided by a prior estimate of speech-target activity.

Intelligibility was assessed using a keyword-recognition task with sentences from the Chilean SHARVARD corpus, presented under controlled virtual acoustic conditions. Twenty-four adults participated in the study, consisting of 12 normal-hearing (NH) listeners and 12 listeners with mild-to-moderate hearing loss (HL). Experimental conditions varied by processing model, masker configuration (Masker Around, Masker Front, and Masker Side), reverberation time (RT = 0.2 s and 0.8 s), and signal-to-noise ratio (− 3, 0, 3, 6, and 9 dB).

Keyword recognition was analyzed using a binomial generalized linear mixed-effects model. Results showed significant effects of SNR, processing model, masker configuration, and reverberation time. Unlike the original hypothesis, the standalone U-Net performed better in keyword recognition than the speech-target-oriented two-stage model, emphasizing the need for perceptual validation in evaluating speech-processing systems.

The authors would like to thank the team at the Laboratory of Experimental Studies of Communication, USS, who helped in the listening test development. The author R.V-M. acknowledges support from ANID FONDECYT Postdoc through grant number 3230356. Departamento de Electrónica e Informática, Universidad Técnica Federico Santa María, 4030000, Concepción, Chile Rhoddy Viveros-Muñoz & Sebastián Guajardo-Herrera Facultad de Ciencias de la Rehabilitación y Calidad de Vida, Universidad San Sebastián, 4030000, Concepción, Chile Universidad de Concepción, 4030000, Concepción, Chile Universidad de la Santísima Concepción, 4030000, Concepción, Chile Universidad del Biobío, 4030000, Concepción, Chile Institute of Acoustics, University Austral of Chile, 5090000, Valdivia, Chile The authors declare no competing interests.

The speech recording protocol was approved by the Institutional Ethics Committee of the University of Santiago, Chile, No. 379/2023 (date of approval 07 July 2023). Informed consent was obtained from all participants before the recording. Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder.

Extract — continue reading at the source.

Read full story