sözaltı news Science
Science
EN AZ

Estimation of voice quality parameters of dysarthric speech with preserved identity using different time-frequency image representations in deep learning network

nature.com 29.09.2026 02:00 3 views

Speech-language pathologists use voice quality metrics of raw speech to assess dysarthria. Instead of using raw speech, a deep learning based diagnostic method that extracts these metrics from time-frequency speech representations will be reliable and preserves speaker identity. In this work a regression-based deep convolutional neural network is experimented to estimate jitter, shimmer, fundamental frequency (F0), and harmonic-to-noise ratio (HNR) employing 6 different time-frequency representations: spectrogram, low-frequency spectrogram, cepstrogram, low-frequency cepstrogram, cochleagram, and Mel scalogram.

The experiment is assessed on VOC-ALS dysarthria speech dataset using RMSE and \(\text ^\) metrics, in which low-frequency cepstrogram performs well across all vowels, with average RMSEs of 0.76% for jitter, 2.36% for shimmer, and 4.075 dB for HNR whereas for F0, the cepstrogram shows the highest precision, with an average RMSE of 21.05 Hz. On Parkinson’s speech dataset (PC-GITA), the cepstrogram yields the lowest average RMSEs for jitter (0.57%), F0 (46.65 Hz), and HNR (3.80 dB), closely followed by the low-frequency cepstrogram with average RMSE for jitter (0.58%), F0 (48.53 Hz), and HNR (4.10 dB). For shimmer, the low-frequency cepstrogram attains the lowest average RMSE (2.64%).

The results show that low-frequency cepstrogram and cepstrogram representation performs the best across all vowels for both dysarthric and Parkinson’s speech. Open access funding provided by Vellore Institute of Technology. Open access funding is provided by the Vellore Institute of Technology.

This study did not receive any particular grant from funding organizations in the public, commercial, or non-profit sectors. School of Electronics Engineering, Vellore Institute of Technology, Vellore, Tamil Nadu, 632014, India Department of Computer Science, VIT Mauritius, 72248, Pierrefonds, Mauritius The authors declare no competing interests. Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder.

To view a copy of this licence, visit http://creativecommons.org/licenses/by/4.0/. S, A., M, R.K. & Ramachandran, P. Estimation of voice quality parameters of dysarthric speech with preserved identity using different time-frequency image representations in deep learning network.

Extract — continue reading at the source.

Read full story