Adaptive speech-to-image translation for impaired speech using parameter-efficient ASR and diffusion-based image generation
Impaired Speech Continues to Pose Challenges for the Automatic Processing of Speech and Multimodal Interaction between Humans and Computers systems caused by distortion of articulation and involuntary repetition. Converting impaired speech into sight allows a complementary system of communicating for individuals who have speech disorders; especially when the textual input is often the only method of expression difficult to interpret. This study proposes a cascaded speech-to-image translation framework which integrates adaptive speech recognition using the diffusion-based image generating.
The speech recognition part is improved based on Silero-based voice activity detection, parameter-efficient fine-tuning using AdaLoRA and shallow fusion using a large language model to improve under dysarthric conditions the robustness of the transcription; The refined textual output is then used to condition a latent diffusion model optimized for controlled image synthesis. Experiments were conducted on the publicly available TORGO dysarthric speech corpus. The proposed configuration achieved a word error rate (WER) of 0.125 and a character error rate (CER) of 0.050, substantially outperforming the baseline Whisper models.
Semantic correspondence between generated images and their conditioning text prompts are quantified using CLIPScore with 30.18 a score of 30.18 for the proposed system. indicating spread of multimodal consistency compared to non-adaptive baselines. Qualitative results further show that the framework maintains important semantic features across diverse categories of prompts such as natural scenes, human portraits and action-oriented layouts. These results indicated that adaptation of speech recognition plays a critical role in stabilizing downstream generating visuals and supporting the feasibility of speech-to-image translation as an assistive communication modality for impaired speech.
Department of Artificial Intelligence and Data Science, Muthoot Institute of Technology and Science, Kochi, Kerala, India Department of Computer Science and Business Systems, KPR Institute of Engineering and Technology, Coimbatore, Tamil Nadu, India The human evaluation conducted in this study involved an anonymous online survey in which adult participants rated generated images for realism, semantic alignment, and visual quality. No personally identifiable or sensitive information was collected. All methods were carried out in accordance with relevant guidelines and regulations.
Participation was voluntary, and informed consent was obtained from all participants prior to completing the survey.The survey involved anonymous, voluntary participation by adult participants, posed minimal risk, and did not involve the collection of sensitive or personally identifiable information. Under these conditions, formal institutional ethical approval was not required. The images presented in this study are synthetic outputs generated by the proposed speech-to-image pipeline and do not correspond to identifiable real individuals.
Therefore, consent for publication of identifiable participant information was not required. The authors declare no competing interests. Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Extract — continue reading at the source.