Salta al contenuto principale
Passa alla visualizzazione normale.

SABATO MARCO SINISCALCHI

Rhythm-Aware Modeling for Speaker-Independent Dysarthric Wake-Up Word Spotting

Abstract

Wake-up word spotting (WWS) for dysarthric speech remains challenging because it results in highly variable articulation, unstable timing, and irregular prosodic rhythm. Moreover, existing systems remain largely speaker-dependent, lacking adequate generalization to speaker-independent settings. Motivated by neurophysiological findings that intentional utterances induce preparatory motor activity that facilitates regular rhythmic patterns, we have developed a Rhythm-Aware Wake-up Word Spotting (RAWS) framework that explicitly leverages these cues for dysarthric speech. RAWS comprises three components: (1) a Temporal Prosody Structure Encoder (TPSE) that models speaking-rate and pause-durations via feature extraction, positional encoding, and transformer-based temporal processing; (2) a Large Language Model–Guided Auxiliary Learning (LLM-GAL) mechanism that provides perceptual rhythm-naturalness scores as auxiliary supervision; and (3) a Progressive Adapter-based Domain Alignment (PADA) strategy that enables effective non-dysarthric-to-dysarthric speech transfer while reducing cross-speaker variability. To the best of our knowledge, this study represents the first investigation of speaker-independent dysarthric WWS, with experimental validations on the Mandarin dysarthric speech corpus (MDSC) and its extended version (MDSC v2) demonstrating that RAWS outperforms strong competitive systems, achieving state-of-the-art results.