ResearchPod Summary
Individuals with dysarthria often struggle to participate in professional speaking environments like conferences or meetings due to the high latency and unnatural output of traditional Augmentative and Alternative Communication (AAC) systems. The authors sought to develop a real-time, low-latency system that improves the intelligibility and naturalness of dysarthric speech without sacrificing the speaker's agency or semantic intent.
The researchers developed Re-Sonance, a three-stage cascaded architecture:
To achieve real-time performance, the system uses an asynchronous, non-blocking pipeline where ASR, LLM, and TTS modules operate in an overlapping, streaming fashion. The system was evaluated using the Chinese Dysarthric Speech Dataset (CDSD) through both subjective human ratings and objective metrics (Word Error Rate, Match Error Rate, and Word Information Lost) across varying levels of dysarthria severity.
Re-Sonance demonstrated significant improvements in intelligibility and naturalness for speakers with mild and moderate dysarthria compared to baseline ASR-TTS frameworks. For these groups, the system successfully reduced transcription errors and increased semantic association. However, the system struggled with severe dysarthria cases, where the initial ASR recognition errors were too high for the LLM to effectively correct. The system maintained a low Real-Time Factor, confirming its viability for live, interactive communication scenarios.
This research demonstrates that LLMs can act as powerful post-processing tools in assistive technology, capable of inferring intent and correcting errors in real-time. By moving away from rigid speech-replacement models toward speech-enhancement frameworks, Re-Sonance offers a more natural communication experience that preserves the speaker's identity and supports participation in professional settings.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.