ResearchPod Summary
Automatic Speech Recognition (ASR) is increasingly used to document psychiatric interviews, yet its reliability in the multilingual and demographically diverse Indian healthcare context remains largely unverified. Psychiatric interviews are particularly challenging due to unique speech patterns associated with mental health conditions, noisy clinical environments, and the inherent power imbalance between clinicians and patients. This study addresses the critical need for accurate and equitable ASR systems in these high-stakes clinical settings.
The authors conducted a comprehensive audit of eight state-of-the-art ASR models—including WhisperLargeV3, Gemini, and IndicWhisper—using a novel dataset of 202 real-world psychiatric interviews spanning Kannada, Hindi, and Indian English. The audit revealed substantial performance variability across models and languages. While some systems performed competitively in Indian English, they struggled significantly with regional languages like Kannada. Furthermore, the researchers identified systematic performance gaps tied to speaker roles (doctor vs. patient) and gender, highlighting concerns regarding the equitable deployment of these tools in clinical practice.
To address these limitations, the authors developed SamaVaani, a unified debiasing framework. The approach uses LoRA (Low-Rank Adaptation) to fine-tune pre-trained models while incorporating two key components: contrastive learning to improve robustness against phonetic variance (such as pitch differences) and a Connectionist Temporal Classification (CTC) head to enhance character-level sequencing accuracy. By combining these techniques, SamaVaani achieves up to a 50% reduction in overall Word Error Rate (WER) compared to base models and demonstrates improved fairness across diverse demographic groups.
This research provides a scalable, actionable pathway for deploying ASR in multilingual clinical environments. By moving beyond general-purpose accuracy metrics to prioritize fairness and robustness in psychiatric contexts, the study offers a blueprint for ensuring that AI-driven clinical documentation does not exacerbate existing healthcare disparities.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.