ResearchPod Summary
Large language models achieve impressive results on English medical benchmarks, but it remains unclear whether these capabilities transfer reliably to other languages. Many existing multilingual medical benchmarks rely on unverified machine translations, which can distort medical terminology and syntax. This study investigates the true multilingual competence of 14 prominent language models using HealMed, a meticulously expert-reviewed medical benchmark developed across nine languages.
HealMed contains 1,000 examples per language drawn from nine source datasets, covering multiple-choice question answering (MCQA), natural language inference (NLI), and open-ended question answering (QA). The evaluated languages include English, German, Spanish, Portuguese, Japanese, Chinese, Thai, Swahili, and Zulu. The benchmark construction involved 23 medical experts across nine countries who performed a rigorous two-stage review of machine-translated instances, scoring and revising them for accuracy, fluency, and completeness.
The evaluation demonstrates that the strongest proprietary models combine high accuracy with remarkable cross-language stability, whereas non-proprietary open-source and medically specialized models experience severe performance drops in lower-resource languages. Crucially, medical specialization alone does not guarantee multilingual robustness. Performance losses were heavily concentrated in Swahili and Zulu, driven by major terminology and fluency issues. Furthermore, paired comparisons against unrevised machine-translated data show that expert review substantially alters measured performance, proving that raw machine translation introduces measurement bias in cross-lingual evaluation.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.