ResearchPod Summary
Metaphors are essential cognitive tools that allow for the communication of complex concepts through familiar, concrete imagery. However, their reliance on cultural context and non-literal meaning makes them notoriously difficult for machine translation (MT) systems. While general translation quality has improved, metaphor-specific performance often lags behind. This paper introduces MetaHOPE, an error-severity-aware annotation framework adapted from the HOPE metric, designed to systematically identify and score metaphor translation errors.
MetaHOPE categorizes errors into five specific types: Impact (IMP), Required Adaptation Missing (RAM), Mistranslation (MIS), Style (STL), and Proofreading Error (PRF). Unlike standard metrics that focus on sentence-level fluency, MetaHOPE requires annotators to consider the full document context before evaluating the translation of metaphor-related words (MRWs). The framework employs a five-level severity scale (minor to critical) to provide a nuanced view of how translation systems handle figurative language.
In a pilot study using English-Chinese and Chinese-English news corpora, the authors evaluated GoogleMT, GPT-5.4, and Hunyuan-7b. The results indicate that metaphor-related errors account for over 90% of translation penalties in the tested systems, confirming that figurative language remains a significant hurdle. Qualitative analysis revealed a trade-off between systems: GoogleMT and GPT-5.4 tend to be more conservative and literal, whereas Hunyuan-7b exhibits greater flexibility in localization but is more prone to hallucinations or omissions. The study also highlights that inter-annotator agreement is lowest for stylistic and impact-related errors, suggesting that these aspects of metaphor translation are highly subjective.
MetaHOPE provides a much-needed bridge between cognitive metaphor theory and empirical MT evaluation. By moving beyond coarse-grained metrics like BLEU or COMET, this framework allows researchers to pinpoint exactly where and how metaphorical meaning is lost or distorted. The provided parallel corpora and error analysis offer a valuable resource for future research aimed at improving the cross-cultural and cross-linguistic capabilities of LLMs.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.