ResearchPod Summary
Language models are increasingly integrated into global workflows, yet they often struggle with nuanced socio-cultural contexts outside of mainstream Western culture. Traditional dance, which serves as a vital repository of history, identity, and social values, is frequently marginalized in AI training data. To address this, the authors present NRITYAM, a new benchmark specifically designed to test the cultural comprehension of AI systems regarding traditional dance forms from 12 countries across 5 continents.
The dataset was constructed through a rigorous, multi-phase process involving 36 native-speaking domain experts. These annotators, who possess formal backgrounds in dance, sourced information from diverse materials—including government cultural websites, scholarly journals, and local blogs—to ensure authenticity. The resulting 9,260 question-answer pairs are divided into text-based and image-based modalities, covering three primary reasoning categories: history-based, rule-based, and scenario-based. The benchmark supports 12 languages, emphasizing the inclusion of mid- to low-resource languages to challenge the linguistic biases of current foundation models.
By evaluating a broad suite of Large Language Models (LLMs), Small Language Models (SLMs), and Multimodal Language Models (MLMs), this work highlights critical gaps in how AI systems perceive and reason about non-Western cultural practices. NRITYAM provides a standardized, high-quality evaluation framework that encourages the development of more inclusive and culturally competent AI, moving beyond the current focus on popular global culture and toward a more equitable representation of human heritage.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.