ResearchPod Summary
Large language models (LLMs) are predominantly trained on English-centric web data, which often leads to a misalignment with the cultural norms and societal values of non-Western, multilingual nations. Sri Lanka, a country with a unique blend of multi-ethnic, religious, and post-conflict social dynamics, lacks a dedicated framework for evaluating or fine-tuning models to reflect its specific cultural context. Existing benchmarks often focus on broad cross-cultural awareness or basic language proficiency, failing to capture the nuanced, situated values that guide social behavior in Sri Lanka.
To address this, the researchers developed a survey-driven methodology to identify and operationalize 40 majority-endorsed Sri Lankan societal values. The process involved a trilingual survey of 205 participants, combining established international frameworks (such as the World Values Survey) with LLM-assisted elicitation to capture local constructs.
Using these 40 values, the team constructed two primary resources:
The researchers evaluated several proprietary and open-weight LLMs, finding that even large, modern models struggle with cultural value alignment in Sinhala. By fine-tuning Qwen-family models using a mixture of LKvaluesIT and general-purpose Sinhala instruction data, the authors demonstrated a significant reduction in invalid outputs and cross-lingual disparities. However, the study highlights that these improvements are not universal; the effectiveness of the fine-tuning recipe varies significantly depending on the model architecture, suggesting that successful low-resource value alignment requires careful, model-specific adaptation.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.