ResearchPod Summary
India's rapid digital transformation, driven by massive platforms like Aadhaar (biometric identity) and UPI (payments), has created a unique environment where LLMs are increasingly deployed for government and public services. However, standard AI benchmarks like MMLU and TruthfulQA are heavily Western-centric, failing to account for India's 22 official languages, complex social structures (such as the caste system), and specific risks related to Digital Public Infrastructure (DPI). This paper introduces Inspect India Evals, an open-source framework designed to bridge this gap by evaluating models on their performance within the Indian linguistic and cultural landscape.
Built on the UK AISI's Inspect AI platform, the framework provides a modular, reproducible suite of six benchmarks. These include:
The researchers evaluated five open-weight models (8B–32B parameters). Sarvam-M 24B and Gemma 2 27B emerged as the top performers, achieving an 80% score on the composite India Fairness Index. Notably, Sarvam-M outperformed larger 32B models in cultural knowledge and DPI safety compliance. While all models successfully refused harmful prompts in the Multilingual Safety test, their ability to handle DPI-related queries varied significantly (20% to 100% compliance), highlighting a critical need for domain-specific safety tuning in Indian deployments.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.