ResearchPod Summary
Mathematical reasoning is a critical benchmark for evaluating the cognitive depth of Large Language Models (LLMs). While English-language benchmarks are well-established, there is a severe lack of resources for low-resource languages like Bengali, which is spoken by over 230 million people. This study addresses the gap in assessing model robustness in Bengali by introducing GSM-PLUS-BN, a comprehensive, perturbation-based benchmark designed to test whether models truly understand mathematical structures or merely rely on pattern recognition.
The researchers developed GSM-PLUS-BN by translating and perturbing 1,000 seed questions from the English GSM-Plus dataset, resulting in 9,000 evaluation samples. They applied eight distinct perturbation types—including numerical substitution, digit expansion, and distractor insertion—to test model resilience. The team evaluated six open-source LLMs (ranging from 8B to 120B parameters) using both Standard Prompting and Chain-of-Thought (CoT) prompting to compare base reasoning capabilities against structured, step-by-step reasoning.
The experiments revealed that while GPT-OSS-20B achieved the highest accuracy on seed questions (96.08%) under standard prompting, larger models like Llama-3.3-70B and GPT-OSS-120B demonstrated superior robustness when faced with perturbed inputs. Although CoT prompting consistently improved performance across most models, a persistent performance gap remains when compared to English benchmarks. This suggests that the inherent complexity of perturbed Bengali text poses a significant challenge that current LLMs have yet to fully overcome, regardless of their parameter scale.
This research provides a foundational resource for the equitable development of AI in linguistically diverse regions. By establishing a standardized, perturbation-based benchmark for Bengali, the study enables researchers to move beyond simple accuracy metrics and rigorously evaluate the reliability of LLMs in real-world, non-English production environments where input variations are common.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.