Back to feed
arXiv cs.CL
arXiv cs.CL
7/16/2026
GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models

Short summary

This paper introduces GSM-Plus-BN, a perturbation-based Bengali mathematical reasoning benchmark derived from the English GSM-Plus dataset, comprising 9,000 evaluation samples across six open-source LLMs. GPT-OSS-20B achieved the highest seed-question accuracy (96.08%) under standard prompting, while larger models like Llama-3.3-70B and GPT-OSS-120B showed superior robustness to perturbations. CoT prompting improved reasoning for most models, but a notable performance gap persisted relative to English benchmarks, highlighting the inherent difficulty of perturbed Bengali text.

  • GSM-Plus-BN is a new 9,000-sample perturbed Bengali math reasoning benchmark for LLMs
  • GPT-OSS-20B leads seed accuracy at 96.08%; larger models show better perturbation robustness
  • CoT prompting helps but a significant gap remains vs English benchmarks across all models

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more