arXiv cs.CL
7/1/2026

Indi-RomCoM: Code-Mixed Benchmark for Evaluating LLMs on Romanized Indic-English Instructions
Short summary
Researchers introduce Indi-RomCoM, a systematic benchmark for evaluating Large Language Models on romanized code-mixed instructions—a dominant communication pattern where bilingual speakers blend Indic languages with English in Roman script. Testing across seven instruction tasks and four Indian languages shows LLMs consistently underperform on code-mixed content, with degradation increasing alongside code-mixing density. The benchmark highlights significant gaps in multilingual AI systems and offers a framework to guide development of more inclusive language models for diverse linguistic communities.
- •New Indi-RomCoM benchmark evaluates LLMs on romanized code-mixed (Indic + English) instructions across 7 tasks and 4 languages
- •LLMs significantly underperform on code-mixed content; performance degrades as code-mixing intensity increases
- •Reasoning tasks degrade less than detection tasks, revealing how LLMs handle mixed-language reasoning versus detection
Generated with AI, which can make mistakes.
Is this a good recommendation for you?