Dev.to
7/21/2026

Why I Built a 23,000+ Record Open Dataset for Gujarati AI
Short summary
The author built GGJI v1, a 23,181-record Gujarati instruction-tuning dataset covering 18 task categories, released on Hugging Face under Apache 2.0. Existing open-source models struggle with Gujarati, often mixing languages or defaulting to Hindi/English despite 60+ million speakers. The dataset targets fine-tuning, benchmarking, and experimentation, with acknowledged gaps in dialect coverage that the author plans to address in future versions.
- •23,181 Gujarati instruction-response pairs across 18 task categories released on Hugging Face
- •Addresses gap in low-resource language datasets for 60M+ Gujarati speakers
- •Apache 2.0 licensed, suitable for research and commercial use
Generated with AI, which can make mistakes.
Is this a good recommendation for you?



