Back to feed
Dev.to
Dev.to
7/21/2026
Why I Built a 23,000+ Record Open Dataset for Gujarati AI

Why I Built a 23,000+ Record Open Dataset for Gujarati AI

Short summary

The author built GGJI v1, a 23,181-record Gujarati instruction-tuning dataset covering 18 task categories, released on Hugging Face under Apache 2.0. Existing open-source models struggle with Gujarati, often mixing languages or defaulting to Hindi/English despite 60+ million speakers. The dataset targets fine-tuning, benchmarking, and experimentation, with acknowledged gaps in dialect coverage that the author plans to address in future versions.

  • 23,181 Gujarati instruction-response pairs across 18 task categories released on Hugging Face
  • Addresses gap in low-resource language datasets for 60M+ Gujarati speakers
  • Apache 2.0 licensed, suitable for research and commercial use

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more