arXiv cs.CL
7/31/2026

HSS-Synth: Humanities and Social Sciences Data Synthesis for LLMs
Short summary
HSS-Synth is the first data synthesis pipeline for humanities and social sciences, covering 14 fields with a subject-centric paradigm. It constructs seed documents from web corpora, backtranslates them into diverse instructions with Q&A alignment checks, and uses teacher-forced answering to anchor semantics and reduce hallucinations. The pipeline yields 237K instruction-tuning samples; fine-tuned Qwen3-8B-Base sets a new SOTA across 16 benchmarks, approaching the official Qwen3-8B. Code is publicly available.
- •First HSS data synthesis pipeline covering 14 fields with subject-centric paradigm
- •237K instruction-tuning samples; Qwen3-8B-Base sets SOTA on 16 benchmarks
- •Teacher-forced answering reduces hallucinations and preserves tone; code is open-source
Generated with AI, which can make mistakes.
Is this a good recommendation for you?