arXiv cs.CL
7/9/2026

Healthier LLMs: Retrieval-Augmented Generation for Public Health Question Answering
Short summary
Researchers extend PubHealthBench to a retrieval-augmented setting and systematically evaluate retrieval and generation choices for public health QA using UK Government guidance. Hybrid retrieval consistently improves recall and ranking, and providing retrieved context lets smaller open-weight models match or outperform larger models without retrieval. A rubric-based LLM-as-a-judge for free-form answering shows strongest human agreement on faithfulness and completeness, while factual consistency and clarity are less reliably reproduced.
- •Hybrid retrieval consistently improves recall and ranking for public health QA
- •Smaller open-weight models with retrieval match or outperform larger models without retrieval
- •LLM-as-judge rubric validated against human annotations, with strongest agreement on faithfulness and completeness
Generated with AI, which can make mistakes.
Is this a good recommendation for you?