Back to feed
arXiv cs.CL
arXiv cs.CL
7/9/2026
Healthier LLMs: Retrieval-Augmented Generation for Public Health Question Answering

Healthier LLMs: Retrieval-Augmented Generation for Public Health Question Answering

Short summary

Researchers extend PubHealthBench to a retrieval-augmented setting and systematically evaluate retrieval and generation choices for public health QA using UK Government guidance. Hybrid retrieval consistently improves recall and ranking, and providing retrieved context lets smaller open-weight models match or outperform larger models without retrieval. A rubric-based LLM-as-a-judge for free-form answering shows strongest human agreement on faithfulness and completeness, while factual consistency and clarity are less reliably reproduced.

  • Hybrid retrieval consistently improves recall and ranking for public health QA
  • Smaller open-weight models with retrieval match or outperform larger models without retrieval
  • LLM-as-judge rubric validated against human annotations, with strongest agreement on faithfulness and completeness

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more