arXiv cs.CL
7/20/2026

Before the Action: Benchmarking LLMs on Prospective Hypothesis Discovery
Short summary
This paper introduces Prospective Hypothesis Discovery (PHD), a new evaluation paradigm testing whether LLMs can autonomously construct grounded, testable hypotheses from inconclusive evidence. The authors release HypoArena, a benchmark of 988 cases across six domains, with an evaluation framework combining pairwise judgments and rubric scoring. Experiments on 15 frontier LLMs reveal clear capability stratification, supporting PHD as a distinct evaluation target for open-ended scientific reasoning.
- •Introduces Prospective Hypothesis Discovery (PHD) as a new LLM evaluation paradigm
- •Releases HypoArena: 988-case benchmark across six scientific domains with bidirectional evaluation
- •Tests 15 frontier LLMs, finding clear capability stratification and model-dependent effects
Generated with AI, which can make mistakes.
Is this a good recommendation for you?
