Back to feed
arXiv cs.CL
arXiv cs.CL
7/20/2026
Before the Action: Benchmarking LLMs on Prospective Hypothesis Discovery

Before the Action: Benchmarking LLMs on Prospective Hypothesis Discovery

Short summary

This paper introduces Prospective Hypothesis Discovery (PHD), a new evaluation paradigm testing whether LLMs can autonomously construct grounded, testable hypotheses from inconclusive evidence. The authors release HypoArena, a benchmark of 988 cases across six domains, with an evaluation framework combining pairwise judgments and rubric scoring. Experiments on 15 frontier LLMs reveal clear capability stratification, supporting PHD as a distinct evaluation target for open-ended scientific reasoning.

  • Introduces Prospective Hypothesis Discovery (PHD) as a new LLM evaluation paradigm
  • Releases HypoArena: 988-case benchmark across six scientific domains with bidirectional evaluation
  • Tests 15 frontier LLMs, finding clear capability stratification and model-dependent effects

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more