Back to feed
arXiv cs.CL
arXiv cs.CL
7/31/2026
Benchmarking LLM Competence on Logical Inference over Probability Operators

Benchmarking LLM Competence on Logical Inference over Probability Operators

Short summary

This paper introduces a benchmark of 14,320 procedurally generated English prompts to test LLM reasoning over probability operators like 'probably,' 'might,' and 'must' across fifteen inference templates. Evaluating 29 models, the authors find most exhibit systematic answer biases independent of logical form, and only 9 of 29 exceed random chance on a competence-floor metric. Additional tests reveal biases across question form, verb phrases, and name gender or origin.

  • 14,320 prompts testing LLM inference over gradable epistemic modals across 15 templates
  • Only 9 of 29 models exceed random chance; most show systematic Yes/No answer biases
  • Biases found across question form, verb activity, and name gender/origin

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more