Back to feed
arXiv cs.CL
arXiv cs.CL
6/16/2026
Evaluating the Robustness of Proof Autoformalization in Lean 4

Evaluating the Robustness of Proof Autoformalization in Lean 4

Short summary

Researchers benchmarked seven LLM models on proof autoformalization robustness in Lean 4, testing both global perturbations (paraphrasing) and local perturbations (value/step changes). All models showed significant failures—failing to remain consistent under style changes and not faithfully reflecting proof modifications. Code and benchmark released open-source on GitHub.

  • First robustness evaluation of LLM-based proof autoformalization models
  • All 7 tested models failed under global and local perturbations
  • Open-source benchmark and code released for reproducibility

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more