Back to feed
arXiv cs.CL
arXiv cs.CL
8/3/2026
FinIndices: Benchmarking LLM Financial Reasoning over Long-Horizon Statements

FinIndices: Benchmarking LLM Financial Reasoning over Long-Horizon Statements

Original: Are the Financial Reasoning from LLMs Credible? A Real World Test over Long-Horizon Statements

Short summary

This paper introduces FinIndices, a benchmark evaluating LLM financial reasoning over uncropped financial statements up to 32K tokens with adversarial traps. It reveals two key vulnerabilities: a Knowledge Bottleneck where removing formula hints causes performance collapse (e.g., Gemini-3.1-Pro drops from 70.7% to 38.2% on table tasks), and a Structural Bottleneck where multi-metric table generation drains reasoning capacity. Supervised fine-tuning partially restores structured logic, yielding +8.5% and +3.8% gains on single and table tasks respectively.

  • FinIndices benchmarks LLM financial reasoning over full uncropped statements up to 32K tokens
  • Removing formula hints causes severe performance collapse, exposing fragile pattern matching rather than genuine reasoning
  • SFT partially restores structured logic with +8.5% single-index and +3.8% table-index gains

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more