
FinIndices: Benchmarking LLM Financial Reasoning over Long-Horizon Statements
Original: Are the Financial Reasoning from LLMs Credible? A Real World Test over Long-Horizon Statements
Short summary
This paper introduces FinIndices, a benchmark evaluating LLM financial reasoning over uncropped financial statements up to 32K tokens with adversarial traps. It reveals two key vulnerabilities: a Knowledge Bottleneck where removing formula hints causes performance collapse (e.g., Gemini-3.1-Pro drops from 70.7% to 38.2% on table tasks), and a Structural Bottleneck where multi-metric table generation drains reasoning capacity. Supervised fine-tuning partially restores structured logic, yielding +8.5% and +3.8% gains on single and table tasks respectively.
- •FinIndices benchmarks LLM financial reasoning over full uncropped statements up to 32K tokens
- •Removing formula hints causes severe performance collapse, exposing fragile pattern matching rather than genuine reasoning
- •SFT partially restores structured logic with +8.5% single-index and +3.8% table-index gains
Generated with AI, which can make mistakes.
Is this a good recommendation for you?
