arXiv cs.LG
7/17/2026

How Much of a 10-K Matters? Aggregation-Dependent Value of Full-Text versus Risk-Factor Sentiment
Short summary
This paper extends supervised lexicon-learning sentiment extraction to 10-K filings and their Item 1A risk-factor sections, training against both return and volatility labels at sector, portfolio, and firm levels. Full-filing text is more accurate at sector/portfolio level, but the narrower Item 1A section wins at the individual-firm level due to signal-to-volume tradeoffs. A Loughran-McDonald dictionary baseline is consistently negatively correlated with price, validating the supervised approach for regulatory text.
- •Supervised sentiment from 10-K filings outperforms dictionary baselines for volatility and return prediction
- •Full-filing text wins at aggregate levels; Item 1A risk factors win at firm level
- •Study covers 1,383 filings from 94 Nasdaq-100 tech firms (2006–2023)
Generated with AI, which can make mistakes.
Is this a good recommendation for you?
