Dev.to
8/2/2026

Getting clean SEC financials without fighting XBRL
Short summary
Pulling clean financials from SEC EDGAR is painful due to inconsistent XBRL tags, dimensional context, and unit scaling. The author ran S&P 500 filings through a normalization pipeline (Filingrail) and packaged results as a flat CSV with one row per company-statement-period. A free 100-row sample is available, and the post includes pandas code for computing margins and ROE.
- •XBRL filings have inconsistent tags, context dimensions, and unit scales across filers
- •Normalized CSV: one row per (company, statement_type, period) with USD figures and filing URLs
- •Free 100-row sample plus pandas snippets for margin and ROE calculations
Generated with AI, which can make mistakes.
Is this a good recommendation for you?



