arXiv cs.CL
7/15/2026

Scaling Point-in-Time Language Models
Short summary
Researchers trained decoder-only transformers up to 4B parameters on 1 trillion chronologically filtered FineWeb tokens to create point-in-time language models that avoid future-data leakage. Monthly checkpoints spanning 2013-2024 approach the performance of comparable open-weight models like Gemma-3-4B and LLaMA-7B on common-sense reasoning benchmarks. The full pipeline—dataset construction, training, and evaluation code—is released to support reproducible research requiring strict temporal validity.
- •4B-parameter models trained on 1T chronologically filtered tokens eliminate lookahead bias
- •Performance approaches Gemma-3-4B and LLaMA-7B on reasoning benchmarks
- •Full pipeline released for reproducible point-in-time LLM research
Generated with AI, which can make mistakes.
Is this a good recommendation for you?
