Back to feed
arXiv cs.CL
arXiv cs.CL
7/15/2026
Scaling Point-in-Time Language Models

Scaling Point-in-Time Language Models

Short summary

Researchers trained decoder-only transformers up to 4B parameters on 1 trillion chronologically filtered FineWeb tokens to create point-in-time language models that avoid future-data leakage. Monthly checkpoints spanning 2013-2024 approach the performance of comparable open-weight models like Gemma-3-4B and LLaMA-7B on common-sense reasoning benchmarks. The full pipeline—dataset construction, training, and evaluation code—is released to support reproducible research requiring strict temporal validity.

  • 4B-parameter models trained on 1T chronologically filtered tokens eliminate lookahead bias
  • Performance approaches Gemma-3-4B and LLaMA-7B on reasoning benchmarks
  • Full pipeline released for reproducible point-in-time LLM research

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more