arXiv cs.CL
7/21/2026

OpenLanguageModel: Readable and Composable Small-Language-Model Pretraining for Education and Research
Short summary
OpenLanguageModel (OLM) is an MIT-licensed PyTorch library for building and pretraining small language models with transparent, readable code. It supports tokenizers, streaming datasets, mixed precision, and CPU through multi-GPU execution, with 27 presets across nine model families. Validation shows 90.6% four-GPU weak-scaling efficiency for a 348M-parameter workload and close agreement with reference implementations.
- •MIT-licensed PyTorch library for readable, composable small LM pretraining
- •27 presets across nine model families, from teaching notebooks to full training runs
- •90.6% four-GPU weak-scaling efficiency; available via PyPI and GitHub
Generated with AI, which can make mistakes.
Is this a good recommendation for you?

