Back to feed
arXiv cs.CL
arXiv cs.CL
6/30/2026
Legal Domain Adaptation of Modern BERT Models

Legal Domain Adaptation of Modern BERT Models

Short summary

Researchers fine-tuned ModernBERT on US court opinions via domain adaptation, improving performance on legal benchmarks. Released models support sequences up to 8,192 tokens for legal document embedding and retrieval. Further pre-training of existing checkpoints outperformed training from scratch despite ModernBERT's massive base training data.

  • Domain adaptation improved ModernBERT across all legal document datasets
  • Models handle 8K-token sequences for legal passages and ranking
  • Incremental pre-training beats scratch training even with 500x larger base data

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more