AION
Research paperLarge Language Models1 source · Oct 8, 2026

SciTBERT: A family of chronologically consistent language models for scientific and technological language processing

We introduce SciTBERT: a family of chronologically consistent BERT-derived language models trained on text from scientific papers, patents, and high-quality educational web text with training data cutoff dates spanning each year between 2013 and 2025.

Key points

  • Pre-trained transformer models are increasingly being used to study scientific and technological progress.
  • Encoders tuned to paper or patent text outperform general-purpose models on downstream classification, regression, and proximity tasks within science and technology.
  • We find that these models generally outperform predecessor domain-specific encoder models even when training data is limited by early year restrictions in the corpus.
  • To further investigate the extent to which this class of models can learn representations that bridge science and technology, we introduce the PatRepEval benchmark, a suite of patent-related text embedding tasks at the science-technology interface.

Sources (1)

Extractive summary: sentences quoted from the sources.