SciTBERT: A family of chronologically consistent language models for scientific and technological language processing
We introduce SciTBERT: a family of chronologically consistent BERT-derived language models trained on text from scientific papers, patents, and high-quality educational web text with training data cutoff dates spanning each year between 2013 and 2025.
Key points
- Pre-trained transformer models are increasingly being used to study scientific and technological progress.
- Encoders tuned to paper or patent text outperform general-purpose models on downstream classification, regression, and proximity tasks within science and technology.
- We find that these models generally outperform predecessor domain-specific encoder models even when training data is limited by early year restrictions in the corpus.
- To further investigate the extent to which this class of models can learn representations that bridge science and technology, we introduce the PatRepEval benchmark, a suite of patent-related text embedding tasks at the science-technology interface.
Sources (1)
- [1]SciTBERT: A family of chronologically consistent language models for scientific and technological language processingarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 03:57 PM
We introduce SciTBERT: a family of chronologically consistent BERT-derived language models trained on text from scientific papers, patents, and high-quality educational web text with training data cutoff dates spanning each year between 2013 and 2025.
Pre-trained transformer models are increasingly being used to study scientific and technological progress.
Extractive summary: sentences quoted from the sources.