Adapting English Quality Classifiers for Multilingual LLM Pretraining Data Selection
Recent advances in large language model (LLM) pretraining highlight the role of high-quality training data in improving performance.
ProofPaper ↗
Key points
- While model-based filtering has proven effective in selecting high-quality subsets from web-scale corpora, especially for high-resource languages, low-resource languages face challenges due to limited availability of annotated data.
- This work explores extending quality filtering to over 100 languages by proposing a multilingual adaptation approach that converts an existing English quality classifier into a multilingual variant.
- Our approach proposes training a small multi-layer perceptron on top of Transformer encoder-only model embeddings, using multilingual text as input and scores obtained from English classifiers applied to machine-translated text as labels.
- To further evaluate cross-lingual generalization, we compare classifier scores of high-quality synthetic data and web samples, and the correlation of classifier scores with LLM-based ones, revealing that the classifier can learn the scoring criteria of its original English variant, even for languages not included in its training data.
Sources (1)
- [1]Adapting English Quality Classifiers for Multilingual LLM Pretraining Data SelectionarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 09:33 AM
Recent advances in large language model (LLM) pretraining highlight the role of high-quality training data in improving performance.
While model-based filtering has proven effective in selecting high-quality subsets from web-scale corpora, especially for high-resource languages, low-resource languages face challenges due to limited availability of annotated data.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 7, 2026MotherTree: Meta-learning on synthetic data improves decision tree training
- Oct 7, 2026Executing Causal Structure Learning with Linear-Attention Transformers
- Oct 7, 2026HuLiGen: Human LiDAR Generation from Parametric Body Models
- Oct 6, 2026Algorithmic Scratchpads and Curriculum Staging for Arithmetic Reasoning in Tiny Transformers
- Oct 6, 2026When Forgetting is not Catastrophic: On the Mechanics of Spurious Forgetting
- Oct 6, 2026HuC-VideoMAE: Human-Centric Video Masked Autoencoding from synthetic data
