ResearchResearch paperLarge Language Models1 source · Oct 8, 2026

Adapting English Quality Classifiers for Multilingual LLM Pretraining Data Selection

Recent advances in large language model (LLM) pretraining highlight the role of high-quality training data in improving performance.

Key points

  • While model-based filtering has proven effective in selecting high-quality subsets from web-scale corpora, especially for high-resource languages, low-resource languages face challenges due to limited availability of annotated data.
  • This work explores extending quality filtering to over 100 languages by proposing a multilingual adaptation approach that converts an existing English quality classifier into a multilingual variant.
  • Our approach proposes training a small multi-layer perceptron on top of Transformer encoder-only model embeddings, using multilingual text as input and scores obtained from English classifiers applied to machine-translated text as labels.
  • To further evaluate cross-lingual generalization, we compare classifier scores of high-quality synthetic data and web samples, and the correlation of classifier scores with LLM-based ones, revealing that the classifier can learn the scoring criteria of its original English variant, even for languages not included in its training data.

Sources (1)

  • [1]Adapting English Quality Classifiers for Multilingual LLM Pretraining Data Selection
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 09:33 AM
    Recent advances in large language model (LLM) pretraining highlight the role of high-quality training data in improving performance.
    While model-based filtering has proven effective in selecting high-quality subsets from web-scale corpora, especially for high-resource languages, low-resource languages face challenges due to limited availability of annotated data.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 7, 2026MotherTree: Meta-learning on synthetic data improves decision tree training
  2. Oct 7, 2026Executing Causal Structure Learning with Linear-Attention Transformers
  3. Oct 7, 2026HuLiGen: Human LiDAR Generation from Parametric Body Models
  4. Oct 6, 2026Algorithmic Scratchpads and Curriculum Staging for Arithmetic Reasoning in Tiny Transformers
  5. Oct 6, 2026When Forgetting is not Catastrophic: On the Mechanics of Spurious Forgetting
  6. Oct 6, 2026HuC-VideoMAE: Human-Centric Video Masked Autoencoding from synthetic data

Related