AION
Research paperLarge Language Models1 source · Oct 8, 2026

Conditional Transfer from Controlled Pretraining Mixtures to Code

Synthetic tasks are increasingly used both as probes of language-model capability and as pretraining data.

Key points

  • A task is diagnostic when its loss tracks global pretraining progress; it is teachable when its loss responds to its own token budget; and a data source transfers when including it improves a downstream target.
  • We study controlled pretraining in which 70% of the corpus is fixed general Python and the remaining 30% is a simplex over three source families: OpenCodeInstruct, a curated suite of 12 code-adjacent synthetic tasks, and 15 literature-derived probe tasks.
  • On the mixture-simplex edge between the curated suite and OpenCodeInstruct, HumanEval pass@20 after a fixed fine-tuning stage rises from 15.9 at pure curated data to its highest observed value, 22.6, at a mixture that is 75% OpenCodeInstruct, then falls to 19.5 at pure OpenCodeInstruct.
  • Optimizing near-term task-loss reduction moves the mixture away from the region that transfers.

Sources (1)

  • [1]Conditional Transfer from Controlled Pretraining Mixtures to Code
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 09:12 AM
    Synthetic tasks are increasingly used both as probes of language-model capability and as pretraining data.
    A task is diagnostic when its loss tracks global pretraining progress; it is teachable when its loss responds to its own token budget; and a data source transfers when including it improves a downstream target.

Extractive summary: sentences quoted from the sources.