Conditional Transfer from Controlled Pretraining Mixtures to Code
Synthetic tasks are increasingly used both as probes of language-model capability and as pretraining data.
Key points
- A task is diagnostic when its loss tracks global pretraining progress; it is teachable when its loss responds to its own token budget; and a data source transfers when including it improves a downstream target.
- We study controlled pretraining in which 70% of the corpus is fixed general Python and the remaining 30% is a simplex over three source families: OpenCodeInstruct, a curated suite of 12 code-adjacent synthetic tasks, and 15 literature-derived probe tasks.
- On the mixture-simplex edge between the curated suite and OpenCodeInstruct, HumanEval pass@20 after a fixed fine-tuning stage rises from 15.9 at pure curated data to its highest observed value, 22.6, at a mixture that is 75% OpenCodeInstruct, then falls to 19.5 at pure OpenCodeInstruct.
- Optimizing near-term task-loss reduction moves the mixture away from the region that transfers.
Sources (1)
- [1]Conditional Transfer from Controlled Pretraining Mixtures to CodearXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 09:12 AM
Synthetic tasks are increasingly used both as probes of language-model capability and as pretraining data.
A task is diagnostic when its loss tracks global pretraining progress; it is teachable when its loss responds to its own token budget; and a data source transfers when including it improves a downstream target.
Extractive summary: sentences quoted from the sources.