Ambient Discrete Diffusion: Using the Wrong Data at the Right Time for Data Efficient Learning
We introduce RefineMix, a framework for training discrete diffusion models under severe data scarcity, a common constraint in scientific applications.
Key points
- RefineMix uses out-of-distribution data at selected diffusion times to improve generalization without biasing the sampling distribution.
- Although this strategy has been explored in continuous diffusion, discrete diffusion presents a distinct challenge: unlike Gaussian noise, masking preserves domain information in surviving tokens, limiting the use of related data at high noise levels.
- Experimentally, across five domain-shift settings, RefineMix matches or outperforms in-domain finetuning and data mixing.
- For protein sequence generation, finetuning with just 197 in-domain examples nearly doubles the fraction of generated proteins that are simultaneously novel, foldable, and in-family compared to standard finetuning.
Sources (1)
- [1]Ambient Discrete Diffusion: Using the Wrong Data at the Right Time for Data Efficient LearningarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 05:15 PM
We introduce RefineMix, a framework for training discrete diffusion models under severe data scarcity, a common constraint in scientific applications.
RefineMix uses out-of-distribution data at selected diffusion times to improve generalization without biasing the sampling distribution.
Extractive summary: sentences quoted from the sources.