ResearchResearch paperRobotics & Embodied AI · Image, Video & 3D Generation · Efficiency & Inference1 source · Oct 7, 2026

Think Before You Paint: Recursive Latent Reasoning for Diffusion Models

We propose Painter-Thinker (PaTh): a small recursive network (the Thinker) reasons over a grid of learned tokens that encode the noisy image and the conditioning, refines a latent state within every denoising step, and steers a frozen diffusion model (the Painter) through ControlNet adapters.

Key points

  • Diffusion models generate realistic images but often fail on visual reasoning tasks, such as filling in a Sudoku or drawing the path through a maze.
  • When a discrete symbolic representation is available, recursive methods such as the Tiny Recursive Model (TRM) solve even hard instances of these puzzles.
  • We ask how such reasoning can be carried over to pixels, where no symbolic representation is available.
  • Together, these results show that reasoning mechanisms developed for symbolic data can be integrated into pixel-space diffusion without symbolic supervision, opening a path toward generating data under increasingly complex constraints.

Sources (1)

  • [1]Think Before You Paint: Recursive Latent Reasoning for Diffusion Models
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 11:36 AM
    We propose Painter-Thinker (PaTh): a small recursive network (the Thinker) reasons over a grid of learned tokens that encode the noisy image and the conditioning, refines a latent state within every denoising step, and steers a frozen diffusion model (the Painter) through ControlNet adapters.
    Diffusion models generate realistic images but often fail on visual reasoning tasks, such as filling in a Sudoku or drawing the path through a maze.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 7, 2026ORCA: Hunting Compositional Failures in Text-to-Image Diffusion
  2. Oct 7, 2026Relational Abstractions for Spatial Reasoning with Diffusion Models
  3. Oct 7, 2026Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning
  4. Oct 6, 2026Disentangling Dual Image References in Frequency Aware Diffusion Models for Personalized Generation
  5. Oct 6, 2026Uniform Discrete Diffusion Models are Minimax Optimal for Estimating Distributions with Small Effective Support Size
  6. Aug 10, 2026vllm-project/vllm v0.27.0

Related