Fluid-Gen-Zero: Grounding Pretrained Video Generators in Physics without Training
We present Fluid-Gen-Zero, a training-free framework for physics-aware fluid-object interaction video generation that decouples physical reasoning from appearance synthesis.
Key points
- Our key insight is to delegate motion dynamics to a physics simulator while preserving the appearance modeling capacity of pretrained video generators.
- We bridge these two domains through a two-level agentic workflow: generation-time planning, where a vision-language model (VLM) agent interprets intent and the simulation rollout to organize generation clips, and latent-space guidance, which injects simulation signals into denoising through region-aware latent wrapping.
- We further introduce a benchmark for fluid-object interaction video generation.
- Across Tora (CogVideoX-based), VACE and WanMove (Wan-based), Fluid-Gen-Zero consistently improves simulation alignment, reducing object trajectory error by 26.7%-81.5% and fluid fEPE (fluid flow endpoint error) by 67.9%-84.0%, while largely preserving perceptual quality.
Sources (1)
- [1]Fluid-Gen-Zero: Grounding Pretrained Video Generators in Physics without TrainingarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 11:16 PM
We present Fluid-Gen-Zero, a training-free framework for physics-aware fluid-object interaction video generation that decouples physical reasoning from appearance synthesis.
Our key insight is to delegate motion dynamics to a physics simulator while preserving the appearance modeling capacity of pretrained video generators.
Extractive summary: sentences quoted from the sources.