Do Vision Models Learn Physical Constraints or Rendering Shortcuts? A Counterfactual Benchmark for Grounded Physical Consistency
We study physical plausibility diagnosis, detecting whether an edited image violates scene physics, naming the violation type, localizing the affected region, and explaining the failure in language.
ProofPaper ↗
Key points
- Modern image editing models can satisfy a text instruction while breaking the physics of the edited scene.
- A new object may cast no shadow, a mirror may fail to reflect visible geometry, or an object may float above a surface that should support it.
- We introduce a counterfactual benchmark whose controlled synthetic component uses Mitsuba 3 to generate 5,500 images from 500 scene families.
- The renderer pipeline provides category labels, affected-region masks and boxes, scene metadata, and explanation targets.
Sources (1)
- [1]Do Vision Models Learn Physical Constraints or Rendering Shortcuts? A Counterfactual Benchmark for Grounded Physical ConsistencyarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 11:06 PM
We study physical plausibility diagnosis, detecting whether an edited image violates scene physics, naming the violation type, localizing the affected region, and explaining the failure in language.
Modern image editing models can satisfy a text instruction while breaking the physics of the edited scene.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 6, 2026[AINews] Reflection Beam - 501B-A23B American Open Model
- Oct 5, 2026perplexity-ai/pplx-decider-v1.1-27b
- Oct 4, 2026nerkyor/Qwen3.8-27B-Coder390-EfficientThink-Opus5.5-GPT6Astra-Grok4.7-DSV4Pro-K3-SFT-RLOO-MTP-DFlash2
- Oct 2, 2026alesha-pro/Qwen3.8-Flash-Next-abliterated-GSQ-RCO-Strata-GGUF
- Oct 1, 2026nvidia/PixelUMM
- Sep 28, 2026Holo4: powering generalist computer-use agents