WorldSonus: Bringing Sound to Worlds
To address these demands, we introduce WorldSonus, an interactive video-to-audio framework designed for real-time spatial sound synthesis in world models.
ProofPaper ↗
Key points
- Recent advances in world models have enabled increasingly realistic visual synthesis.
- Bringing sound to world models poses three core challenges: real-time generation to keep pace with interactive video streams, interactive control to respond to mid-stream sound instructions, and spatially aligned stereo to reflect scene geometry and camera motion.
- For real-time generation, WorldSonus employs a streaming causal autoregressive diffusion architecture that synthesizes audio chunks at a low real-time factor (RTF) of 0.41.
- Extensive experiments demonstrate that while tailored for world models, WorldSonus generalizes effectively to open-domain video-to-audio benchmarks, matching or outperforming state-of-the-art bidirectional models in both acoustic quality and spatial alignment.
Sources (1)
- [1]WorldSonus: Bringing Sound to WorldsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 05:50 PM
To address these demands, we introduce WorldSonus, an interactive video-to-audio framework designed for real-time spatial sound synthesis in world models.
Recent advances in world models have enabled increasingly realistic visual synthesis.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 6, 2026Independent Multi-Agent Reinforcement Learning with Counterfactual Semantic-Social World Models
- Oct 6, 2026SPW-Nav Streams Language-Guided Panoramic Video in Real Time
- Oct 5, 2026From Traces to Agentic Worlds: Agentic Language World Models for Interactive Environment Simulation
- Sep 29, 2026Introducing Quine: An AI research system designed for the complexity of biology
- Sep 29, 2026LongTake: Learning to Sustain Dynamics in Long-Horizon Video Generation
- Jun 9, 2026Powering the future of robotics in Europe