VIS-Ground: Video Interactive Storytelling with Contextual Grounding
Video interactive storytelling enables viewers to actively steer how a video unfolds.
ProofPaper ↗
Key points
- This leaves a fundamental question underexplored: What context should a generation model ground on during interactive continuation, and how can heterogeneous, unstructured inputs be transformed into such grounding context?
- In this work, we formulate contextual grounding as the process of transforming heterogeneous input context into an executable constraint model for video generation.
- To address this challenge, we introduce VIS-Ground, which performs Structured Context Abstraction to recover grounded states and cross-context dependencies, Generation Constraints Induction to project relevant dependencies into candidate-specific constraints, and Constrained Video Generation to enforce these constraints through planning, verification, revision, and rendering.
- Detailed analysis further shows gains across both narrative and knowledge grounding, and reveals remaining challenges in dependency extraction, and faithful realization during video rendering.
Sources (1)
- [1]VIS-Ground: Video Interactive Storytelling with Contextual GroundingarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 02:31 AM
Video interactive storytelling enables viewers to actively steer how a video unfolds.
This leaves a fundamental question underexplored: What context should a generation model ground on during interactive continuation, and how can heterogeneous, unstructured inputs be transformed into such grounding context?
Extractive summary: sentences quoted from the sources.