ResearchResearch paperImage, Video & 3D Generation1 source · Oct 7, 2026

VIS-Ground: Video Interactive Storytelling with Contextual Grounding

Video interactive storytelling enables viewers to actively steer how a video unfolds.

Key points

  • This leaves a fundamental question underexplored: What context should a generation model ground on during interactive continuation, and how can heterogeneous, unstructured inputs be transformed into such grounding context?
  • In this work, we formulate contextual grounding as the process of transforming heterogeneous input context into an executable constraint model for video generation.
  • To address this challenge, we introduce VIS-Ground, which performs Structured Context Abstraction to recover grounded states and cross-context dependencies, Generation Constraints Induction to project relevant dependencies into candidate-specific constraints, and Constrained Video Generation to enforce these constraints through planning, verification, revision, and rendering.
  • Detailed analysis further shows gains across both narrative and knowledge grounding, and reveals remaining challenges in dependency extraction, and faithful realization during video rendering.

Sources (1)

  • [1]VIS-Ground: Video Interactive Storytelling with Contextual Grounding
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 02:31 AM
    Video interactive storytelling enables viewers to actively steer how a video unfolds.
    This leaves a fundamental question underexplored: What context should a generation model ground on during interactive continuation, and how can heterogeneous, unstructured inputs be transformed into such grounding context?

Extractive summary: sentences quoted from the sources.

Related