ResearchResearch paperMultimodal Models · Speech & Audio1 source · Oct 7, 2026

InterView-C: A Synchronized Multimodal Corpus of VR Avatar-Mediated Survey Interviews

We present InterView-C, a German multimodal corpus of 27 survey interviews conducted entirely in virtual reality, with both interlocutors represented by avatars.

Key points

  • The corpus aligns spoken interaction with synchronized behavioral data, including gaze, head and body movement, facial behavior, hand and finger tracking.
  • InterView-C therefore provides word-timed and manually post-edited verbatim transcripts for all 54 recordings, interview-item timings, questionnaire responses and negation cue and scope annotations for 1,422 sentences, 1,398 of them doubly annotated (α=0.87 for cues; α=0.81 for scopes).
  • We demonstrate both challenges empirically: nine open-weight ASR systems disproportionately misrecognize short closed answers and number words, while negation models trained on existing corpora show lower and highly variable performance on our transcribed interviews than a model trained on the InterView-C annotations.
  • InterView-C thus enables linguistic analyses of spoken interaction while retaining their alignment with rich multimodal behavior.

Sources (1)

  • [1]InterView-C: A Synchronized Multimodal Corpus of VR Avatar-Mediated Survey Interviews
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 02:23 PM
    We present InterView-C, a German multimodal corpus of 27 survey interviews conducted entirely in virtual reality, with both interlocutors represented by avatars.
    The corpus aligns spoken interaction with synchronized behavioral data, including gaze, head and body movement, facial behavior, hand and finger tracking.

Extractive summary: sentences quoted from the sources.

Related