EgoVoice: Proactive Spoken Assistance from Egocentric Multimodal Streams
We introduce EgoVoice, a framework for training and evaluating proactive egocentric spoken assistants.
ProofPaper ↗
Key points
- Wearable augmented reality (AR) assistants are moving toward continuous real-world interaction, where they perceive the user's activity through first-person video and audio and provide timely spoken guidance without being explicitly asked.
- While proactive video assistants, spoken dialog systems, and egocentric task understanding have each advanced rapidly, existing systems do not address the joint problem of deciding when to speak and what to say from continuous first-person streams.
- From HoloAssist video recordings of real human instructors, we construct clean audio streams through source separation and speech resynthesis, and convert each video session into a format where the model must decide at each moment whether to remain silent or provide spoken guidance.
- Experiments across closed and open-source models show that existing systems rarely produce well-timed, meaningful proactive interventions, while EgoVoice yields clear improvements in intervention timing, content relevance, and human preference over the zero-shot backbone.
Sources (1)
- [1]EgoVoice: Proactive Spoken Assistance from Egocentric Multimodal StreamsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 04:24 PM
We introduce EgoVoice, a framework for training and evaluating proactive egocentric spoken assistants.
Wearable augmented reality (AR) assistants are moving toward continuous real-world interaction, where they perceive the user's activity through first-person video and audio and provide timely spoken guidance without being explicitly asked.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 7, 2026Enabling Preference-driven Unlearning in Few-step Distilled Text-to-Image Diffusion Models
- Oct 7, 2026From Digital Human Interactions to Physics-Based Humanoid Skills: Physics-Grounded Post-Training of Interaction Generators
- Oct 7, 2026How to train your model organism
- Oct 6, 2026CM-DPO: Constraint-Margin Direct Preference Optimization for LLM Planning