🎉 Accepted at EMNLP 2026

EgoVoice
Proactive Spoken Assistance from
Egocentric Multimodal Streams

Heeseung Kim
Department of AI, University of Seoul

A framework for training and evaluating proactive egocentric spoken assistants that observe first-person video and audio and provide timely spoken guidance without being explicitly asked.

Paper Soon Code Nov 2026 Dataset Nov 2026

Code, models, and dataset will be released in November 2026.

System Overview

Figure 1. An AR assistant receives first-person video and audio, decides whether to remain silent or intervene, and provides spoken guidance at the right moment.

What is EgoVoice?

Wearable augmented reality (AR) assistants are moving toward continuous real-world interaction, where they perceive the user's activity through first-person video and audio and provide timely spoken guidance without being explicitly asked. We introduce EgoVoice, a framework for training and evaluating proactive egocentric spoken assistants. From HoloAssist video recordings of real human instructors, we construct clean audio streams through source separation, text normalization, and speech resynthesis. We fine-tune an omni-modal LLM and further improve its proactive intervention behavior with direct preference optimization. EgoVoice yields clear improvements in intervention timing, content relevance, and human preference over the zero-shot backbone.

Dataset Samples

Full-length sessions from the benchmark set, showing the complete task with all instructor interventions.

User speech (top subtitle) Assistant speech (bottom subtitle)

Model Comparison

Sessions are randomly selected from the benchmark set and cropped to approximately 60 seconds (similar to training window length).

User speech (top subtitle) Assistant speech (bottom subtitle)