ResearchResearch paperSpeech & Audio · Large Language Models1 source · Oct 6, 2026

PERSIST: Who-What-When Memory Across Sessions for Full-Duplex Spoken Dialogue

We present PERSIST, a persistent memory system for multi-session, multi-speaker spoken dialogue that explicitly models Who, What, and When.

Key points

  • Modern voice assistants may be shared by multiple users and should be able to answer questions about earlier conversations such as "When did I originally plan to leave?" or adapt their behavior to individual users based on past interactions.
  • PERSIST structures cross-session histories into readable event records and retrieves them with a 3W joint scoring mechanism that combines semantic content, acoustic speaker identity, and temporal state.
  • For real-time full-duplex interaction, PERSIST further reuses intermediate representations from the dialogue backbone, avoiding query-audio re-encoding and reducing retrieval latency from 578.42 ms to 7.03 ms.
  • We also introduce SpokenTrace, a diagnostic benchmark that factorizes evaluation along memory tasks and speaker-query types, exposing failures in recall, speaker attribution, and temporal-state tracking.

Sources (1)

  • [1]PERSIST: Who-What-When Memory Across Sessions for Full-Duplex Spoken Dialogue
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 04:21 AM
    We present PERSIST, a persistent memory system for multi-session, multi-speaker spoken dialogue that explicitly models Who, What, and When.
    Modern voice assistants may be shared by multiple users and should be able to answer questions about earlier conversations such as "When did I originally plan to leave?" or adapt their behavior to individual users based on past interactions.

Extractive summary: sentences quoted from the sources.

Related