DSV-Mem: Evaluating Multimodal Memory in Professional Workflows for MLLM Agents
To address these challenges, we introduce DSV-Mem, a benchmark for evaluating Dense Stateful Visual Memory.
ProofPaper ↗
Key points
- Conversational MLLM agents are increasingly expected to assist in professional workflows, from AI research and engineering design to product management and business operations.
- DSV-Mem comprises expert-reviewed scenarios and 1,000 questions across five user-oriented categories (Current State, Past State, Derived State, Change History, and Conflict/Refusal).
- We also introduce a generation harness that produces evaluation suites by decoupling state-transition synthesis from conversation filling.
- Raw conversation/haystack length, OCR, and arithmetic are not the primary bottlenecks; 2) models often fail to verify user premises against prior state updates before answering; 3) increased reasoning effort and memory management methods yield limited gains, whereas state-aware designs prove more effective.
Sources (1)
- [1]DSV-Mem: Evaluating Multimodal Memory in Professional Workflows for MLLM AgentsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 10:30 AM
To address these challenges, we introduce DSV-Mem, a benchmark for evaluating Dense Stateful Visual Memory.
Conversational MLLM agents are increasingly expected to assist in professional workflows, from AI research and engineering design to product management and business operations.
Extractive summary: sentences quoted from the sources.