REVIEW 1 cited by
Memory Helps, but Confabulation Misleads: Understanding Streaming Events in Videos with MLLMs
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Multimodal large language models (MLLMs) have demonstrated strong performance in understanding videos holistically, yet their ability to process streaming videos-videos are treated as a sequence of visual events-remains underexplored. Intuitively, leveraging past events as memory can enrich contextual and temporal understanding of the current event. In this paper, we show that leveraging memories as contexts helps MLLMs better understand video events. However, because such memories rely on predictions of preceding events, they may contain misinformation, leading to confabulation and degraded performance. To address this, we propose a confabulation-aware memory modification method that mitigates confabulated memory for memory-enhanced event understanding.
Forward citations
Cited by 1 Pith paper
-
Machine Mirages: Defining the Undefined
A conceptual paper defines 34 AI failure modes with formula-like conditions, proposes expectile-value-at-risk quantification, and proves mostly tautological or trivial existence and impossibility results without experiments.
Discussion (0). Sign in to comment.