Pith. sign in

REVIEW 2 cited by

OmniMMI: A Comprehensive Multi-modal Interaction Benchmark in Streaming Video Contexts

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.22952 v1 pith:Q5ZTQYBV submitted 2025-03-29 cs.CV

classification cs.CV
keywords multi-modalstreamingvideocontextsomnimmibenchmarkcomprehensivedesigned
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The rapid advancement of multi-modal language models (MLLMs) like GPT-4o has propelled the development of Omni language models, designed to process and proactively respond to continuous streams of multi-modal data. Despite their potential, evaluating their real-world interactive capabilities in streaming video contexts remains a formidable challenge. In this work, we introduce OmniMMI, a comprehensive multi-modal interaction benchmark tailored for OmniLLMs in streaming video contexts. OmniMMI encompasses over 1,121 videos and 2,290 questions, addressing two critical yet underexplored challenges in existing video benchmarks: streaming video understanding and proactive reasoning, across six distinct subtasks. Moreover, we propose a novel framework, Multi-modal Multiplexing Modeling (M4), designed to enable an inference-efficient streaming model that can see, listen while generating.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models

    cs.CV 2025-07 conditional novelty 6.0 of 10

    ProactiveVideoQA is a benchmark for proactive video question answering, and the proposed PAUC metric jointly scores response timing and content, claiming better alignment with human preferences.

  2. Aero Realtime: Fully Aligned Input-Output Streams for Low-Latency Streaming Multimodal Generation

    cs.AI 2026-08 conditional novelty 5.0 of 10

    Aero Realtime aligns continuous video, audio, and text output on one 80ms grid, letting a 4B model generate lexical tokens or silence in a single stream while reusing the KV cache.

Pith tools