Pith. sign in

REVIEW 2 cited by

Memory-efficient Streaming VideoLLMs for Real-time Procedural Video Understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.13915 v1 pith:SC64E5YQ submitted 2025-04-10 cs.CV

classification cs.CV
keywords providellmtokensobservationsproceduralstreamingvideocachelong-term
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce ProVideLLM, an end-to-end framework for real-time procedural video understanding. ProVideLLM integrates a multimodal cache configured to store two types of tokens - verbalized text tokens, which provide compressed textual summaries of long-term observations, and visual tokens, encoded with DETR-QFormer to capture fine-grained details from short-term observations. This design reduces token count by 22x over existing methods in representing one hour of long-term observations while effectively encoding fine-granularity of the present. By interleaving these tokens in our multimodal cache, ProVideLLM ensures sub-linear scaling of memory and compute with video length, enabling per-frame streaming inference at 10 FPS and streaming dialogue at 25 FPS, with a minimal 2GB GPU memory footprint. ProVideLLM also sets new state-of-the-art results on six procedural tasks across four datasets.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Seen, Said, or Forgotten? A Causal Audit of Visual KV Memory Across Dialog Turns

    cs.CV 2026-07 conditional novelty 7.0 of 10

    Current attention is not a safe-forgetting signal for visual KV memory; damage from eviction concentrates in visually dependent turns, and only explicitly verbalized facts are reliably rescued by assistant text.

  2. Diffractive electroproduction of light vector particles: leading Fock-state contribution in the presence of significant higher Fock-state effects

    hep-ph 2025-08 unverdicted novelty 5.0 of 10

    The paper claims the leading quark-antiquark approximation in the color dipole model only matches HERA data for rho/gamma above Q^2 of 20 GeV^2 and for phi above Q^2 of 10 GeV^2, unlike J/psi.

Pith tools