Pith. sign in

Paper Citation Record · LEDGER

MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2307.16449.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2307.16449 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:09:07.794518Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T01:36:44.033728Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 74c5ed68-2940-4fd2-9e99-9bb5617b6b00 · inbound

PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning cites this paper.

PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:21:58.021092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T20:21:57.873354Z digest=sha256:99bdba88f9fd1fd49e06419cd5dbc503bf47e623e02c1f69453b888cefffca0e

Observation 708ac9fc-961d-46f0-b59a-461ad3e71b9f · inbound

MLVU: Benchmarking Multi-task Long Video Understanding cites this paper.

MLVU: Benchmarking Multi-task Long Video Understanding MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:55:26.439473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T19:55:26.333923Z digest=sha256:641cf9703209ed688c9939d9eb4e241a1e195b1478982a942ffad15d7d453680

Observation a97ff8ec-0d4b-486d-9398-4b33d2e93a95 · inbound

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs cites this paper.

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-11T02:44:53.638107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T02:44:53.284345Z digest=sha256:3a9b44fe7ae971d9f70b65da71fe902624f51cc041682f42cfa968c342e969ae

Observation 2418df75-a193-48f6-88f8-3fdb91d97860 · inbound

LVBench: An Extreme Long Video Understanding Benchmark cites this paper.

LVBench: An Extreme Long Video Understanding Benchmark MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:55:30.104117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T11:55:30.048525Z digest=sha256:9374194d331b437a3837c6b8578c620f225261dfa1a816cf2f206f781a17f2f2

Observation 185a7090-b23a-4513-a028-018679817359 · inbound

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output cites this paper.

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 135

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:46:28.677521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T10:46:28.447347Z digest=sha256:1f19870bb141b64cc789e4cf7cb2d34d5b690b50b6d354998490617f82e0ae38

Observation e7b9d41c-affe-4ba0-9377-5a65622b0cfc · inbound

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding cites this paper.

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:02:37.587807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T06:01:00.775721Z digest=sha256:9244ae297ab95b4ceda044569b382e7472eec859b71142ff086899551f90654f

Observation c87f1033-5839-4a62-b3a7-7730ec1d7bda · inbound

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos cites this paper.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:32:41.261120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:86c9914f2c3f13032e6f7a2e9415b7d551a339b940439451e1129e49a0dbd135

Observation 563dd0cc-1e6d-479d-beac-81df15334841 · inbound

SmolVLM: Redefining small and efficient multimodal models cites this paper.

SmolVLM: Redefining small and efficient multimodal models MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:23:51.770625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T20:23:50.552549Z digest=sha256:bd6409c4084220a6da00a2ab735ebd4be3f1f0f203a403c9fb6e13f4d53d5b6a

Observation 2042abbd-d90d-445f-9c54-2a82c3e8af38 · inbound

EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language Models cites this paper.

EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language Models MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:07.794518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:09:07.794518Z digest=sha256:e0be4b7eace74e8ba921a45c9f813f22442c2265d86e05e7410fd9fa797c7871

Observation 7e5fa39a-d49d-44f5-9ae8-4c8effec6e48 · inbound

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding cites this paper.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.967687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.967687Z digest=sha256:323c4b855951cc2962e7dfd3384b7bf2cbc46689ca409b7d93790c380da4b432

Observation 82ed5029-2983-493e-ab0a-c3a858ba68dc · inbound

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification cites this paper.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:07.554817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:07.554817Z digest=sha256:a03c4e8e3a3c1bb8fb6b496f279da7ce4eeff953afe31e22cb1ae9579864a8df

Observation b25ca4dc-5cdc-466e-a600-f35ebdf1f7d5 · inbound

Task-Aware KV Compression For Cost-Effective Long Video Understanding cites this paper.

Task-Aware KV Compression For Cost-Effective Long Video Understanding MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:38.567659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:38.567659Z digest=sha256:579facb6d322956dfa7c4162871239636fdb7818adc80f9feab7320eff3c5ea1

Observation 72519638-fa77-49ec-9ce6-613012807379 · inbound

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs cites this paper.

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T22:24:32.112310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:24:32.112310Z digest=sha256:f57b2ffeb8b7510258674e84d2e2e166e7eb086971a352c6ba63c82ae58826de

Observation 00d8e688-0c33-48cb-ab53-274cf486fbfb · inbound

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding cites this paper.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:54.291657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:54.291657Z digest=sha256:725e7abfb7546ea5997cc7e1172e9150cfee92170ad8367130bbf6212687a2fe

Observation 102ccd94-8ead-4ada-9021-3f44632e9396 · inbound

Infinite Video Understanding cites this paper.

Infinite Video Understanding MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.357356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.357356Z digest=sha256:674a6e20d21cdee7dc2068d25cd638c7107c6bb83c576a6ce76cea31142c3c1c

Observation 19bf8d62-97ee-45f3-899e-35f5b9155968 · inbound

A Benchmark for Omni-Modal Reasoning in Long Videos cites this paper.

A Benchmark for Omni-Modal Reasoning in Long Videos MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T15:26:15.102987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:26:15.102987Z digest=sha256:317620e9c21dd40cd666fecf04f6f9508bfeb231f9b91a3cb6404aaa7a047e6e

Observation 86eb29da-14cc-4620-8d25-37ab3170b275 · inbound

Less Data, Faster Convergence: Goal-Driven Data Optimization for Multimodal Instruction Tuning cites this paper.

Less Data, Faster Convergence: Goal-Driven Data Optimization for Multimodal Instruction Tuning MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-14T22:18:56.115559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T22:18:56.115559Z digest=sha256:09b3393dc6ad98882fd7b484017eaf95d9b89a9f92c5e18c2e3052c0d2807ff4

Observation 9cf2f7bb-d52c-4ac8-b6b2-80c9624b37db · inbound

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark cites this paper.

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:08:04.369292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T22:05:07.326202Z digest=sha256:7af2f30b3f31ac690609d3078e227f07efa5e04e68e7a3f94c8e8cc90ddb41af

Observation 3b51cd7a-8a77-4cb9-82bd-64ac814881d3 · inbound

Diagnosing Long-Video Quantitative Reasoning in Multimodal LLMs via Enumeration and Counting cites this paper.

Diagnosing Long-Video Quantitative Reasoning in Multimodal LLMs via Enumeration and Counting MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-13T15:29:31.834566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T15:29:31.834566Z digest=sha256:de4d2abf032682d30ed0b745fafc6589717697f750a22a642ea65b4d21f7a494

Observation 7aeb8240-0e45-4502-b046-84e21d5d72cc · inbound

PyraVid: Hierarchical Multimodal Memory for Long-Horizon Video Reasoning cites this paper.

PyraVid: Hierarchical Multimodal Memory for Long-Horizon Video Reasoning MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:13:24.841282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-20T15:12:00.408851Z digest=sha256:58cabc316c2f3548c68c3fd9e3f9a811ce115017adacc402623b9f085334c10b

Observation 6de1d70f-fa3b-4953-a909-62890ec284c5 · inbound

VTI-CoT: Visual-Textual Interleaved Chain of Thought for Video Reasoning cites this paper.

VTI-CoT: Visual-Textual Interleaved Chain of Thought for Video Reasoning MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:46:56.837302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T01:52:44.785582Z digest=sha256:43ba046c61f68271e37e01a687ed53d241bb85889cc26ff1e526cb3525100e27

Observation bd657a2e-9658-421b-9a36-6bdfd12428c0 · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 151

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:39:37.661744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:acbbb30665fc67b149f5b50297f4e205f8314ff86a89c86c397c7afb90f18f6d

Observation 8be66e76-6987-4810-81f7-d6d89584efc9 · inbound

Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction cites this paper.

Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T07:54:22.343008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T07:48:01.719339Z digest=sha256:19595662a3dcfb8908d5696c6a45b157093cf9ee63fc61c1394a8a84fe2d6541

Observation 1b204ed7-b544-4d00-b49a-40961db8e10f · inbound

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences cites this paper.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:086e14ed56691f56b592ad7dbd5bc6aa57980e900999e8c6c1cec448aa2dd904

Observation 21063b64-1987-4fb9-add8-35b26a601fe8 · inbound

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents cites this paper.

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 108

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:36:44.035222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T01:26:59.421158Z digest=sha256:b6ec04e76b5213872d68338e419567ff731e48a9707b6187e6444eeec066c645