Pith. sign in

Paper Citation Record · LEDGER

MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2307.16449.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2307.16449 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:09:07.794518Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T01:36:44.033728Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 74c5ed68-2940-4fd2-9e99-9bb5617b6b00 · inbound

PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning cites this paper.

PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:21:58.021092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T20:21:57.873354Z digest=sha256:5b02daf6bd1c51cf2d599447fed4f953a769a1c62e0290192b4de47483e1f5cd

Observation 708ac9fc-961d-46f0-b59a-461ad3e71b9f · inbound

MLVU: Benchmarking Multi-task Long Video Understanding cites this paper.

MLVU: Benchmarking Multi-task Long Video Understanding MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:55:26.439473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T19:55:26.333923Z digest=sha256:ca78dea6282f55ca2d620349a422622f7546988d3a24a51f4fa6bb62b0209600

Observation a97ff8ec-0d4b-486d-9398-4b33d2e93a95 · inbound

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs cites this paper.

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-11T02:44:53.638107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T02:44:53.284345Z digest=sha256:c755b311acbd15497b497296dfe622192525d8d282d5c27fedfc6387597042b9

Observation 2418df75-a193-48f6-88f8-3fdb91d97860 · inbound

LVBench: An Extreme Long Video Understanding Benchmark cites this paper.

LVBench: An Extreme Long Video Understanding Benchmark MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:55:30.104117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T11:55:30.048525Z digest=sha256:3f9385fdf7e2dafba0c34a0b34610c411745f88cf4a24b9c40e605df3b49d3be

Observation 185a7090-b23a-4513-a028-018679817359 · inbound

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output cites this paper.

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 135

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:46:28.677521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T10:46:28.447347Z digest=sha256:c869c71dfc9f49eeea91f5eccf5b3a45393e5de1ea1eb1bfb6afb6b2f3892caf

Observation e7b9d41c-affe-4ba0-9377-5a65622b0cfc · inbound

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding cites this paper.

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:02:37.587807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T06:01:00.775721Z digest=sha256:e4644149e851c493a8de83809982e0ce47a8386a4a0599d43b856b1c370fe839

Observation c87f1033-5839-4a62-b3a7-7730ec1d7bda · inbound

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos cites this paper.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:32:41.261120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:aec2325828bf0c6182ebffa7265d5da8240b61ddfe66bfade5e79a3a013cdc97

Observation 563dd0cc-1e6d-479d-beac-81df15334841 · inbound

SmolVLM: Redefining small and efficient multimodal models cites this paper.

SmolVLM: Redefining small and efficient multimodal models MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:23:51.770625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T20:23:50.552549Z digest=sha256:c7e039fd57cc139ea35972679e0367dcfa9e26dea16abfcc441513a1a4509011

Observation 2042abbd-d90d-445f-9c54-2a82c3e8af38 · inbound

EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language Models cites this paper.

EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language Models MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:07.794518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:09:07.794518Z digest=sha256:db00d64d5bb7d993744e8b5a8c212f0a61bd71f9ae289771e51321d129448639

Observation 7e5fa39a-d49d-44f5-9ae8-4c8effec6e48 · inbound

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding cites this paper.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.967687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.967687Z digest=sha256:323c4b855951cc2962e7dfd3384b7bf2cbc46689ca409b7d93790c380da4b432

Observation 82ed5029-2983-493e-ab0a-c3a858ba68dc · inbound

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification cites this paper.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:07.554817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:07.554817Z digest=sha256:a03c4e8e3a3c1bb8fb6b496f279da7ce4eeff953afe31e22cb1ae9579864a8df

Observation b25ca4dc-5cdc-466e-a600-f35ebdf1f7d5 · inbound

Task-Aware KV Compression For Cost-Effective Long Video Understanding cites this paper.

Task-Aware KV Compression For Cost-Effective Long Video Understanding MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:38.567659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:38.567659Z digest=sha256:579facb6d322956dfa7c4162871239636fdb7818adc80f9feab7320eff3c5ea1

Observation 72519638-fa77-49ec-9ce6-613012807379 · inbound

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs cites this paper.

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T22:24:32.112310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:24:32.112310Z digest=sha256:ec37d77d7efb1281f5999fee9a66349b31afe01724e0900841a1c12d432fab53

Observation 00d8e688-0c33-48cb-ab53-274cf486fbfb · inbound

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding cites this paper.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:54.291657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:54.291657Z digest=sha256:725e7abfb7546ea5997cc7e1172e9150cfee92170ad8367130bbf6212687a2fe

Observation 102ccd94-8ead-4ada-9021-3f44632e9396 · inbound

Infinite Video Understanding cites this paper.

Infinite Video Understanding MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.357356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.357356Z digest=sha256:674a6e20d21cdee7dc2068d25cd638c7107c6bb83c576a6ce76cea31142c3c1c

Observation 19bf8d62-97ee-45f3-899e-35f5b9155968 · inbound

A Benchmark for Omni-Modal Reasoning in Long Videos cites this paper.

A Benchmark for Omni-Modal Reasoning in Long Videos MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T15:26:15.102987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:26:15.102987Z digest=sha256:317620e9c21dd40cd666fecf04f6f9508bfeb231f9b91a3cb6404aaa7a047e6e

Observation 86eb29da-14cc-4620-8d25-37ab3170b275 · inbound

Less Data, Faster Convergence: Goal-Driven Data Optimization for Multimodal Instruction Tuning cites this paper.

Less Data, Faster Convergence: Goal-Driven Data Optimization for Multimodal Instruction Tuning MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-14T22:18:56.115559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T22:18:56.115559Z digest=sha256:09b3393dc6ad98882fd7b484017eaf95d9b89a9f92c5e18c2e3052c0d2807ff4

Observation 9cf2f7bb-d52c-4ac8-b6b2-80c9624b37db · inbound

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark cites this paper.

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:08:04.369292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T22:05:07.326202Z digest=sha256:5c47173bff5c31bf01606288879a1d51c9270e468b0535b523d3db5aa94328ad

Observation 3b51cd7a-8a77-4cb9-82bd-64ac814881d3 · inbound

Diagnosing Long-Video Quantitative Reasoning in Multimodal LLMs via Enumeration and Counting cites this paper.

Diagnosing Long-Video Quantitative Reasoning in Multimodal LLMs via Enumeration and Counting MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-13T15:29:31.834566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T15:29:31.834566Z digest=sha256:de4d2abf032682d30ed0b745fafc6589717697f750a22a642ea65b4d21f7a494

Observation 7aeb8240-0e45-4502-b046-84e21d5d72cc · inbound

PyraVid: Hierarchical Multimodal Memory for Long-Horizon Video Reasoning cites this paper.

PyraVid: Hierarchical Multimodal Memory for Long-Horizon Video Reasoning MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:13:24.841282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T15:12:00.408851Z digest=sha256:dc6f06a596d18ee6ae3cfd8e81b833df2bb28a6c4aaf32186952c4b4227ff8b5

Observation 6de1d70f-fa3b-4953-a909-62890ec284c5 · inbound

VTI-CoT: Visual-Textual Interleaved Chain of Thought for Video Reasoning cites this paper.

VTI-CoT: Visual-Textual Interleaved Chain of Thought for Video Reasoning MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:46:56.837302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T01:52:44.785582Z digest=sha256:94ce92c6d86bce18c6fb6556b6aed9ec755facbc9f7dea4f0e37abfa38da09d5

Observation bd657a2e-9658-421b-9a36-6bdfd12428c0 · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 151

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:39:37.661744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:ef31f186197f41b90fc1039e990bf82d61804091eb7302afc166f58fda3a7b3a

Observation 8be66e76-6987-4810-81f7-d6d89584efc9 · inbound

Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction cites this paper.

Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T07:54:22.343008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T07:48:01.719339Z digest=sha256:c530709f9751d17b351954cbcb1244d0d57475f066707d86800950aaff80ca66

Observation 1b204ed7-b544-4d00-b49a-40961db8e10f · inbound

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences cites this paper.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:086e14ed56691f56b592ad7dbd5bc6aa57980e900999e8c6c1cec448aa2dd904

Observation 21063b64-1987-4fb9-add8-35b26a601fe8 · inbound

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents cites this paper.

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 108

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:36:44.035222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-10T01:26:59.421158Z digest=sha256:2183393de207709f2cf0c6c5cb7d360a954c7ecc44be71aeb1d7003f672ea6c0