Pith. sign in

Paper Citation Record · LEDGER

ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understanding

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2412.20504.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.20504 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:37:22.035335Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T20:01:33.546041Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 26926a99-6fe9-4763-ae56-44f168a64c7d · inbound

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders cites this paper.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understanding

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:22.035335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:22.035335Z digest=sha256:df4908fe91c4cb9da709040ae25ed4e2f47ac4045b7ba2a53b3eede738f049ea

Observation afa52dc7-5d5a-4b51-8d3d-97c898209423 · inbound

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification cites this paper.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understanding

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:10.091131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:10.091131Z digest=sha256:5d7e9e943e1ade9034e09f1087968ff1e05a55fac384719813b037df9f808eb9

Observation 242fa570-3d84-43c4-8f27-c01f37a67a4e · inbound

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams cites this paper.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understanding

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:07.383398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:07.383398Z digest=sha256:092c2d03a361c66b0e82a9425a1fd3bfc620e97c2fb6a42140fc0a0db2fe775b

Observation 3fbb95af-caf3-4d79-a8bc-191d95e8f8e2 · inbound

LongAnimation: Long Animation Generation with Dynamic Global-Local Memory cites this paper.

LongAnimation: Long Animation Generation with Dynamic Global-Local Memory ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T20:44:03.079113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:44:03.079113Z digest=sha256:2c345be56ab4e05b4259ea9e227d315a048189e215f963d73193f043888f7dfd

Observation abc553c2-7d53-4d73-b1c7-5da5c5b76a11 · inbound

A Glimpse to Compress: Dynamic Visual Token Pruning for Large Vision-Language Models cites this paper.

A Glimpse to Compress: Dynamic Visual Token Pruning for Large Vision-Language Models ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T05:37:35.445827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:37:35.445827Z digest=sha256:5716d435eadadeecf0ff0652534b475183e5eb6cae2981c0cfcb38a44ae3c039

Observation dd7380db-5064-4d80-a921-2eb87221f7ce · inbound

A Survey on Video Temporal Grounding with Multimodal Large Language Model cites this paper.

A Survey on Video Temporal Grounding with Multimodal Large Language Model ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understanding

Reference 160

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:18.195646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:18.195646Z digest=sha256:973374078508444dcbf2a2e6f59c5cdc36de676c1f647d1a9becc715a418db40

Observation 4a3f20b1-7c31-4b8a-b082-b5533de0464d · inbound

LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding cites this paper.

LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understanding

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:01:33.550637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T20:01:31.129959Z digest=sha256:fdc8471275c445f91652b00414d9def5ae30b35c3a17fa3dc7d93b76f4934b5b

Observation 0ad13aa1-3d18-4d99-9598-f45c711e9cf3 · inbound

Echoes Over Time: Unlocking Length Generalization in Video-to-Audio Generation Models cites this paper.

Echoes Over Time: Unlocking Length Generalization in Video-to-Audio Generation Models ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understanding

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:56:33.542684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T19:53:18.200223Z digest=sha256:be304c041e383da7a651fd05a5d028b996bde087bd3478d9bcc8efaa9120000b