Pith. sign in

Paper Citation Record · LEDGER

TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2409.01156.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.01156 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:40:23.724449Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T07:34:22.143661Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3b10fefc-2f69-4e0c-ab17-c914a975ecb5 · inbound

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models cites this paper.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T14:40:23.724449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:40:23.724449Z digest=sha256:e46bd1f06faad9e1f690fe38ac10c33437770f90de17a10080d8bea20adab9b0

Observation e0aeff09-969a-4cb3-bafb-926f1a79b5b7 · inbound

HEIE: MLLM-Based Hierarchical Explainable AIGC Image Implausibility Evaluator cites this paper.

HEIE: MLLM-Based Hierarchical Explainable AIGC Image Implausibility Evaluator TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:18.526987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:18.526987Z digest=sha256:6bf5f2eb8d882c785bcd741892838360da8f757d256053b7fdac9b233970c9ea

Observation c22e7847-d52e-41c7-b0c7-009d22ecfc0f · inbound

LinVT: Empower Your Image-level Large Language Model to Understand Videos cites this paper.

LinVT: Empower Your Image-level Large Language Model to Understand Videos TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.416990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.416990Z digest=sha256:b868c92cbfa6f3018f5802028ead48e0fbab49c6847cc75c7423af15c10d7334

Observation 52f69615-1a5e-44bc-b7fc-b0b73babfd78 · inbound

AdaTP: Attention-Debiased Token Pruning for Video Large Language Models cites this paper.

AdaTP: Attention-Debiased Token Pruning for Video Large Language Models TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:08:09.248749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:08:09.248749Z digest=sha256:2f92b5ecd888f1d03ec9569ccfab2dab49bf411eed98592598a56da064b517d9

Observation 20943e42-6839-4b62-ad48-600cbe32f003 · inbound

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval cites this paper.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:38.599960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:38.599960Z digest=sha256:6663f013b830e2e79a0790d74f297c6b871361e036799050674cf1fbad52c5da

Observation c1ba48df-e473-4e59-bdf4-1306f72d7a79 · inbound

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs cites this paper.

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T22:24:31.751666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:24:31.751666Z digest=sha256:253e6609a722b55a9addb32cce168244bc55f1651d34ea1cef3ae391f329d61c

Observation 86d26e7a-4e49-4ef1-8d1a-ee965f131884 · inbound

FastVGGT: Training-Free Acceleration of Visual Geometry Transformer cites this paper.

FastVGGT: Training-Free Acceleration of Visual Geometry Transformer TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-15T23:36:06.973826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T23:36:06.870759Z digest=sha256:2bf3ceaa2ba728398dc90c0daa89505f0f387cf478189075318dbbf0cb140d4c

Observation 7ed984dd-c281-4f1b-bf8f-54d493bc3784 · inbound

TCMA: Text-Conditioned Multi-granularity Alignment for Drone Cross-Modal Text-Video Retrieval cites this paper.

TCMA: Text-Conditioned Multi-granularity Alignment for Drone Cross-Modal Text-Video Retrieval TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T10:22:45.895680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:22:45.895680Z digest=sha256:be1fbe6c7ad291c668c3c4bbe3f33c0e51bdd778e0d214cf6a4aaeff79074d70

Observation 98511968-0902-473b-a2c6-0aeff1f6d256 · inbound

LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding cites this paper.

LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:01:33.518670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T20:01:31.129959Z digest=sha256:86e6ea45abc5c8da88ba72a4c16e62142b8b5a06483f3847e6e7a071b2b7065b

Observation e6373ec4-0774-4ee1-8c9a-40bb5e22f222 · inbound

Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models cites this paper.

Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:26:27.050959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T18:25:21.621268Z digest=sha256:e463faad9a20a79c3a198d49ee9ce52e507cb99739a219ab06fcd1d8eb026e1e

Observation 4d8556ba-eb10-4551-832a-2fafc5d97a49 · inbound

Understanding the Performance Plateau in Text-to-Video Retrieval: A Comprehensive Empirical and Linguistic Analysis cites this paper.

Understanding the Performance Plateau in Text-to-Video Retrieval: A Comprehensive Empirical and Linguistic Analysis TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval

Reference 67

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T15:06:09.663380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T15:05:37.964883Z digest=sha256:7da31f6519bcc5f49198ef644b7d8f7087fcbb97f4d7b5b2ff7fed48f5fe46cc

Observation 405e9bd1-0db9-4bfc-8bf8-6a6e5c2d1341 · inbound

Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs cites this paper.

Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:34:22.145236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T07:24:59.159037Z digest=sha256:8517cc32ed50a95647134f685320a31aa45423b868d91daaf4cc36460f32cb01

Observation 54e0cb18-43af-407f-b79f-dc31dac84aca · inbound

STAC: Selective Spatiotemporal Aggregation and Compression for Video Reasoning Segmentation cites this paper.

STAC: Selective Spatiotemporal Aggregation and Compression for Video Reasoning Segmentation TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval

Reference 65

Resolution
unresolved
no resolver link, observed 2026-07-12T06:06:47.233814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:06:47.233814Z digest=sha256:1d2466f86f067347ec3d25553acb45a66f282f270e7191c848ebd01f33b3cfd8

Observation fe4e68f3-7808-4951-9552-1dcae442deca · inbound

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation cites this paper.

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-31T23:59:17.042425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:59:17.042425Z digest=sha256:cd923694d99ff7b9e615fbfc08676a46ec246ad2da462336aca164dc4ef47fc2

Observation bd2233ae-0764-4622-ba01-0bc71e99655c · inbound

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation cites this paper.

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T04:01:48.573191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:01:48.573191Z digest=sha256:645bd697aa85dd7dafe00928489d06efa779aca8e4c99b054cab37dd039231f5