Pith. sign in

Paper Citation Record · LEDGER

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs

As of 17 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 4 inbound Pith citation observations for arXiv:2507.07990.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.07990 v1

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:32:50.801112Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T23:46:51.654336Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T06:58:05.988167Z

Reference resolution

63 of 63 outbound references displayed

  • verified exact0
  • verified fuzzy37
  • unresolved25
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 246b891c-bda5-48f4-b44b-df2299f204f7 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:44.920094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:44.920094Z digest=sha256:1acfd0553e548cb84585490aaf648f4a3d74bc55cd3dd4928200f3677773bd5a

Observation 7d0f0eb4-1003-46b9-83f2-6d8fd01db424 · outbound

This paper cites Token merging: Your vit but faster.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Token merging: Your vit but faster

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.444166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:32:45.007491Z digest=sha256:2c92fd2b7159b629d79ac8a800ad587bf5bab33979dc0fae38c8693256ccbb18

Observation 1d098822-e46e-43ba-98be-c0a9f7a2e67f · outbound

This paper cites PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:45.136591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:45.136591Z digest=sha256:6d5b6b0d3574296ee30f26c676f98dd0d5969889f749869f4f335de7258bb1df

Observation 0670cd7b-274a-459c-8110-fa130e3564d1 · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Quo vadis, action recognition? a new model and the kinetics dataset

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.411633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:32:45.252215Z digest=sha256:46558864e230182147d6c9adddaf73b1214d89c774188e659c14ab880559545c

Observation 19b8f3b1-b4e2-4330-b2cc-7d25e356367c · outbound

This paper cites Honeybee: Locality-enhanced projector for multimodal llm.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Honeybee: Locality-enhanced projector for multimodal llm

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.392591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:32:45.386772Z digest=sha256:3dd400aa94d712c51108067e412c25f3d28960b2e98b302b3942b2ae384d5765

Observation 37f2bc3d-cab8-4313-b194-437375aacf32 · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.374669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:32:45.517307Z digest=sha256:71a8b24943ee6f74e45b7eba4d13534e219c977d968baaa0c5ea349ec0ef5f15

Observation b7c8df05-599c-413e-b4df-e5ea529487ce · outbound

This paper cites Longvila: Scaling long-context visual language models for long videos.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Longvila: Scaling long-context visual language models for long videos

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.352568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:32:45.618165Z digest=sha256:2487a59733b391e32ec1a0dd081d59b922010041f13035b4fcf0d1d6fe8ac843

Observation e62eee1b-61e5-463a-9b3c-2e7f17199433 · outbound

This paper cites vid-tldr: Training free token merging for light-weight video transformer.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs vid-tldr: Training free token merging for light-weight video transformer

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.331836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:32:45.710034Z digest=sha256:dd0891a5be8933c81a807ec738849376c66b4a841c57ead898958e8262870bd2

Observation 5e573a74-dfd0-4e4c-b898-554117a8c69a · outbound

This paper cites FlashAttention-2: Faster attention with better par- allelism and work partitioning.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs FlashAttention-2: Faster attention with better par- allelism and work partitioning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.306539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:32:45.804392Z digest=sha256:5cfa95802283451855244e6b887f49937016c935a6af080896a96689ca366199

Observation be0085a3-55d6-4f57-a88d-a988eef926be · outbound

This paper cites Fu, Stefano Ermon, Atri Rudra, and Christopher R´e.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Fu, Stefano Ermon, Atri Rudra, and Christopher R´e

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.284441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:32:45.866227Z digest=sha256:864933bfb64d3c5d20ba265917dedc0282b4eef32bc4f31d7b1b4a0588da4a01

Observation ec769102-d618-4f13-a71a-79fbae6b2218 · outbound

This paper cites The Llama 3 Herd of Models.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs The Llama 3 Herd of Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:45.999654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:45.999654Z digest=sha256:515d92f8593313b979303a9ff34c7d937a3c2144151225bfe92e78352457f0ff

Observation 2f038acb-cd79-4778-b467-fe202eed1361 · outbound

This paper cites Masked autoencoders as spatiotemporal learners.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Masked autoencoders as spatiotemporal learners

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.265252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:32:46.168771Z digest=sha256:eb594bfe54bac29ca2cce022d5594116a4b67c7d221fe66b413a1f8b8741108c

Observation 77307511-04a3-486b-80c8-5018115eb2da · outbound

This paper cites Quad trees a data structure for retrieval on composite keys.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Quad trees a data structure for retrieval on composite keys

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.245530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:32:46.273906Z digest=sha256:c46322ed0f21608dd450b2389aad0c72c4dbc5838d5ff5dcfbbafe7bc2606094

Observation 86f94159-5ea6-41b5-a195-dcd31f95a052 · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.222087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:32:46.372576Z digest=sha256:2b778b73413f8398a3c81a3b7ffc96bc9f54ace784deb9c9899ff9a031eb730e

Observation e842939a-c212-4e01-b7fd-dfc93bcc6dee · outbound

This paper cites FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:46.455680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:46.455680Z digest=sha256:02e1ce6b323118890a1514d75879d9ada8d203e971d04a477bd3d269c1fb02ae

Observation f045812b-d2eb-4998-a0d9-76435809e32a · outbound

This paper cites Caching — google ai, 2024.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Caching — google ai, 2024

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.204484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:32:46.509746Z digest=sha256:7b71c609c7578d503a9cdba96ef7123434a66f83331ae78420a820c4d795abe5

Observation dcec1c7b-82b6-43cc-bda6-1910dc2673e5 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:46.642984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:46.642984Z digest=sha256:612bfdeee8183e8f05c0f91fe24fa92ff31b9a798e37176cfe7831a13da70124

Observation e5902d98-fde2-470b-9f0e-c750d81456e5 · outbound

This paper cites Deep residual learning for image recognition.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Deep residual learning for image recognition

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:46.759667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:46.759667Z digest=sha256:75a53f31d348340ca74762b6df1c1d8bb2fdbd600bc6c8301a45f82f7d70548a

Observation 0d4efc77-4d15-4bd9-85ff-deabb6392d3a · outbound

This paper cites Masked autoencoders are scalable vision learners.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Masked autoencoders are scalable vision learners

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.171666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:32:46.855039Z digest=sha256:445d7f93c87cbed9acc16c31342a30b52edc052f191b11654a869f09dca7ade3

Observation d544c777-c601-45fb-a23e-d49639e58b38 · outbound

This paper cites PruneVid: Visual Token Pruning for Efficient Video Large Language Models.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs PruneVid: Visual Token Pruning for Efficient Video Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:46.995554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:46.995554Z digest=sha256:f4719cde00df27d642c6bffd3908131eb9830d07955ebcd8b534cba44c368960

Observation 39e5f196-45bd-42c2-a28a-cb5acb60cc83 · outbound

This paper cites Chat-univi: Unified visual representation em- powers large language models with image and video under- standing.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Chat-univi: Unified visual representation em- powers large language models with image and video under- standing

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.148978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:32:47.097387Z digest=sha256:69fe3fd69f9dab8c718c6aab7615f32fa326bde1a7d9081dd7ce65d1727fb34e

Observation 8d0fd6b9-ffdf-4105-9e21-fbf443daf535 · outbound

This paper cites Needle in a haystack – pressure testing llms.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Needle in a haystack – pressure testing llms

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.130345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:32:47.206402Z digest=sha256:75cd46a11d9533ec4d4d7b3ea0293a78a8e38524bada1dd3c0cd7d6fd863b3d0

Observation e7b38e68-d904-4f5a-b31d-54e76909e05b · outbound

This paper cites Handwritten digit recognition with a back- propagation network.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Handwritten digit recognition with a back- propagation network

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.113062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:32:47.301474Z digest=sha256:01f941c0e0407cf7e1ba3d57baaa59e13a740341d0c57738cd70efef5704d72f

Observation 85685f60-2fb4-4cc9-86df-c99a101a9a0c · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:47.367536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:47.367536Z digest=sha256:636220cb9215c32346e325cc9752ad31b9ea479cb05a6588eef8138630f72087

Observation 8b0f4a29-485c-4221-ac3b-4e0064bc0f6f · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.081714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:32:47.470093Z digest=sha256:ef98bf72d289f831e938a64156b7a7143cc74463e2d0a4552d0f66fe9446da58

Observation 3c22a84a-1f9d-4894-bd7a-ddec8a09b6be · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs VideoChat: Chat-Centric Video Understanding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:47.575406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:47.575406Z digest=sha256:ac51a4463c28108b56f6a0d4f33c28d4fe04efe413d641353fb44f68308b5f36

Observation e2e21bf8-0ea7-422d-92da-9f74675e0312 · outbound

This paper cites Mvbench: A comprehensive multi- modal video understanding benchmark.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Mvbench: A comprehensive multi- modal video understanding benchmark

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.063042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:32:47.618941Z digest=sha256:4df88a1ca0bd6e485d15c14f93c5ba8f4702967538e3bb02b2ef7a15b4e5d241

Observation 77420b45-60dd-4a2a-ae06-5c850c902b2e · outbound

This paper cites Video-llava: Learning united visual repre- sentation by alignment before projection.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Video-llava: Learning united visual repre- sentation by alignment before projection

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.042764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:32:47.732124Z digest=sha256:01bda0db253d47c97d65f0883511e944d8569a17565edab6bea7e36284733c3e

Observation bf4da92a-b536-4555-9aed-2d58320ef6c8 · outbound

This paper cites Visual instruction tuning.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Visual instruction tuning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.017634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:32:47.844102Z digest=sha256:a882c95d38f8c46273eee3c40f174d7f715d2f9972a39a5cd2cded0b5237dc23

Observation 1e5324cc-5866-4d60-a22a-132504cbcc9d · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:47.962155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:47.962155Z digest=sha256:1055f7e10fbe318e5dfc273f342c4ca136e7923986bdfe256ff504c6e90240ec

Observation caeac0d6-d89a-49eb-b1f9-09ba2d3d5223 · outbound

This paper cites Ring atten- tion with blockwise transformers for near-infinite context.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Ring atten- tion with blockwise transformers for near-infinite context

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.989405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:32:48.093263Z digest=sha256:aefab361b61debd7bf3515c3fc9bca9f814fd6efb16bc49ce83655d829a6afd5

Observation ed03f2a1-23f5-4a19-b587-574980db0d02 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:48.187988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:48.187988Z digest=sha256:a2ed22ccaeb74309e40076ed7fdbc81c05ccd9880aea0d5e29ddbe60fffcf7f9

Observation 16437856-d4e5-4593-8847-7fa463692fc7 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long- form video language understanding.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Egoschema: A diagnostic benchmark for very long- form video language understanding

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.947353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:32:48.260745Z digest=sha256:85ca5706e8a968906b177a5d15623400318b8dd2cfbf6c5866f8bef4d3ba124a

Observation 4b8f8318-4e79-4561-9f1e-926a1142560c · outbound

This paper cites Efficiently scaling transformer inference.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Efficiently scaling transformer inference

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.915585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:32:48.312111Z digest=sha256:520549279a1b76f9adf38b83c66e686740f23ed7228b45b9ee23ed76d49d0dbe

Observation 687cec00-dd84-4cf7-ad26-d5f4fadf6b97 · outbound

This paper cites Timechat: A time-sensitive multimodal large language model for long video understanding.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Timechat: A time-sensitive multimodal large language model for long video understanding

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.887336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:32:48.362612Z digest=sha256:7d775297a8749da092bb7d4f96f48bbcff075097ffc890834d93b49307f5bbb9

Observation 83a72fcd-4822-4217-91f4-a59a1769c279 · outbound

This paper cites xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:48.415433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:48.415433Z digest=sha256:a2b05edef2158fabc974cb5f30a61f81a2c9e887b9c29195e5a3c9c0019552ca

Observation 88fcb5da-9be6-487d-b531-a752d5631b86 · outbound

This paper cites The quadtree and related hierarchical data structures.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs The quadtree and related hierarchical data structures

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.860677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:32:48.491626Z digest=sha256:46d9d8d83e8908e1cb28e4a26bd0e0530b4b84c88c9e1806e2665e02b331895d

Observation bbe2353b-5f2b-47e9-b11b-23ed357065d0 · outbound

This paper cites Llava-prumerge: Adaptive token reduction for efficient large multimodal models.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Llava-prumerge: Adaptive token reduction for efficient large multimodal models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:48.573421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:48.573421Z digest=sha256:fa991ac2b169d78600cb97793ea6496cfdb5c3fe607d7a64301c27321796d9cf

Observation 7ee7ef49-1ddd-424c-848b-70815a92fcd9 · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Moviechat: From dense token to sparse memory for long video understanding

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.830881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:32:48.647081Z digest=sha256:cebe7409a001fb12d24ce314d05b47cd155c02325cd9a93233698261754d8ed1

Observation afe23cbc-779d-4373-be94-3f1a48a468b9 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Roformer: Enhanced transformer with rotary position embedding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:48.725013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:48.725013Z digest=sha256:64a91473d9eb87cbfa9edfb2e8b91e98ff6bc0a3ea4a06737bffdcda27422278

Observation 1026ae0b-c4e0-46ae-a06a-306d747fec8e · outbound

This paper cites Efficient quadtree cod- ing of images and video.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Efficient quadtree cod- ing of images and video

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.786169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:32:48.768904Z digest=sha256:fe61a91844b8d8ed6905252c7d5176a2818b7cf3720544674d356fa2480a26a6

Observation 827c1b70-8999-44b3-845f-e9a464a8003c · outbound

This paper cites Overview of the high efficiency video coding (hevc) standard.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Overview of the high efficiency video coding (hevc) standard

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.766030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:32:48.839547Z digest=sha256:297b5e7a0e1a52e438c4ddc4e99eba6a427714e18221a7400d18234b36e25ef3

Observation 1e0ee72e-7428-48f7-b981-87a41aa44524 · outbound

This paper cites Dycoke: Dynamic compression of tokens for fast video large language models.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Dycoke: Dynamic compression of tokens for fast video large language models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.739311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:32:48.887333Z digest=sha256:b4b4e7daf512cd8d4ec8554e8eb47c711dda2e5438674f7df059b4472ad6ee7c

Observation 1928e642-4168-443d-9c79-469fbc97316f · outbound

This paper cites Efficiency of a good but not linear set union algorithm.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Efficiency of a good but not linear set union algorithm

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.619484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:32:48.958480Z digest=sha256:12b4a87be3ca59e695278865e25baeaee2b37bb8f7c815bc923d2944a9d8f650

Observation b26d5fc4-c105-4f25-9614-c9a0da4e329b · outbound

This paper cites GPT-4o System Card.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs GPT-4o System Card

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:49.021249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:49.021249Z digest=sha256:246d3050545529e19de64533cef51ee271173761006e13e93688adec48f07988

Observation c7ba62c1-f17b-40e1-b81c-0cf3e1fadad7 · outbound

This paper cites Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.511623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:32:49.081969Z digest=sha256:b22bbccd2ede1efc7fa3a0f0dadcc7ad083a08d8b1c325c308c06a1cca55e2d5

Observation f213867e-cd85-461f-836d-1cc6c5668753 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs LLaMA: Open and Efficient Foundation Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:49.147349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:49.147349Z digest=sha256:62b7b4cfefc73d4ef5a29037d9985681bad03165c0b3206f35f10e58ef6e29a2

Observation 56b383c2-f3a5-4120-99e2-96f1213e87d4 · outbound

This paper cites Learning spatiotemporal features with 3d convolutional networks.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Learning spatiotemporal features with 3d convolutional networks

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.371957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:32:49.185062Z digest=sha256:e022d5b7527b638774dddefd3074744fe82e7ba4091607f7a3fc8115aa089645

Observation 50d62950-433e-436a-9662-4ce9c855f577 · outbound

This paper cites LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:49.280218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:49.280218Z digest=sha256:e2bb6e0fee30463b2ff51f86d3679c8f698549167a90db3b9fef491f26383937

Observation c5a0000c-cdb7-4c5a-8f89-03803715ded6 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:49.345961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:49.345961Z digest=sha256:357ff06a18801f6c76386f05ce01bb336e5ca32a3971180e1509f38fad959fe1

Observation 0778abac-b905-4d0b-ab9a-8dd7e3b71467 · outbound

This paper cites VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:49.423515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:49.423515Z digest=sha256:f64af6753ce180d45f45b319b89f210df9cd640f60ee772bfbbfb990b9ff5930

Observation 8884aa8f-17de-46c5-99ab-cd1704ee6913 · outbound

This paper cites Sullivan, Gisle Bjontegaard, and Ajay Luthra.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Sullivan, Gisle Bjontegaard, and Ajay Luthra

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.254156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:32:49.484414Z digest=sha256:94b47040635acd58503675838179711aab9a3b5034797574bb7b1dfcfbbcc87b

Observation abe2f266-b145-441e-9c51-376e10e9557d · outbound

This paper cites Longvideobench: A benchmark for long-context interleaved video-language understanding.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Longvideobench: A benchmark for long-context interleaved video-language understanding

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.120020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:32:49.577022Z digest=sha256:fb27620ca2a53b5b7b49737f8920b619a3d63b0c5849675e07b4424a0346fdf9

Observation 94cc60c0-1937-4263-a654-9ed27861d065 · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Next-qa: Next phase of question-answering to explaining temporal actions

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:49.636803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:49.636803Z digest=sha256:a9167323c5710f81c42ec9d362a43ba06b3a0097dd3ca4a74b017116bc656b84

Observation 7560d55b-5148-499d-92b8-ca763da865a8 · outbound

This paper cites Pyramiddrop: Accelerating your large vision-language models via pyramid visual redundancy re- duction.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Pyramiddrop: Accelerating your large vision-language models via pyramid visual redundancy re- duction

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:51.945416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:32:49.707972Z digest=sha256:685fd7659c4397c3cf6947a773065fae26538a9b03c65d02eb4137a450b34925

Observation 45a6af3d-703e-4459-8708-f0a4d375049e · outbound

This paper cites Qwen2 Technical Report.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Qwen2 Technical Report

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:49.783006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:49.783006Z digest=sha256:7d09d824e66d94bfa45d6284de2e7cb836f0e189cc114c22d753aad70b56058d

Observation 5c385049-5d20-4389-bb81-fc98a8695fd0 · outbound

This paper cites PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:49.863938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:49.863938Z digest=sha256:17f088c6c5c4480cc5d638603954a083d8f07aac620c9d7d101d803cbe0c9eea

Observation e182b257-b922-4262-ad32-3d431f68bda0 · outbound

This paper cites Video-llama: An instruction-tuned audio-visual language model for video un- derstanding.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Video-llama: An instruction-tuned audio-visual language model for video un- derstanding

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:51.847754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:32:50.034259Z digest=sha256:b84215fd0af7687e62204a364b7653a020c1195f6f0abe0c2954caaf4724a9b9

Observation 602199ed-14c7-4fbf-9815-939352dc8fa4 · outbound

This paper cites Long Context Transfer from Language to Vision.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Long Context Transfer from Language to Vision

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:50.185274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:50.185274Z digest=sha256:5d86925cbfecc48d16f54ce074d293b2111f56bb71c3312fff2838484d1c3db2

Observation dd7c2631-a2ca-43c3-be08-e90da971380a · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:50.312891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:50.312891Z digest=sha256:cbea77868565aaa3ad8a830bd297cb43da48f6f4ae841335306480903f70602f

Observation f3fd2480-95fc-42f5-8e71-53514a96b734 · outbound

This paper cites Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:50.442322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:50.442322Z digest=sha256:d0a04cc2a4d65cad0ac920e7fc0322ce6dc342d3810c1f92d35711b57254f729

Observation a7e64a4e-366b-4623-9790-6ec1045e116f · outbound

This paper cites Mlvu: Benchmarking multi-task long video understanding.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Mlvu: Benchmarking multi-task long video understanding

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:51.646767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:32:50.667041Z digest=sha256:a305680f1612c926a13d5fd14955b0fbf1693ae585cc00e566c2709284ccd9d0

Observation d05ef8e8-af0a-4b5f-8c80-030feee8b836 · outbound

This paper cites 11 to 14 show the absolute values for the main comparison results.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs 11 to 14 show the absolute values for the main comparison results

Reference 2025

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T18:32:51.576968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:32:50.801112Z digest=sha256:b6b1aade57d5b3c5216abccba7dfba4227eb8e5040e4fc1e91ff0cf18cf73427

Pith citing papers

Observation e99a8d45-4dcb-453b-80d4-725afbba76f0 · inbound

Video Parallel Scaling: Aggregating Diverse Frame Subsets for VideoLLMs cites this paper.

Video Parallel Scaling: Aggregating Diverse Frame Subsets for VideoLLMs Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:36:44.191878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-18T18:35:01.328250Z digest=sha256:4cb0dedf5091078a4d589e3e6789cb2c27ce91a8353d87cba6ed4888faa848ba

Observation 8cb5c6ec-fea4-4edd-8ed6-17874b0c7018 · inbound

VLMaxxing through FrameMogging Training-Free Anti-Recomputation for Video Vision-Language Models cites this paper.

VLMaxxing through FrameMogging Training-Free Anti-Recomputation for Video Vision-Language Models Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:11:13.772797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-08T01:30:15.463051Z digest=sha256:52b45862f752ecc84793e595c55771340f3cd8da728eab69e65f4fa4e976551d

Observation 27a489f0-2b34-46c3-b16b-76908f60e426 · inbound

DynaTok: Temporally Adaptive and Positional Bias-Aware Token Compression for Video-LLMs cites this paper.

DynaTok: Temporally Adaptive and Positional Bias-Aware Token Compression for Video-LLMs Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:58:05.989905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T06:55:01.619441Z digest=sha256:1a4865db589f6f34038dd735ef5d78b64ab4d64008b8640c9eb4778082e40e70

Observation d890eb34-3b8c-4d7c-8973-4ab572bb59d4 · inbound

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models cites this paper.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.654336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.654336Z digest=sha256:85bd960cf7a2a8102e97d1463864283f6b594e83e58b16f758daa1083786df05