Pith. sign in

Paper Citation Record · LEDGER

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs

As of 10 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 4 inbound Pith citation observations for arXiv:2507.07990.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.07990 v1

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:32:50.801112Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T23:46:51.654336Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T06:58:05.988167Z

Reference resolution

63 of 63 outbound references displayed

  • verified exact0
  • verified fuzzy37
  • unresolved25
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 246b891c-bda5-48f4-b44b-df2299f204f7 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:44.920094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:44.920094Z digest=sha256:169a1c3fafdc5417bbb791b530271cefcdcf61fe9f79cc13e2abb9afec87ff8f

Observation 7d0f0eb4-1003-46b9-83f2-6d8fd01db424 · outbound

This paper cites Token merging: Your vit but faster.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Token merging: Your vit but faster

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.444166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:32:45.007491Z digest=sha256:8170f52fa978cf266d55ffc2ca920f4634aad049c8503c025ebd384244c760ee

Observation 1d098822-e46e-43ba-98be-c0a9f7a2e67f · outbound

This paper cites PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:45.136591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:45.136591Z digest=sha256:8fe724718af08b6d8cc99640993d77bf450b72c457dd42fbe746947922b4ba3a

Observation 0670cd7b-274a-459c-8110-fa130e3564d1 · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Quo vadis, action recognition? a new model and the kinetics dataset

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.411633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:32:45.252215Z digest=sha256:1e3c3d0b28b2dc933f9d85fae98d84967de29e75a43246b024fcca7cc15842db

Observation 19b8f3b1-b4e2-4330-b2cc-7d25e356367c · outbound

This paper cites Honeybee: Locality-enhanced projector for multimodal llm.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Honeybee: Locality-enhanced projector for multimodal llm

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.392591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:32:45.386772Z digest=sha256:f3890c976d5c4a721f077654c8297e2c6e0f3a320768fc55aef810dc943772aa

Observation 37f2bc3d-cab8-4313-b194-437375aacf32 · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.374669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:32:45.517307Z digest=sha256:b546645224a88892923291da3b425cabde2f456b195b0918900ce17e758ee6b8

Observation b7c8df05-599c-413e-b4df-e5ea529487ce · outbound

This paper cites Longvila: Scaling long-context visual language models for long videos.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Longvila: Scaling long-context visual language models for long videos

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.352568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:32:45.618165Z digest=sha256:3bbd5e8e1c65e98508d495639892a3948143ecdb121064d043853f1735c82d2d

Observation e62eee1b-61e5-463a-9b3c-2e7f17199433 · outbound

This paper cites vid-tldr: Training free token merging for light-weight video transformer.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs vid-tldr: Training free token merging for light-weight video transformer

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.331836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:32:45.710034Z digest=sha256:aaa0c5fdd3a2e8c0b81d86d5d58dcc6bac5322d5d08d2a9bd0c92a90cafd564b

Observation 5e573a74-dfd0-4e4c-b898-554117a8c69a · outbound

This paper cites FlashAttention-2: Faster attention with better par- allelism and work partitioning.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs FlashAttention-2: Faster attention with better par- allelism and work partitioning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.306539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:32:45.804392Z digest=sha256:22688c6dcb586f9ca575d0a430a266520fb57c3810d08ee7f140f4c86bf60a49

Observation be0085a3-55d6-4f57-a88d-a988eef926be · outbound

This paper cites Fu, Stefano Ermon, Atri Rudra, and Christopher R´e.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Fu, Stefano Ermon, Atri Rudra, and Christopher R´e

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.284441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:32:45.866227Z digest=sha256:1c9bf01a1ec395855e59276f8ad0d2677c646907034a35b4911603c790b1e125

Observation ec769102-d618-4f13-a71a-79fbae6b2218 · outbound

This paper cites The Llama 3 Herd of Models.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs The Llama 3 Herd of Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:45.999654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:45.999654Z digest=sha256:15acb97480e5e81ed784e4ece3988eab37a187fb34d5e3c3ce292210289df048

Observation 2f038acb-cd79-4778-b467-fe202eed1361 · outbound

This paper cites Masked autoencoders as spatiotemporal learners.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Masked autoencoders as spatiotemporal learners

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.265252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:32:46.168771Z digest=sha256:7a07ccdc13c90e3738100479f0d88d1da846c34b0e2214ea671d666e651473ee

Observation 77307511-04a3-486b-80c8-5018115eb2da · outbound

This paper cites Quad trees a data structure for retrieval on composite keys.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Quad trees a data structure for retrieval on composite keys

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.245530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:32:46.273906Z digest=sha256:d8efabe7ef99675a2254c56a60ec89f2752fd27fc7c1b00170ca66550b88c1e0

Observation 86f94159-5ea6-41b5-a195-dcd31f95a052 · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.222087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:32:46.372576Z digest=sha256:39c6febc4f79c7cf559d1bdf14b8fc0f3a2c1c73e017217a5b4f98a208415247

Observation e842939a-c212-4e01-b7fd-dfc93bcc6dee · outbound

This paper cites FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:46.455680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:46.455680Z digest=sha256:587cf79070d813f2c3c49d2928b32e57c8d24d34453cea968954cc54882de3aa

Observation f045812b-d2eb-4998-a0d9-76435809e32a · outbound

This paper cites Caching — google ai, 2024.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Caching — google ai, 2024

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.204484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:32:46.509746Z digest=sha256:2b5409fc899c3a41cb0f9a3f97a91b74f4f16602a3b00d739f13e2445f0e64e1

Observation dcec1c7b-82b6-43cc-bda6-1910dc2673e5 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:46.642984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:46.642984Z digest=sha256:02735f0a885c87df9880c5e083320052d377fb7c639648b659249617ca8cfe68

Observation e5902d98-fde2-470b-9f0e-c750d81456e5 · outbound

This paper cites Deep residual learning for image recognition.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Deep residual learning for image recognition

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:46.759667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:46.759667Z digest=sha256:500bcbbb14b8ab78d98955bc0d79583b58af31975ac1bf09defd633ea7710715

Observation 0d4efc77-4d15-4bd9-85ff-deabb6392d3a · outbound

This paper cites Masked autoencoders are scalable vision learners.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Masked autoencoders are scalable vision learners

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.171666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:32:46.855039Z digest=sha256:838af3cccd1216b128ef085e34bd90af26d87f912ceedf090b9afdc92363684c

Observation d544c777-c601-45fb-a23e-d49639e58b38 · outbound

This paper cites PruneVid: Visual Token Pruning for Efficient Video Large Language Models.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs PruneVid: Visual Token Pruning for Efficient Video Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:46.995554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:46.995554Z digest=sha256:ff2b8fc874b472c1077468e747fd7639fa02d3542c3a0ee02daee6fed4bf10e0

Observation 39e5f196-45bd-42c2-a28a-cb5acb60cc83 · outbound

This paper cites Chat-univi: Unified visual representation em- powers large language models with image and video under- standing.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Chat-univi: Unified visual representation em- powers large language models with image and video under- standing

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.148978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:32:47.097387Z digest=sha256:c361d605805c746129a9acff34e25291c7f134ee0aea32e962964cf1fdced7b2

Observation 8d0fd6b9-ffdf-4105-9e21-fbf443daf535 · outbound

This paper cites Needle in a haystack – pressure testing llms.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Needle in a haystack – pressure testing llms

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.130345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:32:47.206402Z digest=sha256:f36dff202278d0a784ba3b791c5e0b381c10176a22a2441f66876160e99babef

Observation e7b38e68-d904-4f5a-b31d-54e76909e05b · outbound

This paper cites Handwritten digit recognition with a back- propagation network.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Handwritten digit recognition with a back- propagation network

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.113062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:32:47.301474Z digest=sha256:b481b28653dec3c95efd1b8ec2285dc6a684c9048f0764eec18c334f9c427041

Observation 85685f60-2fb4-4cc9-86df-c99a101a9a0c · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:47.367536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:47.367536Z digest=sha256:2ed33e7a16840bfc42b2ec0cda8ce828992c0ecaeb21f21349f0ea6d68a281d9

Observation 8b0f4a29-485c-4221-ac3b-4e0064bc0f6f · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.081714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:32:47.470093Z digest=sha256:9bc850ef1dd7cd52f622620149c54e2d031ffe1b368d22d8e5b108b6e5a6c93e

Observation 3c22a84a-1f9d-4894-bd7a-ddec8a09b6be · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs VideoChat: Chat-Centric Video Understanding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:47.575406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:47.575406Z digest=sha256:e94b0fedc7b125e3de54795ed8de6e7609d3c583a7792173dcff27a8d90c7b89

Observation e2e21bf8-0ea7-422d-92da-9f74675e0312 · outbound

This paper cites Mvbench: A comprehensive multi- modal video understanding benchmark.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Mvbench: A comprehensive multi- modal video understanding benchmark

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.063042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:32:47.618941Z digest=sha256:9098051e7b7c554ef7e4c8c4be06c67bd6bd438539031539d131e5e8808b6f57

Observation 77420b45-60dd-4a2a-ae06-5c850c902b2e · outbound

This paper cites Video-llava: Learning united visual repre- sentation by alignment before projection.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Video-llava: Learning united visual repre- sentation by alignment before projection

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.042764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:32:47.732124Z digest=sha256:c621ea0cd1434fba6e41ecd73fa3c756fee9550fca792b5f31781fab94bd4390

Observation bf4da92a-b536-4555-9aed-2d58320ef6c8 · outbound

This paper cites Visual instruction tuning.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Visual instruction tuning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.017634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:32:47.844102Z digest=sha256:1b0eddcadfda28781b0ac79597f4d66b2e51c52bd6c381a8664392fd7bc468eb

Observation 1e5324cc-5866-4d60-a22a-132504cbcc9d · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:47.962155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:47.962155Z digest=sha256:f96c93fb77b689b8730247fe8db6af254a5dc04c6d05e6e693b4eef748647f06

Observation caeac0d6-d89a-49eb-b1f9-09ba2d3d5223 · outbound

This paper cites Ring atten- tion with blockwise transformers for near-infinite context.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Ring atten- tion with blockwise transformers for near-infinite context

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.989405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:32:48.093263Z digest=sha256:50eeae1bdddf8960900792d381e736d2baa6e0ef628e70749d359b3014358064

Observation ed03f2a1-23f5-4a19-b587-574980db0d02 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:48.187988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:48.187988Z digest=sha256:baa38515dac2a2977042e7cadcedb3fc3547aa019cad5029e5e7881d23f80848

Observation 16437856-d4e5-4593-8847-7fa463692fc7 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long- form video language understanding.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Egoschema: A diagnostic benchmark for very long- form video language understanding

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.947353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:32:48.260745Z digest=sha256:b2f107a749311a5b53cd9cfd376750c97640d22c4612307f52d6f7f0f36df39a

Observation 4b8f8318-4e79-4561-9f1e-926a1142560c · outbound

This paper cites Efficiently scaling transformer inference.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Efficiently scaling transformer inference

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.915585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:32:48.312111Z digest=sha256:8ab8a164e25ef61ce9c7e3c11bbea0ea022a62376a316ae34b5061a768b55298

Observation 687cec00-dd84-4cf7-ad26-d5f4fadf6b97 · outbound

This paper cites Timechat: A time-sensitive multimodal large language model for long video understanding.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Timechat: A time-sensitive multimodal large language model for long video understanding

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.887336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:32:48.362612Z digest=sha256:21d58841a60058f8e1b2a57140c14869108c47ca67d71be26c024dfad4c9d9c9

Observation 83a72fcd-4822-4217-91f4-a59a1769c279 · outbound

This paper cites xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:48.415433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:48.415433Z digest=sha256:9828e603e84ed549853c7d6a3db05fe4276cf174429b9df19886330164330213

Observation 88fcb5da-9be6-487d-b531-a752d5631b86 · outbound

This paper cites The quadtree and related hierarchical data structures.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs The quadtree and related hierarchical data structures

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.860677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:32:48.491626Z digest=sha256:f6143969b6a5d65eb0186f29ef6461ff33030cc11059d8f4ddd26836847b7eed

Observation bbe2353b-5f2b-47e9-b11b-23ed357065d0 · outbound

This paper cites Llava-prumerge: Adaptive token reduction for efficient large multimodal models.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Llava-prumerge: Adaptive token reduction for efficient large multimodal models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:48.573421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:48.573421Z digest=sha256:2aabc2f01e882c53185c6daeadf3efc20a9cea45d8cc2b73b1f3b89ed44b346d

Observation 7ee7ef49-1ddd-424c-848b-70815a92fcd9 · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Moviechat: From dense token to sparse memory for long video understanding

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.830881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:32:48.647081Z digest=sha256:71f6e3bb2a875b9093a6045f5e941106a5f0b1450959ca09f84b9482429ba7eb

Observation afe23cbc-779d-4373-be94-3f1a48a468b9 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Roformer: Enhanced transformer with rotary position embedding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:48.725013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:48.725013Z digest=sha256:00da253a69613c428be156528bb72d9a4acf48fbf00d7ea048356dee529f1244

Observation 1026ae0b-c4e0-46ae-a06a-306d747fec8e · outbound

This paper cites Efficient quadtree cod- ing of images and video.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Efficient quadtree cod- ing of images and video

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.786169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:32:48.768904Z digest=sha256:11382390ee5a5ae35fa78cc65c749796559674f42e6abc8d32a295eaf90cde51

Observation 827c1b70-8999-44b3-845f-e9a464a8003c · outbound

This paper cites Overview of the high efficiency video coding (hevc) standard.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Overview of the high efficiency video coding (hevc) standard

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.766030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:32:48.839547Z digest=sha256:607cfa4b4f7b2723a2285556205d8b54e10aeca8ccbb656746e54e8cab028cda

Observation 1e0ee72e-7428-48f7-b981-87a41aa44524 · outbound

This paper cites Dycoke: Dynamic compression of tokens for fast video large language models.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Dycoke: Dynamic compression of tokens for fast video large language models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.739311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:32:48.887333Z digest=sha256:d1f03787e5a848299dd5a9b50f5c11c6e2c340aeb19838330f22cf7d0ea05063

Observation 1928e642-4168-443d-9c79-469fbc97316f · outbound

This paper cites Efficiency of a good but not linear set union algorithm.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Efficiency of a good but not linear set union algorithm

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.619484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:32:48.958480Z digest=sha256:593777b21e582bf9dc3e1fbacb29cfdcb916e2290b262fff766183385a52aee7

Observation b26d5fc4-c105-4f25-9614-c9a0da4e329b · outbound

This paper cites GPT-4o System Card.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs GPT-4o System Card

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:49.021249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:49.021249Z digest=sha256:1f2f3da0e8035f711b8b4bbdcb684f1a680091d49294f3abc08ec080cd121cc1

Observation c7ba62c1-f17b-40e1-b81c-0cf3e1fadad7 · outbound

This paper cites Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.511623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:32:49.081969Z digest=sha256:b45eb0fe55cf704ba3fa28184c00ea9c9200aca42cd856c8ce93cd54d7d00dfd

Observation f213867e-cd85-461f-836d-1cc6c5668753 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs LLaMA: Open and Efficient Foundation Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:49.147349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:49.147349Z digest=sha256:88d23e6ca33e99196d1800fd31cca2619490a50e4b4a3e46eb40be1d17204628

Observation 56b383c2-f3a5-4120-99e2-96f1213e87d4 · outbound

This paper cites Learning spatiotemporal features with 3d convolutional networks.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Learning spatiotemporal features with 3d convolutional networks

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.371957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:32:49.185062Z digest=sha256:68c40bd45358bd0953eb294f5e2b60360f84655bbcc3c4f6e8058549786668d6

Observation 50d62950-433e-436a-9662-4ce9c855f577 · outbound

This paper cites LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:49.280218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:49.280218Z digest=sha256:c7c8689cb63798952e577ff7c2c19e058851f0d04516ea2d1f78e1f6db13a644

Observation c5a0000c-cdb7-4c5a-8f89-03803715ded6 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:49.345961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:49.345961Z digest=sha256:482f307823a38b969cdabc9192fd4935347dd213e15df30c42fd267c523a11bc

Observation 0778abac-b905-4d0b-ab9a-8dd7e3b71467 · outbound

This paper cites VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:49.423515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:49.423515Z digest=sha256:0542459e0c569716170196063754825e4e398ef03ba733b9910e3508e9382c79

Observation 8884aa8f-17de-46c5-99ab-cd1704ee6913 · outbound

This paper cites Sullivan, Gisle Bjontegaard, and Ajay Luthra.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Sullivan, Gisle Bjontegaard, and Ajay Luthra

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.254156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:32:49.484414Z digest=sha256:2acfa186fbecd42391017c50fdc22f7bcef408b31a272ce4ba547a2b91bf7f32

Observation abe2f266-b145-441e-9c51-376e10e9557d · outbound

This paper cites Longvideobench: A benchmark for long-context interleaved video-language understanding.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Longvideobench: A benchmark for long-context interleaved video-language understanding

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.120020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:32:49.577022Z digest=sha256:6948002421c2fdd41a6ee178e03b7abf79719cbe79c9d6d809a0d58ee44f0599

Observation 94cc60c0-1937-4263-a654-9ed27861d065 · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Next-qa: Next phase of question-answering to explaining temporal actions

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:49.636803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:49.636803Z digest=sha256:27410179639037c5e5cf6296990e9d5a07d43ae5362d68811f56c2e2bf39f566

Observation 7560d55b-5148-499d-92b8-ca763da865a8 · outbound

This paper cites Pyramiddrop: Accelerating your large vision-language models via pyramid visual redundancy re- duction.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Pyramiddrop: Accelerating your large vision-language models via pyramid visual redundancy re- duction

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:51.945416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:32:49.707972Z digest=sha256:fd49aa8c76f8af8553ee28ef83bef82e6516bfdf78ce53ae12b17cc70636b169

Observation 45a6af3d-703e-4459-8708-f0a4d375049e · outbound

This paper cites Qwen2 Technical Report.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Qwen2 Technical Report

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:49.783006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:49.783006Z digest=sha256:7a946d670fcceea366de7bb4530cadd15ce40efe47c256cc110822ebe07bcbaa

Observation 5c385049-5d20-4389-bb81-fc98a8695fd0 · outbound

This paper cites PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:49.863938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:49.863938Z digest=sha256:aa263ee6012f19b036a3b9c58b71cc0cabb0c04d2e22cd3219240dd25fc28a0e

Observation e182b257-b922-4262-ad32-3d431f68bda0 · outbound

This paper cites Video-llama: An instruction-tuned audio-visual language model for video un- derstanding.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Video-llama: An instruction-tuned audio-visual language model for video un- derstanding

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:51.847754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:32:50.034259Z digest=sha256:7cc5bff69c4b93f152613424060440b31590bbf3c6f2d3e1153d3c99c0ed28d3

Observation 602199ed-14c7-4fbf-9815-939352dc8fa4 · outbound

This paper cites Long Context Transfer from Language to Vision.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Long Context Transfer from Language to Vision

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:50.185274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:50.185274Z digest=sha256:ddee5d22538b6de87a96065d9f9d0fedcc82e9e22612cf1a5c20bbef1031a55a

Observation dd7c2631-a2ca-43c3-be08-e90da971380a · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:50.312891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:50.312891Z digest=sha256:a01dc08eac5dddfbc1304fd9a1a8a901523762b2958ddc2434f64f272f6758a6

Observation f3fd2480-95fc-42f5-8e71-53514a96b734 · outbound

This paper cites Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:50.442322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:50.442322Z digest=sha256:8140258689cebd441b3ba69f0e7f1b255c78ee98591e8a060657245315fc604b

Observation a7e64a4e-366b-4623-9790-6ec1045e116f · outbound

This paper cites Mlvu: Benchmarking multi-task long video understanding.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Mlvu: Benchmarking multi-task long video understanding

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:51.646767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:32:50.667041Z digest=sha256:384916a42066046f30e2ec87275830822128ebd14277716282bf7ccfb60a9a01

Observation d05ef8e8-af0a-4b5f-8c80-030feee8b836 · outbound

This paper cites 11 to 14 show the absolute values for the main comparison results.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs 11 to 14 show the absolute values for the main comparison results

Reference 2025

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T18:32:51.576968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:32:50.801112Z digest=sha256:1cfd5798cbfdbd32a4f1f0551a3e69530bdeeee6f59b455dfc88cd17c3cb2125

Pith citing papers

Observation e99a8d45-4dcb-453b-80d4-725afbba76f0 · inbound

Video Parallel Scaling: Aggregating Diverse Frame Subsets for VideoLLMs cites this paper.

Video Parallel Scaling: Aggregating Diverse Frame Subsets for VideoLLMs Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:36:44.191878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T18:35:01.328250Z digest=sha256:528d85e56ed0fc5f4dc4552352ab0286e6a5a52e7659af76e75d584a3a984296

Observation 8cb5c6ec-fea4-4edd-8ed6-17874b0c7018 · inbound

VLMaxxing through FrameMogging Training-Free Anti-Recomputation for Video Vision-Language Models cites this paper.

VLMaxxing through FrameMogging Training-Free Anti-Recomputation for Video Vision-Language Models Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:11:13.772797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T01:30:15.463051Z digest=sha256:213d01ab5cc8d0d6430763c4fde0b78b64ceae4c085a680f9d7662c3837d234e

Observation 27a489f0-2b34-46c3-b16b-76908f60e426 · inbound

DynaTok: Temporally Adaptive and Positional Bias-Aware Token Compression for Video-LLMs cites this paper.

DynaTok: Temporally Adaptive and Positional Bias-Aware Token Compression for Video-LLMs Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:58:05.989905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T06:55:01.619441Z digest=sha256:a0b9060adc0cb6fd507032e48b843b1987b9b8f28a7738bc66acf9369311a3e6

Observation d890eb34-3b8c-4d7c-8973-4ab572bb59d4 · inbound

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models cites this paper.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.654336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.654336Z digest=sha256:38c64678496a80dfa01a692d8737639944dac5614373ad4a862a39a1b9457bd1