Pith. sign in

Paper Citation Record · LEDGER

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models

As of 17 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 0 inbound Pith citation observations for arXiv:2608.03083.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.03083 v1

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:57:56.752601Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

63 of 63 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved62
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e1085018-dfde-44b0-85f7-78b5c614ad57 · outbound

This paper cites an unresolved cited work.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:56.108091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:56.108091Z digest=sha256:d3645a99ef69ddb8c44683ef602490e5b456f77f28c6b5bd7766b64448f8cda7

Observation ff5ee652-4d49-40f9-9b74-17d6e282f1d8 · outbound

This paper cites an unresolved cited work.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:56.131532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:56.131532Z digest=sha256:dcbb10fb7a5f197fdef13e44c9da648b4d4e94daee30b2038a8140d306350fe9

Observation 8c807c83-78d9-46cc-8132-fc2d4d69d817 · outbound

This paper cites an unresolved cited work.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:56.152004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:56.152004Z digest=sha256:f108272d49cff2944c19e3d5f3e054c3bac24cb2bddc6a2888efdacf8619379c

Observation 2752700b-4ebb-45cb-9490-5e443217e22a · outbound

This paper cites MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:56.162092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:56.162092Z digest=sha256:419089403acca1f9657478244afa845528432eb8fa1ce897f9a17372f4aaa2ee

Observation 47de14fd-faf0-4cb9-8de4-a139f1b4bfb3 · outbound

This paper cites Qwen3-VL Technical Report.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Qwen3-VL Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:56.170077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:56.170077Z digest=sha256:71c24f47af54cfc57a0586acadcf9809217572f18b6067e3c6f960eb0a0a0a8a

Observation 33ee5007-8e23-48a8-99d7-2d9d0a657a08 · outbound

This paper cites Qwen2.5-VL Technical Report.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Qwen2.5-VL Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:56.181207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:56.181207Z digest=sha256:69bc710948b610085bb3ba38e2ddb7a3ac31f1fe15993ea4468866bd751e31aa

Observation 976a2390-5358-40c2-8220-69286d6b0dbc · outbound

This paper cites Token Merging: Your ViT But Faster.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Token Merging: Your ViT But Faster

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:56.189069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:56.189069Z digest=sha256:b4676fe42d24124bbb3bf638bc515cd4dccce26aa9e5f32759b224ab36a6985a

Observation 22c1062b-1245-4720-88ab-07791839a334 · outbound

This paper cites VideoLLM: Modeling Video Sequence with Large Language Models.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models VideoLLM: Modeling Video Sequence with Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:56.198538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:56.198538Z digest=sha256:9e3b2af9d6f0ddde77eeb8539826b66a9ad61ad52aa0a6b326c8692bf9c10338

Observation cb9c2de3-7353-4897-8dc9-929353f72484 · outbound

This paper cites an unresolved cited work.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:56.205337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:56.205337Z digest=sha256:2e4a4a781280b0ba2bed1274a3a4852fd263dfd100479b56da0b15f381d1bb51

Observation acc0232e-b8b9-46d0-87e4-a5a1681fad42 · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:56.212904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:56.212904Z digest=sha256:d4aa7133a3cbe0d4f7e4484cc996a28c569b588d7f1f2773d37bdcd668ad5d4f

Observation 4b8d239a-2792-41d6-939e-32466093dcbf · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:56.227468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:56.227468Z digest=sha256:173b2c9454b2fc12ade115c7569b7f794541acd3d61dee2b94949029decf20df

Observation 036ab384-5f20-4baa-bcdc-384ca0e7152d · outbound

This paper cites an unresolved cited work.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:56.235912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:56.235912Z digest=sha256:eb2f60dad118e406b25f3014c67a41102ed1678605d50501a6775a1127f316cb

Observation a25a4001-a7cf-4b3b-8750-1cb6cfb6d7db · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:56.242973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:56.242973Z digest=sha256:d29f2fb23b602012c23b51790ec40ac779bb9f9802dd63e24bf54112fdcd0e39

Observation 8402874c-e6eb-499e-98d4-d89bb27acbb8 · outbound

This paper cites an unresolved cited work.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-15T14:57:59.393789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T14:57:56.256934Z digest=sha256:d0735a5057b54457563e59f61015f9780566bcf522e471bc0e18d1b6d6047b3f

Observation 213f6d5c-f975-441d-bd33-15a5cde08464 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:56.266257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:56.266257Z digest=sha256:468f994c62b895be3f915f355772c04f14e7a99506b30af6e4ff0a2bbf2b7f4f

Observation 1b02549a-428a-4f5d-afbc-31f50c60bb1f · outbound

This paper cites an unresolved cited work.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-15T14:57:59.345547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T14:57:56.274197Z digest=sha256:11161b2d169622df8a69e7cb3378dbe9564d64ec4fca0b7854457ff3e49d6e84

Observation 165bbe61-23e5-4761-91bc-dfea5340b00e · outbound

This paper cites an unresolved cited work.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:56.283633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:56.283633Z digest=sha256:dcc90beed7eaa6bd3fb710035ac436a862ccf4c2a4d4e643189804432a0e2d71

Observation 0003ab70-5e6e-44b2-a1a1-5853ce217d7b · outbound

This paper cites an unresolved cited work.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-15T14:57:59.274925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T14:57:56.297668Z digest=sha256:462a13dfe42a7551dc99a930aaa5abf8987366a0ef3e047093ece864bd6ccfd8

Observation 1153f228-79c5-4b5a-aff2-6b76bb7f6cf1 · outbound

This paper cites an unresolved cited work.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-15T14:57:59.224719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T14:57:56.306431Z digest=sha256:e797a7fe41c64776ff9a9b482fcb2b024186464622be7d724b7db3b0838d92de

Observation 86489d16-dca3-438b-89c5-24aa30bcae3a · outbound

This paper cites an unresolved cited work.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-15T14:57:59.174796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T14:57:56.315528Z digest=sha256:48a5f7724c767e890b3568c9519c36997850a2e43afe31efd57256520e7edada

Observation 70854148-96bf-40bf-93b3-4d22dce6448a · outbound

This paper cites an unresolved cited work.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:56.324734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:56.324734Z digest=sha256:c89f6a69e03ad4237d9a5297d4c80c557dab68491fd76823cd4ce97e7b61cf53

Observation f182365d-6de8-44a6-adb0-f3e5b6b9c24c · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:56.332299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:56.332299Z digest=sha256:dc0a1fcb2d6f25b706ad02627b6e809c2681af9506ae2f0639d1e0d55dd29dc1

Observation a631fc30-95a3-4236-afde-8d3068e58e9d · outbound

This paper cites an unresolved cited work.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:56.343809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:56.343809Z digest=sha256:15d95c5416e1fc0a2e6c7c2a6a358489aab72a02b363a4781eb6ee6d70c1d785

Observation a3e910f7-1986-4a4c-85af-70f08a47fa12 · outbound

This paper cites an unresolved cited work.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:56.354162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:56.354162Z digest=sha256:4ec7e563c091c01bbfa84e976a81b6754b3e7816c5e80cb55e6625336d113b25

Observation 9c9da7ad-f05a-4c11-8056-47f981526d11 · outbound

This paper cites an unresolved cited work.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:56.362695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:56.362695Z digest=sha256:1f8af32d770b76d3d272df9a8f3d6e6cac777eaeb05e449bc45598b3f8297593

Observation 0215bf0b-ad13-4201-bc61-b19e7b86591d · outbound

This paper cites an unresolved cited work.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:56.375209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:56.375209Z digest=sha256:b9e1706cb5e395ad8643111154a3384259f5ad62bb6bfdfecc8b5781f51ba687

Observation 775cee6d-21f0-400f-b0e0-ce21db608cea · outbound

This paper cites Not All Patches are What You Need: Expediting Vision Transformers via Token Reorganizations.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Not All Patches are What You Need: Expediting Vision Transformers via Token Reorganizations

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:56.388709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:56.388709Z digest=sha256:6ea903c22b6b3a31e81512c949845e9998fe4606303c608c57636a8e02c4ca8a

Observation 4ef048d8-0f0a-4a72-973e-fa2a13665635 · outbound

This paper cites an unresolved cited work.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:56.403321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:56.403321Z digest=sha256:5184cf99a7dc9d63626bb5ecfe2ac0b1f053f5b96b7dc7ac807be5d62a325312

Observation c251361c-8b88-451e-85a8-1924fa62f407 · outbound

This paper cites an unresolved cited work.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:56.418822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:56.418822Z digest=sha256:6e358ee772e461cc2eb616b192128eef0a29d84d6854e545e102cb3b6f15079a

Observation 1c2ff6b0-2fca-4ce6-9d22-c5e484402030 · outbound

This paper cites an unresolved cited work.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:56.430736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:56.430736Z digest=sha256:05f22edd9799f91f274f91b9993ed066c08393fb8c7f2c5d78e149ad7f9e36d3

Observation 4b391219-50fa-46eb-9ead-85bf81236d22 · outbound

This paper cites an unresolved cited work.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-15T14:57:58.930653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T14:57:56.446762Z digest=sha256:40a4a56cf8431792fe6337777f3d088ef104103c93237e747c0348b6d87ada8f

Observation 4060da20-332d-4247-9ca4-fd48cd18631e · outbound

This paper cites an unresolved cited work.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-15T14:57:58.895761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T14:57:56.460789Z digest=sha256:f7b9d7034b6af2dea22f2b2eee1f91d5f1bb6d9d2ea614498a4b83628bda251e

Observation 9ab596e9-d35f-4d94-a3f3-cf32b1e0939d · outbound

This paper cites an unresolved cited work.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:56.477466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:56.477466Z digest=sha256:56fad5792e1f333109a7a160a857b259f810de815fa089ca4d3069a2cfc06d65

Observation 5453f54a-217e-4f0b-bae0-b276527533bf · outbound

This paper cites an unresolved cited work.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:56.490173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:56.490173Z digest=sha256:0e98a9a774bf5e7a5ceb18ad17ee0791565373d005c6c3b395f243e6095bb366

Observation 089bbc5a-ef76-484f-9d0b-149371639113 · outbound

This paper cites an unresolved cited work.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:56.499506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:56.499506Z digest=sha256:96733d63b547cd25809490d9e6359c0ccb9274482d63f39cd95deb0854062162

Observation 9cfc4506-8302-406f-999f-a53d56ef2b64 · outbound

This paper cites an unresolved cited work.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:56.513049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:56.513049Z digest=sha256:0dc73770fa90b72498b352108a1ddff5e4ba714c55a8a79e9250e4fcee3e6adb

Observation f92f55b0-5846-4249-99ec-12f84baaf23b · outbound

This paper cites an unresolved cited work.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:56.535717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:56.535717Z digest=sha256:52840ae4ef1c7007333332fff463be5cb2bc06607f624e422ca1f78de2908491

Observation ddfc45c4-336a-41da-8c5c-8bd60da3a61f · outbound

This paper cites an unresolved cited work.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:56.543695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:56.543695Z digest=sha256:8a7b02e73309c847d1b69ce67d6179c9e35e69fd685ae6f663c544c8c2e22e8c

Observation c44ac986-dd29-4347-af25-a97e532d0fcb · outbound

This paper cites an unresolved cited work.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-15T14:57:58.697057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T14:57:56.555312Z digest=sha256:1dd3b41238b09230bda05f328a2c2714381b598bfa1ab1a06942a1aaebb31d64

Observation 7088f210-d73a-4f6d-ac74-d4d8b01b5243 · outbound

This paper cites an unresolved cited work.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-15T14:57:58.660001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T14:57:56.569779Z digest=sha256:50e911c9ae60af866826066fca30512cd660052b33e3ea066c4c07f8f3f5bd53

Observation 2073a944-5328-40b6-9520-14fe4d6f4c8b · outbound

This paper cites AdaTP: Attention-Debiased Token Pruning for Video Large Language Models.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models AdaTP: Attention-Debiased Token Pruning for Video Large Language Models

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-08-15T14:57:57.269717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T14:57:56.582528Z digest=sha256:4b379928a11b87aa216a86c127e137d45aab55ba77648836db9422e8545e28a0

Observation 772fb8d1-53d1-4470-91ea-b2d273c5ff39 · outbound

This paper cites an unresolved cited work.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:56.589948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:56.589948Z digest=sha256:f5d3fff5153f7999b9d8215c0f5393e8e875d0827a1984a079db7c9ff92869bb

Observation 11f3c715-5a64-47f1-bb66-7ac013f47b7a · outbound

This paper cites an unresolved cited work.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-15T14:57:58.583179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T14:57:56.599390Z digest=sha256:3fe8a7254dfdfb753277e2f7364a043ce8fa62dfd0bde85c7670506c45e30050

Observation a3258b68-4238-4411-aa41-dd5f7abea387 · outbound

This paper cites an unresolved cited work.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-15T14:57:58.550414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T14:57:56.611700Z digest=sha256:fa3911866f15cc02e35b6c2af60981e90c81427c376c600025fc2229a53047d3

Observation 64a4e827-a9c3-4bcd-a4c8-082bf7b7ab40 · outbound

This paper cites an unresolved cited work.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:56.619114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:56.619114Z digest=sha256:43f9f86ada4897b6b43460882544100f9fca7a184f8be0ca4f5c574eebb0c64d

Observation 694e1842-7425-4edf-8f27-01d056af4ff8 · outbound

This paper cites PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:56.627432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:56.627432Z digest=sha256:8bd23e5d7e477f51442ea2b81c927c1abdfb5e27778254ea91237b5576bb3bd1

Observation 1e2e6ee1-1712-47ba-b12f-9545d5dba291 · outbound

This paper cites an unresolved cited work.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:56.633694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:56.633694Z digest=sha256:6619e81cc3af157499b8c794d7b8d6b39de8473263d4715fb026b69b36eb5948

Observation 010c3b1e-cb27-49ca-8f15-57d01737d440 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:56.640802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:56.640802Z digest=sha256:92adb675fc11564d41a3cb0977d2e36c0e80ccd83a8e0f0a7158f6d550ec7c1c

Observation 20e58152-9c2d-43ad-8e45-8501ee920757 · outbound

This paper cites an unresolved cited work.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-15T14:57:58.477114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T14:57:56.650251Z digest=sha256:c9d460404663e9e7ccee638fa55c1e87de61086e74f088f797f5bfd5caff4430

Observation 550c3538-91ba-4b70-8cf3-2ab5bec83702 · outbound

This paper cites an unresolved cited work.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-15T14:57:58.430213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T14:57:56.658637Z digest=sha256:ca01d7cda5aff0ff36deaf216ec13d16e8d19a5a1f90caccb9e2ec1c2cd4b772

Observation 75f927b1-301b-44c4-ad50-57b5fb545933 · outbound

This paper cites an unresolved cited work.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-15T14:57:58.388636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T14:57:56.676220Z digest=sha256:a14dd29ead9ccd53785eb33dd939530b805fa4a64e1afb5b145551526670ec96

Observation f8f6d80a-39b2-43df-bc8d-8d8d45ec9044 · outbound

This paper cites an unresolved cited work.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:56.684005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:56.684005Z digest=sha256:e8b1f1871e123b8ebe7b0cc5866a6892a67220a2cd9df5d564a0774270a3ee58

Observation 11101bf1-83a1-4c81-928b-1444ec23b155 · outbound

This paper cites an unresolved cited work.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:56.699482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:56.699482Z digest=sha256:9474a0de03473f1d359e6dc8dcc60a03f886ec24b89af188883d991e1d70b415

Observation 001fd366-5c11-4c09-9234-3350672fa44c · outbound

This paper cites an unresolved cited work.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-15T14:57:58.301627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T14:57:56.712767Z digest=sha256:327ac7d6191c13999fb9b0cc40ff116cab0380fb25445a4eb349f2bd54ba61bd

Observation 14e7701e-2074-4b82-836f-dad26b1bd571 · outbound

This paper cites Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:56.723618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:56.723618Z digest=sha256:7f72128d76f72825c6c525116e664f01c650770bd2aca3b0b23ecd0cf1ca81ca

Observation 3acfe0d8-6ee5-4625-b64e-e8a1c28917ea · outbound

This paper cites an unresolved cited work.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:56.729974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:56.729974Z digest=sha256:1c3246f68ff09f2bed7e342e30e7185ccbba423bc5d379e1c073160a2bc81dca

Observation c1759393-e52c-498c-be32-fd0e4b07dec9 · outbound

This paper cites an unresolved cited work.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:56.743820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:56.743820Z digest=sha256:55118e465fa6448ef72498c5ed6939c48a37822397476f92276f9f28d4c9e649

Observation fe7b203f-cde9-4956-baeb-2cd35b553471 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:56.752601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:56.752601Z digest=sha256:56d4b4564c3d158ae076879047b6125277b80758466f8f7a52e367db6bc48ea4

Observation a6d206d4-edab-46fa-ada6-a614a83cef4d · outbound

This paper cites SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:56.737747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:56.737747Z digest=sha256:3017ba9a33c139e526212415b644bc18a80074c3f1efe2c694053aa8eaf7a848

Observation 9e8d6223-e916-4026-a7cf-c291ee1541fb · outbound

This paper cites Advances in neural information processing systems34 (2021), 13937–13949.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Advances in neural information processing systems34 (2021), 13937–13949

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:56.523306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:56.523306Z digest=sha256:3809701465759dac284493c8c77f6eae223481c8344fb09e6efee5be7f5c2354

Observation 85cc6769-bc5a-4d74-9385-a720e870132c · outbound

This paper cites an unresolved cited work.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:56.120294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:56.120294Z digest=sha256:00709d78c612b6f109195569f891d97fc27b062b8daf9da6c97b705d7df5b35d

Observation a2aee47a-d6c7-43da-bbc9-29666f38fb20 · outbound

This paper cites Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:56.439046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:56.439046Z digest=sha256:eb9772ab8e341d26f9e93d6e5510779cdb9092c931e648c29c22e67d8d34b73d

Observation bab3dea0-b8cc-4aa0-b581-ee049fd4b787 · outbound

This paper cites InProceedings of the Computer Vision and Pattern Recognition Conference.

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models InProceedings of the Computer Vision and Pattern Recognition Conference

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:56.142848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:56.142848Z digest=sha256:3422b29248093fe479d2c7331a2caaef35012fe1ae391dd9f2495c63a5c498a7

Pith citing papers

No inbound Pith citation observations are available.