Pith. sign in

Paper Citation Record · LEDGER

Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models

As of 10 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 0 inbound Pith citation observations for arXiv:2608.03112.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.03112 v1

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T00:55:50.162416Z

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

34 of 34 outbound references displayed

  • verified exact0
  • verified fuzzy9
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6aa95c04-c4fd-424c-bf4e-5f080bb33dc9 · outbound

This paper cites Divprune: Diversity-based visual token pruning for large multimodal models.

Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models Divprune: Diversity-based visual token pruning for large multimodal models

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:55:50.724195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T00:55:50.036588Z digest=sha256:a3c87330ef30456bd7667f37ffd4634e6ba9d158cdf6d44a8a9f43c8dbafa308

Observation 42b15421-c221-4477-b3a9-e3a5bfdbbe09 · outbound

This paper cites Qwen2.5-VL Technical Report.

Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T00:55:50.040897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:55:50.040897Z digest=sha256:997c6c9a996a1bf6ce7bfa8dc50b3053eb1458fd32da8ea802035993ae2c3a1e

Observation 0ec23800-fd5b-4476-8bc9-bb6b9304a4b7 · outbound

This paper cites LLaVA-KD: A Framework of Distilling Multimodal Large Language Models.

Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models LLaVA-KD: A Framework of Distilling Multimodal Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T00:55:50.045305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:55:50.045305Z digest=sha256:5778026b7fc840b801277fec759972f70d0757c1cef4db98c291c043ec0fc8e3

Observation 3a00e5f9-79d9-4cc1-b3a2-4b63c60a11b1 · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models.

Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:55:50.713888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T00:55:50.049443Z digest=sha256:a412605714d107d3714d977c549718962fba668cc9cda5a83ef57b9c5502ddb7

Observation 6aaaff18-9677-42cd-9ac8-98930dc07965 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T00:55:50.053363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:55:50.053363Z digest=sha256:cd5b4c39ce68d6bbb3d035c55457628ea67dccd8e71be700821422292a2ae20e

Observation e04ecb77-5faa-4c3d-a1c5-ab31c2c5dd2b · outbound

This paper cites Instructblip: Towards general-purpose vision- language models with instruction tuning.Advances in neural information processing systems, 36:49250–49267, 2023.

Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models Instructblip: Towards general-purpose vision- language models with instruction tuning.Advances in neural information processing systems, 36:49250–49267, 2023

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T00:55:50.057359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:55:50.057359Z digest=sha256:4639c6bcb041cf5b6a072a3015031da44969d6327d2d4af495db248f874e8af3

Observation 43fb8f65-9a80-46fd-9c21-fc66ef589dfb · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T00:55:50.061082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:55:50.061082Z digest=sha256:13474d117990a44bb2a1882cc8b6c2a2fc9793e9318cd39df55566697a872965

Observation 126021bf-1271-4442-a0ca-2f14caebe633 · outbound

This paper cites Attention Score is not All You Need for Token Importance Indicator in KV Cache Reduction: Value Also Matters.

Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models Attention Score is not All You Need for Token Importance Indicator in KV Cache Reduction: Value Also Matters

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T00:55:50.064934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:55:50.064934Z digest=sha256:9031b4b18e56058ea9ce6818f77ca21c2e108b403985a6cd2262947965ead615

Observation 97536288-417c-4bb4-a920-90d2bcc1e097 · outbound

This paper cites Filter, correlate, compress: Training-free to- ken reduction for mllm acceleration.arXiv preprint arXiv:2411.17686, 2024.

Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models Filter, correlate, compress: Training-free to- ken reduction for mllm acceleration.arXiv preprint arXiv:2411.17686, 2024

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T00:55:50.068691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:55:50.068691Z digest=sha256:e80875b4d5207c7171963e0ba240465fb32d5042879bfade4f0db56216925de5

Observation 7e115841-d1f7-4f52-aae3-970b7392d232 · outbound

This paper cites Efficient Multimodal Learning from Data-centric Perspective.

Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models Efficient Multimodal Learning from Data-centric Perspective

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T00:55:50.072148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:55:50.072148Z digest=sha256:c5d8d61190acd6edd63232af7ad6946d669e2350374abecf0dd2005c9afff44f

Observation 9b88ced9-c6c7-4e39-8b1f-f92e6d2fd0e6 · outbound

This paper cites Ivtp: Instruction-guided visual token pruning for large vision-language models.

Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models Ivtp: Instruction-guided visual token pruning for large vision-language models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:55:50.691372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T00:55:50.076060Z digest=sha256:2cb0860a2a66c47f0a22671c2d9ad6cd6f69b4c3645be3179882ddd4bf108bb2

Observation 475ef900-2adf-4dfa-aa48-3fb750537e2c · outbound

This paper cites Fast pruning using principal components.Advances in neural information processing systems, 6, 1993.

Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models Fast pruning using principal components.Advances in neural information processing systems, 6, 1993

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:55:50.680940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T00:55:50.079301Z digest=sha256:39211eafd46ff4e188a7654d09de21441c1d0f1a39a0d82e051be73dccc2db42

Observation 17d30c23-5d1b-4b62-b66b-5d3912721494 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T00:55:50.082897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:55:50.082897Z digest=sha256:78f52e9ed6f052303182b185b852367dd2bc724e72489a2cdd4d9e9b04bd18c9

Observation a2efe416-7b05-43a0-938b-57e1118d90c7 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models Llama-vid: An image is worth 2 tokens in large language models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T00:55:50.086574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:55:50.086574Z digest=sha256:9df8f8870c9c089a2e1b75db40ce42ed6415dce38b1a5f8a88ae1eceef19f30c

Observation a2fbb7ef-1df6-4de6-a5ce-e13c9d235442 · outbound

This paper cites Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding.

Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T00:55:50.089925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:55:50.089925Z digest=sha256:0df87439b63073a54fa16973a1d40824b5386e49e6a3451a2f4070b75f210ceb

Observation d0860654-ac66-4fdc-9b8d-f70b7d351ac5 · outbound

This paper cites Video detail caption.

Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models Video detail caption

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:55:50.665223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T00:55:50.093552Z digest=sha256:8c7526c1422a539e06941cca1f1c34de76735fe2797ad3402e7799dc7d48eb14

Observation c51946b4-8667-4d8d-8930-2d124f6c9292 · outbound

This paper cites an unresolved cited work.

Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-08T00:55:50.654611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T00:55:50.096902Z digest=sha256:55a9dca35b7bfae65a03e7df9d0e6e8036d08651fe4c8174838937771f25efcf

Observation 08fe127c-bdcf-4157-b0ee-2d503c0556d9 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T00:55:50.100827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:55:50.100827Z digest=sha256:4396686bfb2332ad5d114d6bd7512b876a748f7160e861a2f53ea416d2a3a5a6

Observation 581c3805-4d45-4d1e-bd00-1602d8e60b17 · outbound

This paper cites Per- ception test: A diagnostic benchmark for multimodal video models.Advances in Neural Information Processing Sys- tems, 36:42748–42761, 2023.

Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models Per- ception test: A diagnostic benchmark for multimodal video models.Advances in Neural Information Processing Sys- tems, 36:42748–42761, 2023

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T00:55:50.104780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:55:50.104780Z digest=sha256:af0f93e37db48156ec5bfe422c3ed004728e00b5535c28af9ab489c95f3c55f7

Observation 30bbac77-b536-4487-b583-6426e3c42043 · outbound

This paper cites Llava-prumerge: Adaptive token reduction for efficient large multimodal models.arXiv preprint arXiv:2403.15388,.

Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models Llava-prumerge: Adaptive token reduction for efficient large multimodal models.arXiv preprint arXiv:2403.15388,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T00:55:50.108437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:55:50.108437Z digest=sha256:88884161cbd3df22bebcb98aa55d842af82fd4b21b81679d8cfed4abce4ec23b

Observation 1cad450f-7298-4506-beb1-ab398a6982d8 · outbound

This paper cites Imp: Highly capable large multimodal models for mobile devices.IEEE Transactions on Multime- dia, 2025.

Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models Imp: Highly capable large multimodal models for mobile devices.IEEE Transactions on Multime- dia, 2025

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:55:50.638126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T00:55:50.112464Z digest=sha256:05b0226b468727031d399cb04dd654e613e0460a128f6be961ba7c165511ca6b

Observation 346c318a-9fc1-49a1-8dd2-d59fbe86ab65 · outbound

This paper cites LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T00:55:50.116104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:55:50.116104Z digest=sha256:7f713db397f9b241a19585c6093b7268f90ac4417d2b9d9ab6eb09f55204cd0c

Observation afb7bdd5-f0d4-4314-9070-1171eb84c3f4 · outbound

This paper cites LLaVA-MoD: Making LLaVA Tiny via MoE Knowledge Distillation.

Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models LLaVA-MoD: Making LLaVA Tiny via MoE Knowledge Distillation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T00:55:50.120232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:55:50.120232Z digest=sha256:37f23860ddf7d12dcbe01ca92011f2eccb3c535238db5087c04c1e9f49ac14e4

Observation 636baae2-9a92-4e79-84c1-c37fdd4ea012 · outbound

This paper cites LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs.

Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T00:55:50.124410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:55:50.124410Z digest=sha256:91cc83f143529d30a03e6731a13f6ac0bef7ffe5bf3b7bd5768a9a1ab0d0af17

Observation 5449c859-83cf-4de4-87c8-6b5172039234 · outbound

This paper cites Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM.

Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T00:55:50.128159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:55:50.128159Z digest=sha256:e05d45599a3ccf63ec158f23a66c7218cca6390d25aac87cd2cdb1b19f61d8b5

Observation 10cb08f1-eda7-487a-89fd-613ca4a412ab · outbound

This paper cites VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges.

Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T00:55:50.132180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:55:50.132180Z digest=sha256:7687eced92da0e53b70a286fbb9339ada5dc45ac5b617b68a1bfbc9c34553dc1

Observation 637995de-1493-4280-9a81-4a9e02102bf0 · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models Next-qa: Next phase of question-answering to explaining temporal actions

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T00:55:50.136531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:55:50.136531Z digest=sha256:16e6a193b7387517d4d132e980a168ec3517cb474573851fe42c6834862b45e3

Observation 5c40ef14-ccdd-40e2-ba5e-56e629d10789 · outbound

This paper cites Topv: Compatible token pruning with infer- ence time optimization for fast and low-memory multimodal vision language model.

Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models Topv: Compatible token pruning with infer- ence time optimization for fast and low-memory multimodal vision language model

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:55:50.620506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T00:55:50.140040Z digest=sha256:1e112516365c69d8d34e70e8c059bbc2f5044107c4d30cc17ed953343bcb9fc1

Observation 4f98ddbd-2456-4d52-bcca-2b69b771b108 · outbound

This paper cites Atp-llava: Adaptive token pruning for large vision language models.

Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models Atp-llava: Adaptive token pruning for large vision language models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:55:50.609139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T00:55:50.144028Z digest=sha256:90d76681d8ddd00e0bee11052c5e8c1d7e8128fbe3dde105609208ee2be73170

Observation ed439344-4c14-4165-ae58-c54fe0510026 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T00:55:50.147700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:55:50.147700Z digest=sha256:531bc1895bc6dd07307122bd39d7a286f4d4a426f5d4b116deb777b3b260489b

Observation f2b898cd-650a-4113-89c2-8fac10f34d16 · outbound

This paper cites TinyLLaVA: A Framework of Small-scale Large Multimodal Models.

Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models TinyLLaVA: A Framework of Small-scale Large Multimodal Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T00:55:50.151584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:55:50.151584Z digest=sha256:b694fd4e76f845e00f3cb80830eb8a74672cb6abea543460ee6894b57abb92d4

Observation 89dabbd2-75bb-43a6-9a2e-664507f1caa0 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T00:55:50.155447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:55:50.155447Z digest=sha256:1a7d03218464d82adeeaae904a128cfd2b7009354fbe57694f373b22524a18e6

Observation e854eb82-62e2-4e56-8079-a8cc27b8091f · outbound

This paper cites Also, we use beam size of 1, and the number of maximum new to- kens is capped to 1024.

Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models Also, we use beam size of 1, and the number of maximum new to- kens is capped to 1024

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:55:50.597966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T00:55:50.159058Z digest=sha256:8971f20af1928be04a4e6c34c9ed8a073e8e1a1a777f291d29d416e1311f494a

Observation f25dd4a7-d912-47b3-91ad-b1909b26b579 · outbound

This paper cites an unresolved cited work.

Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-08T00:55:50.586532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T00:55:50.162416Z digest=sha256:33ca0bbb2c066aea33450d471d1712bcf580bed633ad424198855ccb8dd11028

Pith citing papers

No inbound Pith citation observations are available.