Pith. sign in

Paper Citation Record · LEDGER

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding

As of 18 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 3 inbound Pith citation observations for arXiv:2411.14401.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.14401 v2

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:17:19.535023Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-26T00:19:26.153682Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-10T06:15:00.866473Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact0
  • verified fuzzy33
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4775f871-409b-495c-ad49-c996aa385278 · outbound

This paper cites Is space-time attention all you need for video understanding? In ICML, page 4, 2021.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Is space-time attention all you need for video understanding? In ICML, page 4, 2021

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:20.482163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T15:17:19.311077Z digest=sha256:910fd9c1eeecc3cc9684b88d63118378f2440aa9b934bf62aea2f34f82cc54e9

Observation 9894c35c-047e-4413-9367-df150fa01062 · outbound

This paper cites Token merging: Your ViT but faster.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Token merging: Your ViT but faster

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:19.316997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:19.316997Z digest=sha256:21607d47889585593910c13ee5cddae42cdf69999b7690294416e3fe8980e3c6

Observation 65cb87c9-ba34-4c34-ad75-5da0bc51ba73 · outbound

This paper cites Collecting highly parallel data for paraphrase evaluation.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Collecting highly parallel data for paraphrase evaluation

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:20.454974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T15:17:19.321812Z digest=sha256:20161247e7aa571760bcaca4024a22151be2e22812fd91bebcb06e894470e314

Observation df3283eb-4378-402f-ac2e-3d494be2ce4c · outbound

This paper cites Fewer Tokens and Fewer Videos: Extending Video Understanding Abilities in Large Vision-Language Models.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Fewer Tokens and Fewer Videos: Extending Video Understanding Abilities in Large Vision-Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:19.326998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:19.326998Z digest=sha256:70ad57d497f820b38776107af68c75859850291671f31022caf42d36bd0efbaa

Observation 0df4e6c6-1ce9-4d70-8453-9f03fbf54617 · outbound

This paper cites HALC: Object Hallucination Reduction via Adaptive Focal-Contrast Decoding.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding HALC: Object Hallucination Reduction via Adaptive Focal-Contrast Decoding

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:19.332297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:19.332297Z digest=sha256:08049c837458ca91bf19af71336a4b40be7a3a2f2874972b0006dd8be6d41786

Observation 522e2b5f-218d-4e71-a728-a82a41ccd501 · outbound

This paper cites VideoLLaMA 2: Advanc- ing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs, 2024.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding VideoLLaMA 2: Advanc- ing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:20.439033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T15:17:19.337077Z digest=sha256:a0e3cfd23f70433ece8486f6c99b1945b4b1e735cee1c1074b28bac6998bba48

Observation 6584ecec-8fcd-411b-a0b5-d29c20b22968 · outbound

This paper cites Video-MME: The First- Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis, 2024.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Video-MME: The First- Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis, 2024

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:20.422508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T15:17:19.341685Z digest=sha256:acbf74ed7d0459f66c873c7d27ea0f7a39f96a520e64c63f0b7b5d2f732faa1b

Observation e97becb7-a974-43ac-8878-56c3eda69ca5 · outbound

This paper cites VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:19.346184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:19.346184Z digest=sha256:8a04f1c858db47e7e9fec635ccbf6a54bf705dbb5e2bc5b7bdc603fdb76153cc

Observation d13b6b4d-4bfe-4311-b0fd-4d9bbd12c0f1 · outbound

This paper cites Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding, 2024.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:20.404545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T15:17:19.351596Z digest=sha256:5d710395df71000999caca1a2a2a2073ad2f27b39f7191666a02d4595a74f068

Observation 34f17add-0dea-426e-8c3b-acf3a602a184 · outbound

This paper cites Evaluating Open-Domain Question Answering in the Era of Large Language Models.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Evaluating Open-Domain Question Answering in the Era of Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:19.356837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:19.356837Z digest=sha256:c65ad4bcffb42bd0f8e48872ee7252d37ece2a0038816da89616559d59ceb5de

Observation 94a946b4-2d98-48ea-8c81-090b8683b43c · outbound

This paper cites An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM, 2024.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM, 2024

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:20.382311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T15:17:19.362862Z digest=sha256:233938565edeafae9ba4831dbe4a6b74393c463a9b9c22b198d0d7b217f4348e

Observation b368254a-fb71-477d-99c1-1bb6d90eb9d5 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:19.367588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:19.367588Z digest=sha256:b05f046e42ae29c933c72eaddaadfcb6d6deb90e05ce8a01602c98fc14002c97

Observation 126393c9-28ef-420b-9a0b-77f28d11f8cc · outbound

This paper cites Inten- tqa: Context-aware video intent reasoning.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Inten- tqa: Context-aware video intent reasoning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:20.360658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T15:17:19.373377Z digest=sha256:bf2f19daad15021e5cc6c348a2a173968626dbb096529d3059effa365f7cb0d3

Observation 1fb44f72-8ff2-45f3-89fb-3b0efa5067e9 · outbound

This paper cites Uniformer: Unified transformer for efficient spatiotemporal representation learning, 2022.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Uniformer: Unified transformer for efficient spatiotemporal representation learning, 2022

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:20.335870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T15:17:19.379498Z digest=sha256:3e473c208e233d5b2606eef70e066a6bcad875e3a49f7aad9d364b9d6bffef54

Observation 45b39613-a40d-4e1c-becf-74ee4e5336e2 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding, 2024.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding VideoChat: Chat-Centric Video Understanding, 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:20.321606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T15:17:19.384344Z digest=sha256:ea8f855adbc0c9d6fe408b189a3d0d6e4a7dbcbf680db5c7ddfcd299c733c20a

Observation c968a893-9c6c-45af-9fa5-73e2e717b92b · outbound

This paper cites MVBench: A Comprehensive Multi- modal Video Understanding Benchmark, 2024.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding MVBench: A Comprehensive Multi- modal Video Understanding Benchmark, 2024

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:20.302761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T15:17:19.388360Z digest=sha256:826b62815c75fae5dc9d52a5688b410e8b60ec845763d6f3f3263509261af58a

Observation f95d33bb-a7a4-421d-9b07-971045cc1870 · outbound

This paper cites Tgif: A new dataset and benchmark on animated gif description.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Tgif: A new dataset and benchmark on animated gif description

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:20.285900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T15:17:19.392754Z digest=sha256:567b09a841f32baa73351d00b5283a7989042a9892ce6eaf7693fde4759545cb

Observation 40103667-a67b-4cca-955c-0a043a37b599 · outbound

This paper cites LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models, 2023.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models, 2023

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:20.269655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T15:17:19.396866Z digest=sha256:6640f4a2236bcf82ce134339346a122e20cfc0b156f498e358950fa156b4f3d1

Observation 2f2d01cb-2286-4b52-9849-91ba195fdb1c · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Llama-vid: An image is worth 2 tokens in large language models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:20.254266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T15:17:19.400928Z digest=sha256:6885755bc93e1f5e23b01bbdd6322dfa502f616a36d5bb5957c08614e2e34a6f

Observation 2af3598e-f90c-4b78-9ef6-6f98a5f579b8 · outbound

This paper cites Video-LLaV A: Learning United Visual Representation by Alignment Before Projection, 2023.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Video-LLaV A: Learning United Visual Representation by Alignment Before Projection, 2023

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:20.237582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T15:17:19.405590Z digest=sha256:19020fc1f2004830fa17af37cd3c91c5c75d7b64fe4754e6b1fcb3e57067514a

Observation 27d28afd-2203-4e27-afe6-7f9115ebeb05 · outbound

This paper cites Tsm: Temporal shift module for efficient video understanding.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Tsm: Temporal shift module for efficient video understanding

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:20.217520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T15:17:19.409359Z digest=sha256:ac1810518ed37e908e75ff7ca91529bbd1f120bdab22fbe8028373595da54f57

Observation be1c8c6a-6027-42fc-b188-e631b843d082 · outbound

This paper cites Video swin transformer.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Video swin transformer

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:19.413389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:19.413389Z digest=sha256:e5e354c06ef0b96c8afa5eaabcc4f53c956dfa44fe394675ae23b3adb926e1f1

Observation c5bb445c-ae61-4b73-bc79-99f772d81188 · outbound

This paper cites Vista-LLaMA: Reliable Video Narrator via Equal Distance to Visual Tokens, 2023.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Vista-LLaMA: Reliable Video Narrator via Equal Distance to Visual Tokens, 2023

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:20.188660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T15:17:19.417067Z digest=sha256:84ff06ccb816066f02d2897b19055843a8f17ee761b583260ecde0ca8ed5ac77

Observation 292b679e-7e22-440e-9a36-3bc0a2e54d21 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:19.420733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:19.420733Z digest=sha256:bf57ff18fc036a1a5f0d3a21c41c3bf95e971cf1f3cfc970cd584d44e46f52bf

Observation 9c2212ef-61a1-4c2c-9eeb-a7056a73a074 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Mod- els, 2023.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Mod- els, 2023

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:20.170409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T15:17:19.424893Z digest=sha256:c03842a18722bdf62a44f38dc1ecc19c25631c4dc9b073f7562329033c686330

Observation a76213b4-79b5-476e-805c-d43f79e92590 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:20.156272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T15:17:19.428685Z digest=sha256:99f94410f801c4a1457d822659e55aae8ad30c8f0f121d02674345c79e464838

Observation 09290a55-b8f4-42ff-82c4-3002a6a7c6e5 · outbound

This paper cites Foundation Models for Video Understanding: A Survey.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Foundation Models for Video Understanding: A Survey

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:19.433014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:19.433014Z digest=sha256:f55d3a04bc17255437d28c65d34e33cda65f69f331e7bf1d12f0192c24913996

Observation 00d517ab-25a8-4355-8797-38d79fdd9fc8 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long- 9 form video language understanding.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Egoschema: A diagnostic benchmark for very long- 9 form video language understanding

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:20.139469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T15:17:19.437718Z digest=sha256:12e19e032c4fd9f6606211809823735cb450e9c5cc36a4582d17ff0d4746eaad

Observation 3ec775d1-81f7-4664-bf67-6c7f25c5f6df · outbound

This paper cites MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:19.441536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:19.441536Z digest=sha256:ccd71928579dec048eef8eec232d0ebba8b3653ff29659471e1229b3c9751e85

Observation 698362b6-d6f1-4625-83ea-a10dffb579ac · outbound

This paper cites Less is more: Pay less attention in vision transform- ers.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Less is more: Pay less attention in vision transform- ers

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:20.123968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T15:17:19.445865Z digest=sha256:42d12b2b1cf4036861dafdd71a6f5304d1eeec2700ee66a51a12a7a0e376aaeb

Observation 8ba4bb6d-d006-4869-a91d-94c5e1a9bccb · outbound

This paper cites Effi- cient parameter-free clustering using first neighbor relations.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Effi- cient parameter-free clustering using first neighbor relations

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:20.107949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T15:17:19.449954Z digest=sha256:0c3437e4ef1219df73680fb44d7f5775d8c1b7a71f6cac9b9289b2a4f7889407

Observation bd5e2a6f-3fa4-4272-8a7c-5b678a64b268 · outbound

This paper cites MovieChat: From Dense Token to Sparse Memory for Long Video Understanding, 2024.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding MovieChat: From Dense Token to Sparse Memory for Long Video Understanding, 2024

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:20.094030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T15:17:19.454068Z digest=sha256:31a7864af575a3a0836aa1f5a106ecad03a2bd8b362918a2d1143614f44dfbc2

Observation dd73bd8a-9856-4504-a44a-673f2a94970c · outbound

This paper cites Video understanding with large language models: A survey.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Video understanding with large language models: A survey

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:19.457831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:19.457831Z digest=sha256:3827145eb72b81cd2269ba3ed3438d2aa3561d624b925caa2753664ef60eba77

Observation d753d32a-3dbd-45bb-9e13-9d78255acdce · outbound

This paper cites Internvideo2: Scaling foundation models for mul- timodal video understanding.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Internvideo2: Scaling foundation models for mul- timodal video understanding

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:20.079942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T15:17:19.462316Z digest=sha256:59d21cfc29a3ee89d040abf8285570a19f24ef3e1c5d0a2483e0807ef6a70fec

Observation db43c629-b1c1-4ccf-a52b-a567266de72e · outbound

This paper cites Freeva: Offline mllm as training-free video assistant.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Freeva: Offline mllm as training-free video assistant

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:20.064393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T15:17:19.466380Z digest=sha256:69deca607074b40abd3594aeb1cad47ce5f5e92577107c9c8943ecbd7878a73a

Observation 2bed5b56-0624-48f6-aa5b-886f14588489 · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Next-qa: Next phase of question-answering to explaining temporal actions

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:20.047090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T15:17:19.470673Z digest=sha256:6ea6d58dfdd52933d335241118a5510741dc1882016729b18731b219a97135f4

Observation 0d929b18-6a74-46fd-a4ee-455e44dc4ca0 · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Msr-vtt: A large video description dataset for bridging video and language

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:20.027795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T15:17:19.476488Z digest=sha256:d145f5587df3ef6013f7887361db0274c6cc0d09163ba5578a5fca4828a282dc

Observation 14923024-3d8d-418f-8728-03e3bee1f7c7 · outbound

This paper cites PLLaV A : Parameter-free LLaV A Extension from Images to Videos for Video Dense Captioning, 2024.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding PLLaV A : Parameter-free LLaV A Extension from Images to Videos for Video Dense Captioning, 2024

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:20.010444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T15:17:19.481713Z digest=sha256:00f12d125d72223c0ae3cbcf0186a7772530074885da1c6bcc37ca3d98014793

Observation 7399f2ba-5ace-47e2-b8b8-63840eade38a · outbound

This paper cites SlowFast-LLaV A: A Strong Training-Free Baseline for Video Large Language Models, 2024.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding SlowFast-LLaV A: A Strong Training-Free Baseline for Video Large Language Models, 2024

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:19.982895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T15:17:19.485762Z digest=sha256:4743933fe7ae7e1f4d2f6b08da20ee1f64a50d51fd991d8b7b586615b17443fa

Observation 69bd3635-6fdf-4f89-b31f-2dbe3a87845e · outbound

This paper cites Self-chained image-language model for video localization and question answering.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Self-chained image-language model for video localization and question answering

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:19.964391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T15:17:19.493075Z digest=sha256:be510241cd60a6cd1a9fe7eb74aeaa98b4b2858b9b809daf3dca9ba31ddf8bf5

Observation 76c52acc-6394-407b-9ea5-ec27461b406d · outbound

This paper cites Self-chained image-language model for video localization and question answering.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Self-chained image-language model for video localization and question answering

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:19.946256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T15:17:19.498409Z digest=sha256:eb0347f50d63235f214bdf90298465639dddfadce29f14ad0429a716b666bdda

Observation ed617feb-2b64-4461-a4a3-8a9f2adf14dd · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:19.930947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T15:17:19.502684Z digest=sha256:791e36575eaa8cc2113e0af4c5ea2d234f68b1f67af1069451f916d98a1707cb

Observation e2cf9dd5-bb5c-4c4a-8943-43c299e5ecbe · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding, 2023.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding, 2023

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:19.916189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T15:17:19.507722Z digest=sha256:d4227ac7b7fa7aae43bb4a237bd660594483604e78b5becb3442d21de75f4e2e

Observation 84dfa728-6de5-460e-b61f-942252e29a55 · outbound

This paper cites Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:19.513134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:19.513134Z digest=sha256:2c5593cbb08a8a1f6e3af7402b5cbda4004e89a91946889393fa0fe874b4111b

Observation 4ade763f-3bac-4b4c-987a-fe95caf799b4 · outbound

This paper cites Llava- next: A strong zero-shot video understanding model, 2024.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Llava- next: A strong zero-shot video understanding model, 2024

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:19.517778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:19.517778Z digest=sha256:93984dfe748b97d223471dfac2d85dc5fa5e030286e7cb28e87e5136c0fd29cf

Observation caba4785-ef58-4e5d-91a3-f8d139b4d05f · outbound

This paper cites Llava- next: A strong zero-shot video understanding model, 2024.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Llava- next: A strong zero-shot video understanding model, 2024

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:19.522013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:19.522013Z digest=sha256:803061f46dc39af6d9cad71eee46c0f74528d95c554520974a338478a551e41a

Observation c64262c1-0a17-4e9b-be81-f3fbb8d16758 · outbound

This paper cites RankCLIP: Ranking-Consistent Language-Image Pretraining.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding RankCLIP: Ranking-Consistent Language-Image Pretraining

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:19.526139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:19.526139Z digest=sha256:76621d8146d36126bab715c29dad4d38a36772ce8f1f57cd71d64e6cea9bb70d

Observation 46dc1631-eb8a-425b-998d-de6f1d6bcb56 · outbound

This paper cites Multimodal guidance network for missing- modality inference in content moderation.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Multimodal guidance network for missing- modality inference in content moderation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:19.884148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T15:17:19.530734Z digest=sha256:b38ce7fc71e2692d97290d6ad2e4e97af51af38363f4a4ce93ae2d3acf075c42

Observation 23d25d77-fef8-4ec5-8ece-71083439c966 · outbound

This paper cites A Survey on Generative AI and LLM for Video Generation, Understanding, and Streaming.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding A Survey on Generative AI and LLM for Video Generation, Understanding, and Streaming

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:19.535023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:19.535023Z digest=sha256:225e43ab4f59d548adcc0695375cf26a6125761e7b4c27fa9baac0119d20e002

Pith citing papers

Observation 3eea6c55-aa84-4112-8e27-05f5de688bc8 · inbound

Mosaic: Cross-Modal Clustering for Efficient Video Understanding cites this paper.

Mosaic: Cross-Modal Clustering for Efficient Video Understanding Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-10T16:10:34.350836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T16:07:36.133404Z digest=sha256:d25ed0af893fd1464c16920477c646a69718a41ff062c180c63c996f507780b0

Observation 01acc121-5f1e-419a-bd97-6bb4d4ed9dca · inbound

GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models cites this paper.

GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:17:50.839649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-14T19:15:13.205594Z digest=sha256:a63d33827d270315121bef593b94667630a22d64d2ed923eb53128cdbc1ef08d

Observation adcf274d-f2a2-4a79-bf6c-9297ffb9eab4 · inbound

video-SALMONN-R$^3$: Learning to ReWatch, ReAsk, and ReAnswer for Efficient Video Understanding cites this paper.

video-SALMONN-R$^3$: Learning to ReWatch, ReAsk, and ReAnswer for Efficient Video Understanding Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:39:58.352126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T00:19:26.153682Z digest=sha256:2eed0d82befb6fef64e9ca9241e8b3dff560c6d876f75963c56d560af9d12add