Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T15:17:19.535023Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 3 inbound Pith citation observations for arXiv:2411.14401.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T15:17:19.535023Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-26T00:19:26.153682Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-10T06:15:00.866473Z
49 of 49 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4775f871-409b-495c-ad49-c996aa385278 · outbound
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Is space-time attention all you need for video understanding? In ICML, page 4, 2021
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9894c35c-047e-4413-9367-df150fa01062 · outbound
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Token merging: Your ViT but faster
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65cb87c9-ba34-4c34-ad75-5da0bc51ba73 · outbound
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Collecting highly parallel data for paraphrase evaluation
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation df3283eb-4378-402f-ac2e-3d494be2ce4c · outbound
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Fewer Tokens and Fewer Videos: Extending Video Understanding Abilities in Large Vision-Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0df4e6c6-1ce9-4d70-8453-9f03fbf54617 · outbound
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding HALC: Object Hallucination Reduction via Adaptive Focal-Contrast Decoding
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 522e2b5f-218d-4e71-a728-a82a41ccd501 · outbound
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding VideoLLaMA 2: Advanc- ing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs, 2024
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6584ecec-8fcd-411b-a0b5-d29c20b22968 · outbound
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Video-MME: The First- Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis, 2024
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e97becb7-a974-43ac-8878-56c3eda69ca5 · outbound
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d13b6b4d-4bfe-4311-b0fd-4d9bbd12c0f1 · outbound
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding, 2024
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 34f17add-0dea-426e-8c3b-acf3a602a184 · outbound
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Evaluating Open-Domain Question Answering in the Era of Large Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94a946b4-2d98-48ea-8c81-090b8683b43c · outbound
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM, 2024
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b368254a-fb71-477d-99c1-1bb6d90eb9d5 · outbound
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding LLaVA-OneVision: Easy Visual Task Transfer
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 126393c9-28ef-420b-9a0b-77f28d11f8cc · outbound
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Inten- tqa: Context-aware video intent reasoning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1fb44f72-8ff2-45f3-89fb-3b0efa5067e9 · outbound
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Uniformer: Unified transformer for efficient spatiotemporal representation learning, 2022
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 45b39613-a40d-4e1c-becf-74ee4e5336e2 · outbound
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding VideoChat: Chat-Centric Video Understanding, 2024
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c968a893-9c6c-45af-9fa5-73e2e717b92b · outbound
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding MVBench: A Comprehensive Multi- modal Video Understanding Benchmark, 2024
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f95d33bb-a7a4-421d-9b07-971045cc1870 · outbound
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Tgif: A new dataset and benchmark on animated gif description
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 40103667-a67b-4cca-955c-0a043a37b599 · outbound
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models, 2023
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2f2d01cb-2286-4b52-9849-91ba195fdb1c · outbound
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Llama-vid: An image is worth 2 tokens in large language models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2af3598e-f90c-4b78-9ef6-6f98a5f579b8 · outbound
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Video-LLaV A: Learning United Visual Representation by Alignment Before Projection, 2023
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 27d28afd-2203-4e27-afe6-7f9115ebeb05 · outbound
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Tsm: Temporal shift module for efficient video understanding
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation be1c8c6a-6027-42fc-b188-e631b843d082 · outbound
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Video swin transformer
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5bb445c-ae61-4b73-bc79-99f772d81188 · outbound
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Vista-LLaMA: Reliable Video Narrator via Equal Distance to Visual Tokens, 2023
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 292b679e-7e22-440e-9a36-3bc0a2e54d21 · outbound
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c2212ef-61a1-4c2c-9eeb-a7056a73a074 · outbound
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Mod- els, 2023
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a76213b4-79b5-476e-805c-d43f79e92590 · outbound
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Video-chatgpt: Towards detailed video understanding via large vision and language models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 09290a55-b8f4-42ff-82c4-3002a6a7c6e5 · outbound
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Foundation Models for Video Understanding: A Survey
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00d517ab-25a8-4355-8797-38d79fdd9fc8 · outbound
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Egoschema: A diagnostic benchmark for very long- 9 form video language understanding
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3ec775d1-81f7-4664-bf67-6c7f25c5f6df · outbound
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 698362b6-d6f1-4625-83ea-a10dffb579ac · outbound
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Less is more: Pay less attention in vision transform- ers
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8ba4bb6d-d006-4869-a91d-94c5e1a9bccb · outbound
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Effi- cient parameter-free clustering using first neighbor relations
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation bd5e2a6f-3fa4-4272-8a7c-5b678a64b268 · outbound
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding MovieChat: From Dense Token to Sparse Memory for Long Video Understanding, 2024
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation dd73bd8a-9856-4504-a44a-673f2a94970c · outbound
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Video understanding with large language models: A survey
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d753d32a-3dbd-45bb-9e13-9d78255acdce · outbound
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Internvideo2: Scaling foundation models for mul- timodal video understanding
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation db43c629-b1c1-4ccf-a52b-a567266de72e · outbound
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Freeva: Offline mllm as training-free video assistant
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2bed5b56-0624-48f6-aa5b-886f14588489 · outbound
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Next-qa: Next phase of question-answering to explaining temporal actions
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0d929b18-6a74-46fd-a4ee-455e44dc4ca0 · outbound
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Msr-vtt: A large video description dataset for bridging video and language
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 14923024-3d8d-418f-8728-03e3bee1f7c7 · outbound
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding PLLaV A : Parameter-free LLaV A Extension from Images to Videos for Video Dense Captioning, 2024
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7399f2ba-5ace-47e2-b8b8-63840eade38a · outbound
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding SlowFast-LLaV A: A Strong Training-Free Baseline for Video Large Language Models, 2024
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 69bd3635-6fdf-4f89-b31f-2dbe3a87845e · outbound
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Self-chained image-language model for video localization and question answering
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 76c52acc-6394-407b-9ea5-ec27461b406d · outbound
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Self-chained image-language model for video localization and question answering
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ed617feb-2b64-4461-a4a3-8a9f2adf14dd · outbound
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Activitynet-qa: A dataset for understanding complex web videos via question answering
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e2cf9dd5-bb5c-4c4a-8943-43c299e5ecbe · outbound
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding, 2023
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 84dfa728-6de5-460e-b61f-942252e29a55 · outbound
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ade763f-3bac-4b4c-987a-fe95caf799b4 · outbound
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Llava- next: A strong zero-shot video understanding model, 2024
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation caba4785-ef58-4e5d-91a3-f8d139b4d05f · outbound
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Llava- next: A strong zero-shot video understanding model, 2024
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c64262c1-0a17-4e9b-be81-f3fbb8d16758 · outbound
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding RankCLIP: Ranking-Consistent Language-Image Pretraining
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46dc1631-eb8a-425b-998d-de6f1d6bcb56 · outbound
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Multimodal guidance network for missing- modality inference in content moderation
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 23d25d77-fef8-4ec5-8ece-71083439c966 · outbound
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding A Survey on Generative AI and LLM for Video Generation, Understanding, and Streaming
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3eea6c55-aa84-4112-8e27-05f5de688bc8 · inbound
Mosaic: Cross-Modal Clustering for Efficient Video Understanding Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 01acc121-5f1e-419a-bd97-6bb4d4ed9dca · inbound
GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation adcf274d-f2a2-4a79-bf6c-9297ffb9eab4 · inbound
video-SALMONN-R$^3$: Learning to ReWatch, ReAsk, and ReAnswer for Efficient Video Understanding Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.