Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:59:11.354129Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 5 inbound Pith citation observations for arXiv:2506.00993.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:59:11.354129Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T23:42:35.141746Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-12T10:11:28.172880Z
52 of 52 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a2b7e372-c127-4f3a-90d4-efd27c3a854d · outbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Qwen2.5-vl technical report,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 142e73d1-bebc-49bf-aba7-c98074de1e36 · outbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Fast Differentiable Sorting and Ranking
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c9fb958-105e-4065-a867-de5b55fd6aa6 · outbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Principal component analysis.Analytical methods, 6(9): 2812–2831, 2014
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ec529799-0e10-4035-b845-b9cc040f1b06 · outbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fbfa54c-6909-484a-814c-ee2ecabe4087 · outbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 765a8170-44c3-4b5f-8f69-cb30a2eebca8 · outbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d536abaf-3295-4fbb-a432-666bc4a20e34 · outbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Scaling Video-Language Models to 10K Frames via Hierarchical Differential Distillation
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a6a7819-c052-4f17-a2a0-c9e20b22d499 · outbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding ReWind: Understanding Long Videos with Instructed Learnable Memory
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cf1cc51-d3b1-4ccd-95ef-a55c6c0a8e66 · outbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Feather the Throttle: Revisiting Visual Token Pruning for Vision-Language Model Acceleration
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 889779c6-2937-4cab-bb15-295b9ff8aeea · outbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation edb3e71d-43c9-4f30-aeab-8340890c5497 · outbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1499470-2857-43b4-9b4a-d9b606b4758b · outbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Framefusion: Combining similarity and importance for video token reduction on large visual language models, 2024
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7a28d731-d390-4e5c-9c8d-1465ae651bdd · outbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding LinVT: Empower Your Image-level Large Language Model to Understand Videos
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2fd70d1-368d-4c2e-a2d5-30a39a559de3 · outbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Llava-onevision: Easy visual task transfer,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21cc2dfd-dfd4-44a1-b1c5-5e5e52635d85 · outbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cade06a3-70a4-4910-80ba-31e38fac5da2 · outbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0f4db96-2152-4a32-8749-5ef93fdf9da5 · outbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Temporal preference optimization for long-form video understanding, 2025
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 02ef2d55-2f23-47e2-b7f1-85fa1b1795f0 · outbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45e15b40-6cc5-4825-8b1a-1dea06352fcd · outbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b3bc549-41c2-46f4-95bc-16ad15f0b52f · outbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53815c49-612f-4fdc-ace2-72e165abdbdc · outbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding NVILA: Efficient Frontier Visual Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76f94cf2-f665-450c-bbba-a82b44d337a1 · outbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b223a9c7-6b65-4bea-8703-b7752342912c · outbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding QuoTA: Query-oriented Token Assignment via CoT Query Decouple for Long Video Comprehension
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8c88a420-dec8-4220-ae5d-15760d672cf5 · outbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Hello GPT-4o, 5 2024
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b5051d15-0865-43d9-a9a2-a195f343f37b · outbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Yarn: Efficient context window extension of large language models, 2023
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 44eec4db-7d7a-43fd-86e8-4d3f6f8c8d70 · outbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding UI-TARS: Pioneering Automated GUI Interaction with Native Agents
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1367ce42-ec36-4505-b7ec-4aae10fc5a3b · outbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Video-xl: Extra-long vision language model for hour-scale video understanding,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ac8e8494-8784-4eee-8003-c3e14569b47c · outbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding The proof and measurement of association between two things
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44971cd9-b215-45e0-bfe9-0d8d6dc5fc3c · outbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Dycoke: Dynamic compression of tokens for fast video large language models, 2025
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 358355fd-9998-444c-aae7-b87c1418de3a · outbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1b035eb-1fbc-4349-8fb3-6a2e48470e27 · outbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Qwen2.5 Technical Report
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ee70352-c5af-403b-9cb6-f1beddcc5450 · outbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e3b39db-a13d-4434-9faa-b797beceb8ab · outbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Gemini: A Family of Highly Capable Multimodal Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59d56e0e-84dd-4318-9897-74785c7466c4 · outbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding AdaReTaKe: Adaptive Redundancy Reduction to Perceive Longer for Video-language Understanding
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f35f5d8d-1efb-41e6-96bd-6b0121220b77 · outbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Stop Looking for Important Tokens in Multimodal Language Models: Duplication Matters More
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85466586-ec1c-4645-8ef9-f0792b8a866d · outbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding LVBench: An Extreme Long Video Understanding Benchmark
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03995c8e-38cc-42d9-8556-b7994c2b04e9 · outbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdd0506c-935c-43ec-b472-bfb3ddaf2094 · outbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07b2a234-bd0f-4ea4-832d-8429a0bf96d2 · outbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98b0f7b4-8bb5-4331-ba7a-d661b7b8ed23 · outbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 317dbfad-ce0d-4ebf-b807-df11483ea597 · outbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b56cbe40-4a7d-4461-89f2-913ff3f5f7bc · outbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Long context transfer from language to vision,
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation dfa4e81f-0ea0-49af-aad9-41a331e93e78 · outbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8826dfff-bcbe-4a51-859c-962a6a42fd2b · outbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 947f0de2-438a-4d18-be4e-eb00dfad343f · outbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31892af6-6997-4b2c-9b72-0e9cc0709858 · outbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Long Context Transfer from Language to Vision
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28156de9-2eef-4bfc-95fd-494707660083 · outbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding MLVU: Benchmarking Multi-task Long Video Understanding
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb4cd5d7-3937-4748-ae83-c7fe93378ead · outbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f6d8b62-bb23-49b7-b3fc-4b9f9fd99e74 · outbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ff19277-4d81-42ea-a6ca-e0b1fdb91644 · outbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Apollo: An Exploration of Video Understanding in Large Multimodal Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e10be65-1ea8-4a7a-b4ef-a848e5b51ac4 · outbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding LLaVA-OneVision: Easy Visual Task Transfer
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0add14b5-d3c4-4cbf-becc-c9a872c9f851 · outbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Qwen2.5-VL Technical Report
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aab852e4-f310-444f-81f7-d11913bbc228 · inbound
HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes FlexSelect: Flexible Token Selection for Efficient Long Video Understanding
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad69ed62-b0a9-4128-bfef-26435c8257e2 · inbound
Stateful Token Reduction for Long-Video Hybrid VLMs FlexSelect: Flexible Token Selection for Efficient Long Video Understanding
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8232cef7-438a-4279-b437-41f59df8bfab · inbound
Purifying Multimodal Retrieval: Fragment-Level Evidence Selection for RAG FlexSelect: Flexible Token Selection for Efficient Long Video Understanding
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fd526ce5-1a3b-4b14-91b0-56a7c9883c5e · inbound
Reasoning with Memory: A Temporal Granularity-Adaptive Framework for Training-Free Long Video Understanding FlexSelect: Flexible Token Selection for Efficient Long Video Understanding
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b80a305-06b8-4a90-b461-e9b9d104db82 · inbound
Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding FlexSelect: Flexible Token Selection for Efficient Long Video Understanding
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.