Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T20:06:35.255089Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 16 inbound Pith citation observations for arXiv:2502.05173.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T20:06:35.255089Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-08T17:26:28.859252Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T06:39:37.598350Z
39 of 39 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b5837546-afe1-4b76-bd4d-509e4deaf5f0 · outbound
VideoRoPE: What Makes for Good Video Rotary Position Embedding? AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2077bea-ac2a-447b-a03e-e843b71bea30 · outbound
VideoRoPE: What Makes for Good Video Rotary Position Embedding? ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07a9bb7b-6e13-4cfc-b409-9f359f324945 · outbound
VideoRoPE: What Makes for Good Video Rotary Position Embedding? The Llama 3 Herd of Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a484ea4-8fc3-41f7-9399-a6d86d273de6 · outbound
VideoRoPE: What Makes for Good Video Rotary Position Embedding? VideoHallucer classifies hallucinations into two primary types: intrinsic and extrinsic
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d5a95be9-768f-49d8-8e48-89b5465b8e14 · outbound
VideoRoPE: What Makes for Good Video Rotary Position Embedding? TC-LLaVA: Rethinking the Transfer from Image to Video Understanding with Temporal Considerations
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5b65fabf-b039-4a63-85c7-3f1263918e07 · outbound
VideoRoPE: What Makes for Good Video Rotary Position Embedding? Accessed: 2025-01-12
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9e69fc2f-6a95-4bfe-b89b-12d3a77ec5cb · outbound
VideoRoPE: What Makes for Good Video Rotary Position Embedding? RULER: What's the Real Context Size of Your Long-Context Language Models?
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76ce522a-51f9-4e33-8d59-0eba8a33ac6e · outbound
VideoRoPE: What Makes for Good Video Rotary Position Embedding? VTimeLLM: Empower LLM to Grasp Video Moments
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ef1c0de-4222-4e57-86c7-2db6cc064039 · outbound
VideoRoPE: What Makes for Good Video Rotary Position Embedding? OPERA: Alleviating Hallucination in Multi-Modal Large Language Models via Over-Trust Penalty and Retrospection-Allocation
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23cc5c99-1f77-4c8e-bf47-d466a50145bb · outbound
VideoRoPE: What Makes for Good Video Rotary Position Embedding? The Kinetics Human Action Video Dataset
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b932cf2e-e01e-46e7-920c-29be7c7953f8 · outbound
VideoRoPE: What Makes for Good Video Rotary Position Embedding? Less is More: ClipBERT for Video-and-Language Learning via Sparse Sampling
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69c19d6d-75bc-4eab-a81e-bf1ba7233419 · outbound
VideoRoPE: What Makes for Good Video Rotary Position Embedding? Temporal Reasoning Transfer from Text to Video
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff4f5f3d-1bd8-4344-9e74-500a973541fc · outbound
VideoRoPE: What Makes for Good Video Rotary Position Embedding? OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e1d8574-032b-4117-ab75-eb7050bd2c38 · outbound
VideoRoPE: What Makes for Good Video Rotary Position Embedding? Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f387f90c-54ba-4c63-ba1a-e9d1f17789f5 · outbound
VideoRoPE: What Makes for Good Video Rotary Position Embedding? YaRN: Efficient Context Window Extension of Large Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a52dff9-a257-442d-8a57-e889f9000459 · outbound
VideoRoPE: What Makes for Good Video Rotary Position Embedding? Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84713aff-4810-4419-91c5-4a46ad3647e9 · outbound
VideoRoPE: What Makes for Good Video Rotary Position Embedding? Learning Transferable Visual Models From Natural Language Supervision
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17f81594-282b-48cc-9b24-50ae266d3f7a · outbound
VideoRoPE: What Makes for Good Video Rotary Position Embedding? doi: 10.1007/s11633-024-1502-8
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b43b4771-904a-462c-b813-dbae6baa87fc · outbound
VideoRoPE: What Makes for Good Video Rotary Position Embedding? Gemma: Open Models Based on Gemini Research and Technology
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec2a373d-f810-4828-a02f-1d37a7f2e28d · outbound
VideoRoPE: What Makes for Good Video Rotary Position Embedding? ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3afbc6e1-ba68-4238-9002-d88ba44221df · outbound
VideoRoPE: What Makes for Good Video Rotary Position Embedding? PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f45b970-b8b4-4627-9524-83ad999afd2a · outbound
VideoRoPE: What Makes for Good Video Rotary Position Embedding? VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 072338a1-9646-41b7-81b0-5e858610cf05 · outbound
VideoRoPE: What Makes for Good Video Rotary Position Embedding? PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06e9e869-3a66-4c50-8bca-eb4abcc7a6d2 · outbound
VideoRoPE: What Makes for Good Video Rotary Position Embedding? Qwen2 Technical Report
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4118f63-dd22-44cc-943a-d3c28a5d7eef · outbound
VideoRoPE: What Makes for Good Video Rotary Position Embedding? InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b51070d-d63d-4fc0-976a-77c714bd43a5 · outbound
VideoRoPE: What Makes for Good Video Rotary Position Embedding? Long-CLIP: Unlocking the Long-Text Capability of CLIP
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5d1cee3-db0c-4fd8-b453-1aa778b5ed74 · outbound
VideoRoPE: What Makes for Good Video Rotary Position Embedding? LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78137972-cb66-4640-99b0-572721a6272b · outbound
VideoRoPE: What Makes for Good Video Rotary Position Embedding? , y, y, y,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fb0685fe-619e-4eb5-9965-c775bdf2a6bd · outbound
VideoRoPE: What Makes for Good Video Rotary Position Embedding? Extending these models to video requires handling temporal dependencies (Xu et al., 2021; Lei et al., 2021; Bertasius et al., 2021; Huang et al.,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 74a37008-fa0c-4e64-a6d0-9b890e4bda00 · outbound
VideoRoPE: What Makes for Good Video Rotary Position Embedding? Early video LLMs, such as Wang et al
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d6d43b69-749f-471f-88cc-e8ecd5dc073f · outbound
VideoRoPE: What Makes for Good Video Rotary Position Embedding? However, video haystack retrieval still lags in difficulty compared with the QA or retrieval task in NLP (Hsieh et al., 2024; Yuan et al., 2024)
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 208dd9a1-187b-4412-85a0-d387598f8d00 · outbound
VideoRoPE: What Makes for Good Video Rotary Position Embedding? Unresolved cited work
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d8f62f04-0191-4f40-b2c4-170ecfbe2b63 · outbound
VideoRoPE: What Makes for Good Video Rotary Position Embedding? On the other hand, VideoRoPE employs low-frequency temporal modeling, allowing it to capture long-range dependencies and successfully identify the needle for accurate responses
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 42257403-172b-49d9-ab51-fb888b0324de · outbound
VideoRoPE: What Makes for Good Video Rotary Position Embedding? Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 2015
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c564a8a0-6674-415b-bf39-4b706af6b661 · outbound
VideoRoPE: What Makes for Good Video Rotary Position Embedding? Is Space-Time Attention All You Need for Video Understanding?
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc9bac64-023a-4caa-b918-1bec63136a6a · outbound
VideoRoPE: What Makes for Good Video Rotary Position Embedding? InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c48653a7-17e7-4649-b1f8-6273adf43fdc · outbound
VideoRoPE: What Makes for Good Video Rotary Position Embedding? ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5beaa9dd-d781-4b65-9da0-1da05671958c · outbound
VideoRoPE: What Makes for Good Video Rotary Position Embedding? Pixtral 12B
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcb314c8-8f4f-48c6-94cf-730a25f6fc0d · outbound
VideoRoPE: What Makes for Good Video Rotary Position Embedding? LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c8f2209-5d3c-42fd-ab0c-1e4749873c54 · inbound
ClinKD: Cross-Modal Clinical Knowledge Distiller For Multi-Task Medical Images VideoRoPE: What Makes for Good Video Rotary Position Embedding?
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d0f781a-f4a6-4cfc-91ce-56a209cd6d6a · inbound
Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation VideoRoPE: What Makes for Good Video Rotary Position Embedding?
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a78a7883-6d7d-4472-902f-582dae61d412 · inbound
Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models VideoRoPE: What Makes for Good Video Rotary Position Embedding?
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 72d0da19-90bd-49a7-bad9-1dce2fcd7bc2 · inbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization VideoRoPE: What Makes for Good Video Rotary Position Embedding?
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a909771-f00d-4697-acda-ccf2c2a7a875 · inbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification VideoRoPE: What Makes for Good Video Rotary Position Embedding?
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 333a114b-7802-40da-aee7-bd656b079e76 · inbound
Task-Aware KV Compression For Cost-Effective Long Video Understanding VideoRoPE: What Makes for Good Video Rotary Position Embedding?
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f519b7f8-72c3-4347-acda-48ca2a008599 · inbound
Infinite Video Understanding VideoRoPE: What Makes for Good Video Rotary Position Embedding?
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 403e88a9-0bf6-44aa-8100-3224b1e886ed · inbound
Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture VideoRoPE: What Makes for Good Video Rotary Position Embedding?
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bccfc75-5faa-4730-9dad-5d0766c30932 · inbound
Phase-Aligned RoPE for Mixed-Resolution Diffusion Transformer VideoRoPE: What Makes for Good Video Rotary Position Embedding?
Reference 115
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 267f2f9f-3cd4-4d5e-b8d4-ef00846bdc9d · inbound
Adaptive 3D-RoPE: Physics-Aligned Rotary Positional Encoding for Wireless Foundation Models VideoRoPE: What Makes for Good Video Rotary Position Embedding?
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 12933e0b-01a1-4e67-bb01-f81fd46e42b4 · inbound
Diffusing in the Right Space: A Systematic Study of Latent Diffusability VideoRoPE: What Makes for Good Video Rotary Position Embedding?
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2a713525-41ae-4314-9492-796073ef8e62 · inbound
LiteVSR: Lightweight Adaptation of Frozen Diffusion Transformers for Video Super-Resolution VideoRoPE: What Makes for Good Video Rotary Position Embedding?
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation afdacdee-a5a5-485e-8d91-d3428a46e6e3 · inbound
HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning VideoRoPE: What Makes for Good Video Rotary Position Embedding?
Reference 130
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5f6635f3-64b7-44be-b639-c611f485d09c · inbound
ShotPlan: Cinematic Video Generation with Learnable Planning Token VideoRoPE: What Makes for Good Video Rotary Position Embedding?
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac47fea1-aa70-403c-8d35-88a576e228ee · inbound
ChronoStitch: Training-Free Composition of Visual KV Memories for Long-Horizon Temporal Reasoning VideoRoPE: What Makes for Good Video Rotary Position Embedding?
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 387af61f-359b-4b3a-8569-d5f16007ec61 · inbound
RealVDeblur: One-Step Diffusion for Generalizable Real-World Video Deblurring VideoRoPE: What Makes for Good Video Rotary Position Embedding?
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.