REVIEW 10 cited by
VideoRoPE: What Makes for Good Video Rotary Position Embedding?
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
While Rotary Position Embedding (RoPE) and its variants are widely adopted for their long-context capabilities, the extension of the 1D RoPE to video, with its complex spatio-temporal structure, remains an open challenge. This work first introduces a comprehensive analysis that identifies four key characteristics essential for the effective adaptation of RoPE to video, which have not been fully considered in prior work. As part of our analysis, we introduce a challenging V-NIAH-D (Visual Needle-In-A-Haystack with Distractors) task, which adds periodic distractors into V-NIAH. The V-NIAH-D task demonstrates that previous RoPE variants, lacking appropriate temporal dimension allocation, are easily misled by distractors. Based on our analysis, we introduce \textbf{VideoRoPE}, with a \textit{3D structure} designed to preserve spatio-temporal relationships. VideoRoPE features \textit{low-frequency temporal allocation} to mitigate periodic oscillations, a \textit{diagonal layout} to maintain spatial symmetry, and \textit{adjustable temporal spacing} to decouple temporal and spatial indexing. VideoRoPE consistently surpasses previous RoPE variants, across diverse downstream tasks such as long video retrieval, video understanding, and video hallucination. Our code will be available at \href{https://github.com/Wiselnn570/VideoRoPE}{https://github.com/Wiselnn570/VideoRoPE}.
Forward citations
Cited by 10 Pith papers
-
RealVDeblur: One-Step Diffusion for Generalizable Real-World Video Deblurring
RealVDeblur trains a one-step video-diffusion deblurrer on a large 3DGS-based synthetic dataset and stabilizes long-video inference with a temporal window mask, improving perceptual quality on real-world benchmarks.
-
ChronoStitch: Training-Free Composition of Visual KV Memories for Long-Horizon Temporal Reasoning
ChronoStitch re-bases stored video-chunk KV keys into a global multimodal RoPE frame and selectively recomputes a small slice of high-deviation tokens, recovering most of the joint-prefill temporal-reasoning gap at 3....
-
ShotPlan: Cinematic Video Generation with Learnable Planning Token
Learnable planning tokens with fractional positional timestamps let one diffusion pass generate multi-shot video with frame-accurate cuts and timed camera motion.
-
Phase-Aligned RoPE for Mixed-Resolution Diffusion Transformer
Expressing all RoPE positions on the query's grid ('one attention, one scale') plus a small boundary content-exchange step restores mixed-resolution diffusion generation that naive position interpolation destroys.
-
Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture
Spatial understanding in multimodal LLMs plateaus quickly as training data grows, and position encoding in the visual encoder is the more influential factor.
-
Task-Aware KV Compression For Cost-Effective Long Video Understanding
Video-X2L uses bi-level KV compression with task-aware selective reloading to improve long-video QA accuracy and reduce decode-time memory versus uniform KV compression.
-
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification
Video-XL-2 cuts long-video inference cost with chunked pre-filling and query-gated dense-or-sparse KV reloading, reporting half the FLOPs and a third less decoding memory at roughly equal benchmark scores.
-
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization
EVA02-AT combines full-dimension spatial and temporal rotary position embeddings with a symmetric multi-similarity loss to improve egocentric video-text retrieval.
-
ClinKD: Cross-Modal Clinical Knowledge Distiller For Multi-Task Medical Images
ClinKD combines a modified rotary position embedding, confidence-weighted pseudo-label distillation, and CLIP-based answer selection, reporting state-of-the-art scores on Med-GRIT and LLaVA-Med-QA benchmarks.
-
Infinite Video Understanding
The paper argues that video understanding research should aim at processing streams of arbitrary, unbounded duration and outlines the challenges, directions, and metrics needed.
Discussion (0). Continue with ORCID to comment.