Pith. sign in

Paper Citation Record · LEDGER

VideoRoPE: What Makes for Good Video Rotary Position Embedding?

As of 8 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 16 inbound Pith citation observations for arXiv:2502.05173.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.05173 v3

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T20:06:35.255089Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T17:26:28.859252Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:39:37.598350Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact1
  • verified fuzzy7
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b5837546-afe1-4b76-bd4d-509e4deaf5f0 · outbound

This paper cites AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.088592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.088592Z digest=sha256:2cf1632186ff17ae3c6d74f257d3de76e10fed5ed3d63ec0dafaa3482a49de08

Observation a2077bea-ac2a-447b-a03e-e843b71bea30 · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.093340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.093340Z digest=sha256:562bddb600be2ee64c33a0765477ae69b6c17e7ee90e11fe3d0a129874ad7188

Observation 07a9bb7b-6e13-4cfc-b409-9f359f324945 · outbound

This paper cites The Llama 3 Herd of Models.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? The Llama 3 Herd of Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.107957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.107957Z digest=sha256:dc089dea6f38baddaab155454fbed085ae7cc7bd8c28c7eb098b43f958e671ad

Observation 7a484ea4-8fc3-41f7-9399-a6d86d273de6 · outbound

This paper cites VideoHallucer classifies hallucinations into two primary types: intrinsic and extrinsic.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? VideoHallucer classifies hallucinations into two primary types: intrinsic and extrinsic

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:06:35.887283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:06:35.231736Z digest=sha256:d04b7caac41cd239a5cb7074828c72d7f50a2151f8fb13dd4abef32af6bf00a4

Observation d5a95be9-768f-49d8-8e48-89b5465b8e14 · outbound

This paper cites TC-LLaVA: Rethinking the Transfer from Image to Video Understanding with Temporal Considerations.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? TC-LLaVA: Rethinking the Transfer from Image to Video Understanding with Temporal Considerations

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-08T20:06:35.593964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:06:35.117016Z digest=sha256:588542645e411fb47da87dedfaa0be98d9037201fef43ecaaccb1eab6508386e

Observation 5b65fabf-b039-4a63-85c7-3f1263918e07 · outbound

This paper cites Accessed: 2025-01-12.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? Accessed: 2025-01-12

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:06:35.914386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:06:35.121719Z digest=sha256:8cd559005038ed1bb1b797a38a74b13e1d7a8c1fc4970ecc386998133efd70b8

Observation 9e69fc2f-6a95-4bfe-b89b-12d3a77ec5cb · outbound

This paper cites RULER: What's the Real Context Size of Your Long-Context Language Models?.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? RULER: What's the Real Context Size of Your Long-Context Language Models?

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.126322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.126322Z digest=sha256:95d25363bdcab6121a6ea34da8ee438ed8734e327f8fb97ac5e1d7e0361cf27a

Observation 76ce522a-51f9-4e33-8d59-0eba8a33ac6e · outbound

This paper cites VTimeLLM: Empower LLM to Grasp Video Moments.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? VTimeLLM: Empower LLM to Grasp Video Moments

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.130897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.130897Z digest=sha256:8e5100e5afb9867fc214a7e54b34a4a83e638f7c36d0ad087d21dcd7b8b2e5f5

Observation 0ef1c0de-4222-4e57-86c7-2db6cc064039 · outbound

This paper cites OPERA: Alleviating Hallucination in Multi-Modal Large Language Models via Over-Trust Penalty and Retrospection-Allocation.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? OPERA: Alleviating Hallucination in Multi-Modal Large Language Models via Over-Trust Penalty and Retrospection-Allocation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.135448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.135448Z digest=sha256:e46c7b6c7e6c95fc68fb465efa83c9190a2fe9c74bbcf77f8e4e34877a9310ff

Observation 23cc5c99-1f77-4c8e-bf47-d466a50145bb · outbound

This paper cites The Kinetics Human Action Video Dataset.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? The Kinetics Human Action Video Dataset

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.140418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.140418Z digest=sha256:0a5bf297702333d879de1afc26a9afa858ca00e040779f3a6d2e989934aa18be

Observation b932cf2e-e01e-46e7-920c-29be7c7953f8 · outbound

This paper cites Less is More: ClipBERT for Video-and-Language Learning via Sparse Sampling.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? Less is More: ClipBERT for Video-and-Language Learning via Sparse Sampling

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.144695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.144695Z digest=sha256:19b68f8e0b9f6f62797b5f50d8f96f8c31b1cb8bd44ca8cf31518a24d3fd47a2

Observation 69c19d6d-75bc-4eab-a81e-bf1ba7233419 · outbound

This paper cites Temporal Reasoning Transfer from Text to Video.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? Temporal Reasoning Transfer from Text to Video

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.149060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.149060Z digest=sha256:fd7a09b74606b9491847795fab769654dafdc82b104f60ba64505cab06df1097

Observation ff4f5f3d-1bd8-4344-9e74-500a973541fc · outbound

This paper cites OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.153650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.153650Z digest=sha256:e544111460e62ca23279a7be3046004bf5f3f93980ad17d3fd6b8b5a78a36924

Observation 0e1d8574-032b-4117-ab75-eb7050bd2c38 · outbound

This paper cites Valley: Video Assistant with Large Language model Enhanced abilitY.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.158259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.158259Z digest=sha256:bf760ce233690c4fca341a81b932ec8a3226d676e127b30547d91c55d7290ec0

Observation f387f90c-54ba-4c63-ba1a-e9d1f17789f5 · outbound

This paper cites YaRN: Efficient Context Window Extension of Large Language Models.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? YaRN: Efficient Context Window Extension of Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.162915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.162915Z digest=sha256:cd08fa8c528957dd3ae6df6006e8fa580395b52d650b3aa9151190a00df75e1d

Observation 8a52dff9-a257-442d-8a57-e889f9000459 · outbound

This paper cites Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.167362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.167362Z digest=sha256:d52c918b63a096bb0515b87747e845cc5466732348ff6b8618748017074bdd35

Observation 84713aff-4810-4419-91c5-4a46ad3647e9 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? Learning Transferable Visual Models From Natural Language Supervision

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.172033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.172033Z digest=sha256:e13acacdbf8d18fe2c71169c7b6a8753715edb6b1931563bb929b79582e48363

Observation 17f81594-282b-48cc-9b24-50ae266d3f7a · outbound

This paper cites doi: 10.1007/s11633-024-1502-8.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? doi: 10.1007/s11633-024-1502-8

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.176439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.176439Z digest=sha256:82b1cf2f77cbaf7329d34cb38203f621e6c1f8880d6484ff22f0f05c324bb21e

Observation b43b4771-904a-462c-b813-dbae6baa87fc · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? Gemma: Open Models Based on Gemini Research and Technology

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.180843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.180843Z digest=sha256:9773ac7ba7b8c1fceaa4d92a5cc7e8470a0cabe11cbaa58365bef448ccb48010

Observation ec2a373d-f810-4828-a02f-1d37a7f2e28d · outbound

This paper cites ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.185621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.185621Z digest=sha256:2b80cb84905a445db06ba82fd42ca731b1ec0d5ac75c3f7f7e228a8ee31fc482

Observation 3afbc6e1-ba68-4238-9002-d88ba44221df · outbound

This paper cites PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.194856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.194856Z digest=sha256:14cbcf87c6258039cef16d3c021c912bfa6ebdf13242a89b4d23c4fcad68b069

Observation 8f45b970-b8b4-4627-9524-83ad999afd2a · outbound

This paper cites VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.199614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.199614Z digest=sha256:1b52bab2a6da5ec2780606b424845b176dc0c4809d117640713165eff560b2e6

Observation 072338a1-9646-41b7-81b0-5e858610cf05 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.204077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.204077Z digest=sha256:20d6121cf53b1003ad962f590f7ef7f191e3cd2afdb0d8a4586473a347484a67

Observation 06e9e869-3a66-4c50-8bca-eb4abcc7a6d2 · outbound

This paper cites Qwen2 Technical Report.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? Qwen2 Technical Report

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.208740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.208740Z digest=sha256:00ea17ea76904cd9147b7022ee4de40df786357a52d9c9dbb95d783192c47509

Observation d4118f63-dd22-44cc-943a-d3c28a5d7eef · outbound

This paper cites InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.213153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.213153Z digest=sha256:bfe9dfc62026bee8ff836771c5c7299835b3bfe57f59898fdb2762ba738bd448

Observation 5b51070d-d63d-4fc0-976a-77c714bd43a5 · outbound

This paper cites Long-CLIP: Unlocking the Long-Text Capability of CLIP.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.217443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.217443Z digest=sha256:ea9680ee2d75ba7393334c81a7dc1ad1a3c5b0f5a252a8297f64a02d14b2bf85

Observation c5d1cee3-db0c-4fd8-b453-1aa778b5ed74 · outbound

This paper cites LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.222155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.222155Z digest=sha256:8d5f50bc8ff6b261330d9f51c6fc93b306c92e57621c7397c28107ada9a4e014

Observation 78137972-cb66-4640-99b0-572721a6272b · outbound

This paper cites , y, y, y,.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? , y, y, y,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:06:35.901230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:06:35.226952Z digest=sha256:ef5cefed5eabbb13155aa10ab232dc8b8c6ed8064d96723040e99748668ff5df

Observation fb0685fe-619e-4eb5-9965-c775bdf2a6bd · outbound

This paper cites Extending these models to video requires handling temporal dependencies (Xu et al., 2021; Lei et al., 2021; Bertasius et al., 2021; Huang et al.,.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? Extending these models to video requires handling temporal dependencies (Xu et al., 2021; Lei et al., 2021; Bertasius et al., 2021; Huang et al.,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:06:35.872041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:06:35.236494Z digest=sha256:da5a3f40178784c7d92509314a9695c4ddf56201ef82cfec09a7ab6df9e90238

Observation 74a37008-fa0c-4e64-a6d0-9b890e4bda00 · outbound

This paper cites Early video LLMs, such as Wang et al.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? Early video LLMs, such as Wang et al

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:06:35.858338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:06:35.241583Z digest=sha256:30337cd1eab158b44829b9d8e22c53363e9d375ab5d2fd0951a76c9706185728

Observation d6d43b69-749f-471f-88cc-e8ecd5dc073f · outbound

This paper cites However, video haystack retrieval still lags in difficulty compared with the QA or retrieval task in NLP (Hsieh et al., 2024; Yuan et al., 2024).

VideoRoPE: What Makes for Good Video Rotary Position Embedding? However, video haystack retrieval still lags in difficulty compared with the QA or retrieval task in NLP (Hsieh et al., 2024; Yuan et al., 2024)

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:06:35.844073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:06:35.246387Z digest=sha256:9ed1b803509b9e4b418d20102821e93839756f53a0bdefb8230d28086e9f44d5

Observation 208dd9a1-187b-4412-85a0-d387598f8d00 · outbound

This paper cites an unresolved cited work.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:06:35.830372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:06:35.250707Z digest=sha256:5753c8d2de90e468ac633ee8c244967c84091100a3a3bb945ef44342e37ae8e5

Observation d8f62f04-0191-4f40-b2c4-170ecfbe2b63 · outbound

This paper cites On the other hand, VideoRoPE employs low-frequency temporal modeling, allowing it to capture long-range dependencies and successfully identify the needle for accurate responses.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? On the other hand, VideoRoPE employs low-frequency temporal modeling, allowing it to capture long-range dependencies and successfully identify the needle for accurate responses

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:06:35.717033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:06:35.255089Z digest=sha256:3b500ea823ed323f0bcd16ef013c9dc855d8cd5496afbfc0d12d0206cea4a73d

Observation 42257403-172b-49d9-ab51-fb888b0324de · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.112479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.112479Z digest=sha256:46436aa55717ce6cc6e7528f035d8be114240ab289f1e1f2db45b56e94e88f5b

Observation c564a8a0-6674-415b-bf39-4b706af6b661 · outbound

This paper cites Is Space-Time Attention All You Need for Video Understanding?.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? Is Space-Time Attention All You Need for Video Understanding?

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.083369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.083369Z digest=sha256:9d13c6e0e50dd128d5e4d25986baa711a519659b721b6ba201547e60536ccd20

Observation dc9bac64-023a-4caa-b918-1bec63136a6a · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.190284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.190284Z digest=sha256:01eee4abb038b2215298ccab9d16751ccc1db2320905e26c9a484d8ab2cd8276

Observation c48653a7-17e7-4649-b1f8-6273adf43fdc · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.097848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.097848Z digest=sha256:ace3e0b57c287d6b3032f6469b25b3f5b927ff85a09e06c341cafa473dfc5c11

Observation 5beaa9dd-d781-4b65-9da0-1da05671958c · outbound

This paper cites Pixtral 12B.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? Pixtral 12B

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.077993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.077993Z digest=sha256:6b52bdf0df60463fc8cb79e5c0a004dc06c9a4bf6d7079830cba571f24353640

Observation bcb314c8-8f4f-48c6-94cf-730a25f6fc0d · outbound

This paper cites LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.102689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.102689Z digest=sha256:12941493d462c4e4082c0cdabe7ed484a9a3ed4a4338e784a35971dcf554adc8

Pith citing papers

Observation 0c8f2209-5d3c-42fd-ab0c-1e4749873c54 · inbound

ClinKD: Cross-Modal Clinical Knowledge Distiller For Multi-Task Medical Images cites this paper.

ClinKD: Cross-Modal Clinical Knowledge Distiller For Multi-Task Medical Images VideoRoPE: What Makes for Good Video Rotary Position Embedding?

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T17:26:28.859252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:26:28.859252Z digest=sha256:aa465401c3662bf117a3c7fdbebff8f3f5b5083cde038f50a347f8d6c35b4c14

Observation 3d0f781a-f4a6-4cfc-91ce-56a209cd6d6a · inbound

Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation cites this paper.

Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation VideoRoPE: What Makes for Good Video Rotary Position Embedding?

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-17T07:24:04.687193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T07:24:04.460276Z digest=sha256:057b62088f3247ce4d6fac8c7ea750d997cdac350fc291b233caa9bada515f6d

Observation a78a7883-6d7d-4472-902f-582dae61d412 · inbound

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models cites this paper.

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models VideoRoPE: What Makes for Good Video Rotary Position Embedding?

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:21:39.810267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T14:19:34.622854Z digest=sha256:ec4527b090ef299874733266eb3758192b389c608e99697024ed6381f06a6ba1

Observation 72d0da19-90bd-49a7-bad9-1dce2fcd7bc2 · inbound

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization cites this paper.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization VideoRoPE: What Makes for Good Video Rotary Position Embedding?

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:03.315947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:03.315947Z digest=sha256:5d855d96d9c9eb8c81c3d9ab7d6fa7735af1a6756a145db9c03343708a2aa068

Observation 4a909771-f00d-4697-acda-ccf2c2a7a875 · inbound

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification cites this paper.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification VideoRoPE: What Makes for Good Video Rotary Position Embedding?

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:11.715462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:11.715462Z digest=sha256:271612d876dc79002eff15a29555c067f6e623f354896cdd216778cf05ba4a57

Observation 333a114b-7802-40da-aee7-bd656b079e76 · inbound

Task-Aware KV Compression For Cost-Effective Long Video Understanding cites this paper.

Task-Aware KV Compression For Cost-Effective Long Video Understanding VideoRoPE: What Makes for Good Video Rotary Position Embedding?

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:38.968114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:38.968114Z digest=sha256:2a683df69e124eb096e2665052210583113c460038266edea222b2e62867caf6

Observation f519b7f8-72c3-4347-acda-48ca2a008599 · inbound

Infinite Video Understanding cites this paper.

Infinite Video Understanding VideoRoPE: What Makes for Good Video Rotary Position Embedding?

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.257348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.257348Z digest=sha256:5912eddcc1ed36d208a12a95830c837269b674d5d23846daf85da5a9bcdb48f8

Observation 403e88a9-0bf6-44aa-8100-3224b1e886ed · inbound

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture cites this paper.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture VideoRoPE: What Makes for Good Video Rotary Position Embedding?

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T11:40:23.835237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:40:23.835237Z digest=sha256:05eb0d4110a87e313ec748a4295f6d4b89ffe4bcc23d56709517118b32741609

Observation 1bccfc75-5faa-4730-9dad-5d0766c30932 · inbound

Phase-Aligned RoPE for Mixed-Resolution Diffusion Transformer cites this paper.

Phase-Aligned RoPE for Mixed-Resolution Diffusion Transformer VideoRoPE: What Makes for Good Video Rotary Position Embedding?

Reference 115

Resolution
unresolved
no resolver link, observed 2026-08-03T20:30:38.274053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:30:38.274053Z digest=sha256:6962c8276323b6910fd19b8e7513a1616438e714afaab59307da870cd1ee96e1

Observation 267f2f9f-3cd4-4d5e-b8d4-ef00846bdc9d · inbound

Adaptive 3D-RoPE: Physics-Aligned Rotary Positional Encoding for Wireless Foundation Models cites this paper.

Adaptive 3D-RoPE: Physics-Aligned Rotary Positional Encoding for Wireless Foundation Models VideoRoPE: What Makes for Good Video Rotary Position Embedding?

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:06:07.573803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-09T18:48:03.437788Z digest=sha256:56e30f56bcfcace321ee9b9f1d483aff2063ef7d8d8395d8310a731975fd0dd6

Observation 12933e0b-01a1-4e67-bb01-f81fd46e42b4 · inbound

Diffusing in the Right Space: A Systematic Study of Latent Diffusability cites this paper.

Diffusing in the Right Space: A Systematic Study of Latent Diffusability VideoRoPE: What Makes for Good Video Rotary Position Embedding?

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:36:27.600377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T10:44:24.318786Z digest=sha256:1eef8e5f918e358f7e6facb7ad749e28e4ba2001274b5901564ded94c41418a1

Observation 2a713525-41ae-4314-9492-796073ef8e62 · inbound

LiteVSR: Lightweight Adaptation of Frozen Diffusion Transformers for Video Super-Resolution cites this paper.

LiteVSR: Lightweight Adaptation of Frozen Diffusion Transformers for Video Super-Resolution VideoRoPE: What Makes for Good Video Rotary Position Embedding?

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:17:29.629230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T17:16:44.500338Z digest=sha256:b2684e6cf5e2bc03b33bdc60648d0bd4bc2605f41feee1a7e2c1abd5a3a0c5dc

Observation afdacdee-a5a5-485e-8d91-d3428a46e6e3 · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning VideoRoPE: What Makes for Good Video Rotary Position Embedding?

Reference 130

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:39:37.599608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:e15ed9c3e481cae4ba727d0fa18aa8d8522191692277b5d7d41e72ff00faf612

Observation 5f6635f3-64b7-44be-b639-c611f485d09c · inbound

ShotPlan: Cinematic Video Generation with Learnable Planning Token cites this paper.

ShotPlan: Cinematic Video Generation with Learnable Planning Token VideoRoPE: What Makes for Good Video Rotary Position Embedding?

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T17:22:32.782701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:22:32.782701Z digest=sha256:3baa789437c130c449562dbd6f1565ec0448ca96c452410a9e53e2bae7f59da8

Observation ac47fea1-aa70-403c-8d35-88a576e228ee · inbound

ChronoStitch: Training-Free Composition of Visual KV Memories for Long-Horizon Temporal Reasoning cites this paper.

ChronoStitch: Training-Free Composition of Visual KV Memories for Long-Horizon Temporal Reasoning VideoRoPE: What Makes for Good Video Rotary Position Embedding?

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T12:28:00.977584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:28:00.977584Z digest=sha256:6a8ecef1bec7b8edbf73cae69174acb12e2683182bdcf69b9deb3fd23c640cd1

Observation 387af61f-359b-4b3a-8569-d5f16007ec61 · inbound

RealVDeblur: One-Step Diffusion for Generalizable Real-World Video Deblurring cites this paper.

RealVDeblur: One-Step Diffusion for Generalizable Real-World Video Deblurring VideoRoPE: What Makes for Good Video Rotary Position Embedding?

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T09:52:44.291024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T09:52:44.291024Z digest=sha256:dbd9ff698d5044ace74ab33b4d3d3bc4979b1709e05e542070189fda5595ff7c