Pith. sign in

Paper Citation Record · LEDGER

VideoRoPE: What Makes for Good Video Rotary Position Embedding?

As of 10 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 17 inbound Pith citation observations for arXiv:2502.05173.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.05173 v3

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T20:06:35.255089Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T05:06:51.998314Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:39:37.598350Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact1
  • verified fuzzy7
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b5837546-afe1-4b76-bd4d-509e4deaf5f0 · outbound

This paper cites AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.088592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.088592Z digest=sha256:7ab4d5fe938098166d521302a0febb5850f946d1f558d7884cf67c25165547d2

Observation a2077bea-ac2a-447b-a03e-e843b71bea30 · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.093340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.093340Z digest=sha256:ffccea8f680a39f196d34e1918cca76e5415584b9bf74062d9efb460e759d378

Observation 07a9bb7b-6e13-4cfc-b409-9f359f324945 · outbound

This paper cites The Llama 3 Herd of Models.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? The Llama 3 Herd of Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.107957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.107957Z digest=sha256:d779dc84ecfe0f34639dcac6918f19861b943c75f7e260b8f55270beaede0897

Observation 7a484ea4-8fc3-41f7-9399-a6d86d273de6 · outbound

This paper cites VideoHallucer classifies hallucinations into two primary types: intrinsic and extrinsic.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? VideoHallucer classifies hallucinations into two primary types: intrinsic and extrinsic

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:06:35.887283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T20:06:35.231736Z digest=sha256:8c07749eb0cf143e6e4aaed4dceec474b095095a642c5c1abdce98fe7f8e622c

Observation d5a95be9-768f-49d8-8e48-89b5465b8e14 · outbound

This paper cites TC-LLaVA: Rethinking the Transfer from Image to Video Understanding with Temporal Considerations.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? TC-LLaVA: Rethinking the Transfer from Image to Video Understanding with Temporal Considerations

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-08T20:06:35.593964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T20:06:35.117016Z digest=sha256:93b7e82458c3e97b2a3618b4620905e792aace37196069112ed6901ff79592a6

Observation 5b65fabf-b039-4a63-85c7-3f1263918e07 · outbound

This paper cites Accessed: 2025-01-12.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? Accessed: 2025-01-12

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:06:35.914386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T20:06:35.121719Z digest=sha256:5c69e2251a607d3c9a1704daf597cc35dfed25a75074e2357c1834d3af20a92c

Observation 9e69fc2f-6a95-4bfe-b89b-12d3a77ec5cb · outbound

This paper cites RULER: What's the Real Context Size of Your Long-Context Language Models?.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? RULER: What's the Real Context Size of Your Long-Context Language Models?

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.126322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.126322Z digest=sha256:398be58ccae76208f1cde3d3d4d1b32dbe3c4013d6a03faa5cd0ef11282dfdc0

Observation 76ce522a-51f9-4e33-8d59-0eba8a33ac6e · outbound

This paper cites VTimeLLM: Empower LLM to Grasp Video Moments.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? VTimeLLM: Empower LLM to Grasp Video Moments

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.130897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.130897Z digest=sha256:ebe6e2b1b9bb68abdde3f77cb22b0c8105c1c914f6745c49ed1c2d46aab3f6ea

Observation 0ef1c0de-4222-4e57-86c7-2db6cc064039 · outbound

This paper cites OPERA: Alleviating Hallucination in Multi-Modal Large Language Models via Over-Trust Penalty and Retrospection-Allocation.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? OPERA: Alleviating Hallucination in Multi-Modal Large Language Models via Over-Trust Penalty and Retrospection-Allocation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.135448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.135448Z digest=sha256:49b80b6cefac2f691b0fb6d43fb8513f1330b483412ea1d48502ab6fcc44607a

Observation 23cc5c99-1f77-4c8e-bf47-d466a50145bb · outbound

This paper cites The Kinetics Human Action Video Dataset.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? The Kinetics Human Action Video Dataset

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.140418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.140418Z digest=sha256:5dd793c162217da9977ddb9e244ec09d34eca646e7edd7a95a0072be4815564f

Observation b932cf2e-e01e-46e7-920c-29be7c7953f8 · outbound

This paper cites Less is More: ClipBERT for Video-and-Language Learning via Sparse Sampling.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? Less is More: ClipBERT for Video-and-Language Learning via Sparse Sampling

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.144695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.144695Z digest=sha256:51fce5ac39a4e4cc09ba6e941b084e0a0bb8c1d5e870bd0c3cfc72749d372c3b

Observation 69c19d6d-75bc-4eab-a81e-bf1ba7233419 · outbound

This paper cites Temporal Reasoning Transfer from Text to Video.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? Temporal Reasoning Transfer from Text to Video

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.149060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.149060Z digest=sha256:bcf825682677f31045cc5d4118415baba8cc16332623976f6dbdce1346ff9b60

Observation ff4f5f3d-1bd8-4344-9e74-500a973541fc · outbound

This paper cites OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.153650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.153650Z digest=sha256:a0908da158d69a61b48c11a86124ba8ac6030f33d668aa182c54483c3c2921cd

Observation 0e1d8574-032b-4117-ab75-eb7050bd2c38 · outbound

This paper cites Valley: Video Assistant with Large Language model Enhanced abilitY.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.158259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.158259Z digest=sha256:55898a1d42a3e9374399be64bfe9c57aef152613bb386c0e3404a272d2bdbe97

Observation f387f90c-54ba-4c63-ba1a-e9d1f17789f5 · outbound

This paper cites YaRN: Efficient Context Window Extension of Large Language Models.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? YaRN: Efficient Context Window Extension of Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.162915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.162915Z digest=sha256:82c019c08109e29a0570cfcc492b2e69f66d83b9076fdd7cc875a52984f622f9

Observation 8a52dff9-a257-442d-8a57-e889f9000459 · outbound

This paper cites Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.167362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.167362Z digest=sha256:394e68a4b94dd9ee8f97be8bd2f92831a745535aece29b3c615e51db607355c4

Observation 84713aff-4810-4419-91c5-4a46ad3647e9 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? Learning Transferable Visual Models From Natural Language Supervision

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.172033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.172033Z digest=sha256:759294d9266cb712220db22626acc3e5cb4cc775b0e2554710f94ddc66f85d67

Observation 17f81594-282b-48cc-9b24-50ae266d3f7a · outbound

This paper cites doi: 10.1007/s11633-024-1502-8.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? doi: 10.1007/s11633-024-1502-8

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.176439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.176439Z digest=sha256:5eccb185a72231bca158c825b5ca7e860827e7e22c457dcc842b98bbc5a36670

Observation b43b4771-904a-462c-b813-dbae6baa87fc · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? Gemma: Open Models Based on Gemini Research and Technology

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.180843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.180843Z digest=sha256:2740765995e3c04d6a2173cb7deb356b9343c5310a27c6e4b0af26648abd020a

Observation ec2a373d-f810-4828-a02f-1d37a7f2e28d · outbound

This paper cites ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.185621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.185621Z digest=sha256:39018be5f2efab7f7c21c0ca422b1e5beb92e5dc5d57cb18f37b1e62214f9fce

Observation 3afbc6e1-ba68-4238-9002-d88ba44221df · outbound

This paper cites PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.194856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.194856Z digest=sha256:3b1a7f90eadbf3a8b92ae8767c484e41140265144ca0450625a6126b8486808d

Observation 8f45b970-b8b4-4627-9524-83ad999afd2a · outbound

This paper cites VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.199614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.199614Z digest=sha256:e08e86495e1defa499568f544745e3915dc972978e2b14c39c994a93c180bd30

Observation 072338a1-9646-41b7-81b0-5e858610cf05 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.204077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.204077Z digest=sha256:1737d83c93aaec0ed9bbaef25d02046f6616c0db7d3375bf61e116ecc643fff9

Observation 06e9e869-3a66-4c50-8bca-eb4abcc7a6d2 · outbound

This paper cites Qwen2 Technical Report.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? Qwen2 Technical Report

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.208740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.208740Z digest=sha256:c33ee30a9dffa73e5b8c0029d4ffa6c70fc107b9b8fc8ac5be22caaa7bbf631f

Observation d4118f63-dd22-44cc-943a-d3c28a5d7eef · outbound

This paper cites InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.213153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.213153Z digest=sha256:4055d814cb7b8f1e058bfd13b7d364fc3c4e0e2f99292fdcca74e7d91b48e516

Observation 5b51070d-d63d-4fc0-976a-77c714bd43a5 · outbound

This paper cites Long-CLIP: Unlocking the Long-Text Capability of CLIP.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.217443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.217443Z digest=sha256:1fc9fc44501d643a7eba8745a42f281e6383689cc0a6c1984171390e9fb00053

Observation c5d1cee3-db0c-4fd8-b453-1aa778b5ed74 · outbound

This paper cites LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.222155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.222155Z digest=sha256:c6877caee246415acff4fa25fcecad3793e1866c412c58b28a406c91838dcec0

Observation 78137972-cb66-4640-99b0-572721a6272b · outbound

This paper cites , y, y, y,.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? , y, y, y,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:06:35.901230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T20:06:35.226952Z digest=sha256:af2bb057a59607ea9f3725c6feda765b8e39e4e40417699315b62379c1a39c7b

Observation fb0685fe-619e-4eb5-9965-c775bdf2a6bd · outbound

This paper cites Extending these models to video requires handling temporal dependencies (Xu et al., 2021; Lei et al., 2021; Bertasius et al., 2021; Huang et al.,.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? Extending these models to video requires handling temporal dependencies (Xu et al., 2021; Lei et al., 2021; Bertasius et al., 2021; Huang et al.,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:06:35.872041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T20:06:35.236494Z digest=sha256:912c20d6b3a53ef0ff613f8569deb242e53bfaa33f56a59246a53350a51d9f71

Observation 74a37008-fa0c-4e64-a6d0-9b890e4bda00 · outbound

This paper cites Early video LLMs, such as Wang et al.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? Early video LLMs, such as Wang et al

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:06:35.858338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T20:06:35.241583Z digest=sha256:67c36e15df82106b34100c256e9d9896b317cf0dbf2e31c15561e4276140a9f0

Observation d6d43b69-749f-471f-88cc-e8ecd5dc073f · outbound

This paper cites However, video haystack retrieval still lags in difficulty compared with the QA or retrieval task in NLP (Hsieh et al., 2024; Yuan et al., 2024).

VideoRoPE: What Makes for Good Video Rotary Position Embedding? However, video haystack retrieval still lags in difficulty compared with the QA or retrieval task in NLP (Hsieh et al., 2024; Yuan et al., 2024)

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:06:35.844073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T20:06:35.246387Z digest=sha256:2273e779b073deeaf7ade2dc30581bff743157945dac68cfb01cd0d920e64eb3

Observation 208dd9a1-187b-4412-85a0-d387598f8d00 · outbound

This paper cites an unresolved cited work.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:06:35.830372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T20:06:35.250707Z digest=sha256:738e7abd68107abcc2fb5bdcce11ed87b04b23ecf9df7109dec0985138bef7ae

Observation d8f62f04-0191-4f40-b2c4-170ecfbe2b63 · outbound

This paper cites On the other hand, VideoRoPE employs low-frequency temporal modeling, allowing it to capture long-range dependencies and successfully identify the needle for accurate responses.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? On the other hand, VideoRoPE employs low-frequency temporal modeling, allowing it to capture long-range dependencies and successfully identify the needle for accurate responses

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:06:35.717033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T20:06:35.255089Z digest=sha256:4760cd64ae55aa1dc9b6ea87aba738d5af828adb3a23a45758dc8d3526daad3e

Observation 42257403-172b-49d9-ab51-fb888b0324de · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.112479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.112479Z digest=sha256:a02d00278d5f0eeb69727f207bc2c12798232d94a7df2721a4ac635a87f8660c

Observation c564a8a0-6674-415b-bf39-4b706af6b661 · outbound

This paper cites Is Space-Time Attention All You Need for Video Understanding?.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? Is Space-Time Attention All You Need for Video Understanding?

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.083369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.083369Z digest=sha256:3f1c84f8a4e236362ba74855d526df32a8fa70cfeaade151b8927137ed60ed58

Observation dc9bac64-023a-4caa-b918-1bec63136a6a · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.190284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.190284Z digest=sha256:13a655fb4d8d700efcca5af361a6efeecf093f974441236419efd6ed56836f92

Observation c48653a7-17e7-4649-b1f8-6273adf43fdc · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.097848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.097848Z digest=sha256:e3abce7e5143278addb790c40e2eef36602619a9dcb9101e804cd0a107bad36c

Observation 5beaa9dd-d781-4b65-9da0-1da05671958c · outbound

This paper cites Pixtral 12B.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? Pixtral 12B

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.077993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.077993Z digest=sha256:d23412d33bc4463c1ee520233f572a2138aed43bd11de9e2ea95dffa615f35d6

Observation bcb314c8-8f4f-48c6-94cf-730a25f6fc0d · outbound

This paper cites LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.102689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.102689Z digest=sha256:53a603c2a7b0b10f7b2fc8cc7c3742ea799111a982685a6da81b75630e28ade8

Pith citing papers

Observation 0c8f2209-5d3c-42fd-ab0c-1e4749873c54 · inbound

ClinKD: Cross-Modal Clinical Knowledge Distiller For Multi-Task Medical Images cites this paper.

ClinKD: Cross-Modal Clinical Knowledge Distiller For Multi-Task Medical Images VideoRoPE: What Makes for Good Video Rotary Position Embedding?

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T17:26:28.859252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:26:28.859252Z digest=sha256:8f2d5959e2561fb4e4a54f4c2dde8dda4c95adcd2f1489d0540b260fa3fc65b0

Observation 3d0f781a-f4a6-4cfc-91ce-56a209cd6d6a · inbound

Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation cites this paper.

Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation VideoRoPE: What Makes for Good Video Rotary Position Embedding?

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-17T07:24:04.687193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T07:24:04.460276Z digest=sha256:e18d549b41e5bfc61a391bc942bdb38f1398e6cfa17d6ac0dfe5e8a03f892dc0

Observation a78a7883-6d7d-4472-902f-582dae61d412 · inbound

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models cites this paper.

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models VideoRoPE: What Makes for Good Video Rotary Position Embedding?

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:21:39.810267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T14:19:34.622854Z digest=sha256:df409d4c98919dab49841a7e9f32fd4f665d4baf591fb668da0e7bdbe6ea84c3

Observation 72d0da19-90bd-49a7-bad9-1dce2fcd7bc2 · inbound

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization cites this paper.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization VideoRoPE: What Makes for Good Video Rotary Position Embedding?

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:03.315947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:03.315947Z digest=sha256:c14e01032e3e07be4eda2bfd6c2b87f8952eabcf9f4214339a3bce20e484cfc6

Observation 4a909771-f00d-4697-acda-ccf2c2a7a875 · inbound

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification cites this paper.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification VideoRoPE: What Makes for Good Video Rotary Position Embedding?

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:11.715462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:11.715462Z digest=sha256:cc101f3ce02d703bd8004f8ce095ee8b960a3a08880ad9604c678f8d24fd8205

Observation 333a114b-7802-40da-aee7-bd656b079e76 · inbound

Task-Aware KV Compression For Cost-Effective Long Video Understanding cites this paper.

Task-Aware KV Compression For Cost-Effective Long Video Understanding VideoRoPE: What Makes for Good Video Rotary Position Embedding?

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:38.968114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:38.968114Z digest=sha256:6d9331ec38df31529145f0f78df0fe1c6fb1862f1d028362b499a1ed1a975c81

Observation f519b7f8-72c3-4347-acda-48ca2a008599 · inbound

Infinite Video Understanding cites this paper.

Infinite Video Understanding VideoRoPE: What Makes for Good Video Rotary Position Embedding?

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.257348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.257348Z digest=sha256:5d906b31d5dd70949244db967ef898479ea4b45ba94b6f51585878d2be1a641b

Observation 403e88a9-0bf6-44aa-8100-3224b1e886ed · inbound

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture cites this paper.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture VideoRoPE: What Makes for Good Video Rotary Position Embedding?

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T11:40:23.835237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:40:23.835237Z digest=sha256:518b434b0b0a48b88a5014bbb19d63bb5aa9915a21c925575a318cee7956cdb2

Observation 1bccfc75-5faa-4730-9dad-5d0766c30932 · inbound

Phase-Aligned RoPE for Mixed-Resolution Diffusion Transformer cites this paper.

Phase-Aligned RoPE for Mixed-Resolution Diffusion Transformer VideoRoPE: What Makes for Good Video Rotary Position Embedding?

Reference 115

Resolution
unresolved
no resolver link, observed 2026-08-03T20:30:38.274053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:30:38.274053Z digest=sha256:98060c333b6a11c07d7585f735ff7be756a6e2e1b86f98ef5acdbd2f4e82263a

Observation 267f2f9f-3cd4-4d5e-b8d4-ef00846bdc9d · inbound

Adaptive 3D-RoPE: Physics-Aligned Rotary Positional Encoding for Wireless Foundation Models cites this paper.

Adaptive 3D-RoPE: Physics-Aligned Rotary Positional Encoding for Wireless Foundation Models VideoRoPE: What Makes for Good Video Rotary Position Embedding?

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:06:07.573803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T18:48:03.437788Z digest=sha256:b84963cba7210e0f7dcf938721e65412e36df1a22b3d962bf6f7cd9aad3f1826

Observation 12933e0b-01a1-4e67-bb01-f81fd46e42b4 · inbound

Diffusing in the Right Space: A Systematic Study of Latent Diffusability cites this paper.

Diffusing in the Right Space: A Systematic Study of Latent Diffusability VideoRoPE: What Makes for Good Video Rotary Position Embedding?

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:36:27.600377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T10:44:24.318786Z digest=sha256:6d8e40fc45046c8b1786254265985b3fdcb8a6042dd0e2588b036aaa37313fc8

Observation 2a713525-41ae-4314-9492-796073ef8e62 · inbound

LiteVSR: Lightweight Adaptation of Frozen Diffusion Transformers for Video Super-Resolution cites this paper.

LiteVSR: Lightweight Adaptation of Frozen Diffusion Transformers for Video Super-Resolution VideoRoPE: What Makes for Good Video Rotary Position Embedding?

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:17:29.629230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T17:16:44.500338Z digest=sha256:8aaa36436c4a04d4a18c1e02339094b2d2dae4ada535004b486de6f186b784e6

Observation afdacdee-a5a5-485e-8d91-d3428a46e6e3 · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning VideoRoPE: What Makes for Good Video Rotary Position Embedding?

Reference 130

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:39:37.599608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:932a37add955bd2b3cc1b287c30c00755b8a56849ae6622ab1d40b2d5df73b31

Observation 5f6635f3-64b7-44be-b639-c611f485d09c · inbound

ShotPlan: Cinematic Video Generation with Learnable Planning Token cites this paper.

ShotPlan: Cinematic Video Generation with Learnable Planning Token VideoRoPE: What Makes for Good Video Rotary Position Embedding?

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T17:22:32.782701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:22:32.782701Z digest=sha256:d54d5b63813689d2a04cce528900c43c1dccfc21656db56d3cf2e7d7903aa269

Observation ac47fea1-aa70-403c-8d35-88a576e228ee · inbound

ChronoStitch: Training-Free Composition of Visual KV Memories for Long-Horizon Temporal Reasoning cites this paper.

ChronoStitch: Training-Free Composition of Visual KV Memories for Long-Horizon Temporal Reasoning VideoRoPE: What Makes for Good Video Rotary Position Embedding?

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T12:28:00.977584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:28:00.977584Z digest=sha256:c396ac1746705ad0cf597afe7a826043e443426234347274d5a693f023b080a7

Observation 387af61f-359b-4b3a-8569-d5f16007ec61 · inbound

RealVDeblur: One-Step Diffusion for Generalizable Real-World Video Deblurring cites this paper.

RealVDeblur: One-Step Diffusion for Generalizable Real-World Video Deblurring VideoRoPE: What Makes for Good Video Rotary Position Embedding?

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T09:52:44.291024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T09:52:44.291024Z digest=sha256:bb321e1a943ff9e5a75b4d4d865e6d38b93d7d343846dbcff54d670ddca908bb

Observation bf6b92ab-aa87-4a59-8ff8-96e0a87c7b1b · inbound

Addressable Memory for Video World Models cites this paper.

Addressable Memory for Video World Models VideoRoPE: What Makes for Good Video Rotary Position Embedding?

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-10T05:06:51.998314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:06:51.998314Z digest=sha256:d0e33e34a639c63a070a626dd9a0bc9bcdb1f61d97019bbaaf0c530d35e26bbe