Pith. sign in

Paper Citation Record · LEDGER

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation

As of 14 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 10 inbound Pith citation observations for arXiv:2505.14640.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.14640 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:34:44.884301Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:50:02.846860Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T19:07:17.426534Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved48
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6a88210c-b53b-4908-b4e2-4d404068d158 · outbound

This paper cites LVBench: An Extreme Long Video Understanding Benchmark.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation LVBench: An Extreme Long Video Understanding Benchmark

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:40.208786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:40.208786Z digest=sha256:ce65b2a1ca477b3a6da60ed3d1ba300f7b6f2a441ac716ab59ffa8c92bf98130

Observation 10ed969e-dfad-49c1-9133-7e0750fa635e · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:40.317680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:40.317680Z digest=sha256:ee3dea2c0cbe303d72f8d174a51882a3b4766c6d245b0c131d6b2a713bbf5fd7

Observation 1e7ad542-9a07-4b88-bba4-3cd40cdc2b28 · outbound

This paper cites Video Anomaly Detection and Explanation via Large Language Models.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Video Anomaly Detection and Explanation via Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:40.477549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:40.477549Z digest=sha256:a22c9cfe9885c1ce19b23ab14fad3b952fc7456398cbd57a8e75cd6d1d240c7d

Observation 9d5f0c6b-03c9-4506-9af8-f0464c479552 · outbound

This paper cites Large scale interactive motion forecasting for autonomous driving: The waymo open motion dataset.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Large scale interactive motion forecasting for autonomous driving: The waymo open motion dataset

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:40.586160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:40.586160Z digest=sha256:76906aecb5cb7af300ea06cbb278856fab0e094ff3ea3fcb4ddac19727aeed58

Observation e77b6f11-b337-4d0a-9a24-bfdfca92ecb7 · outbound

This paper cites Towards automatic learning of procedures from web instructional videos.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Towards automatic learning of procedures from web instructional videos

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:40.675590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:40.675590Z digest=sha256:e416987eb8d408d8a94d85ec89231f6e37d9f6e19063caf66d4ef91a900aab03

Observation fbb17e75-ffea-475a-947c-10ab364bd87b · outbound

This paper cites Long Context Transfer from Language to Vision.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Long Context Transfer from Language to Vision

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:40.769437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:40.769437Z digest=sha256:44dd517656fa61618eabe2cef217f23150a27d0ed3a936603a722d87269c771d

Observation f52197b2-3d0b-4a6e-8276-e907aa5a7935 · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:40.848837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:40.848837Z digest=sha256:1b27ada110757e1024baa192c12a69932ee38335c1548bab039f9c308b75369f

Observation 0e55e230-5270-4ea2-ba9d-10a36c347294 · outbound

This paper cites VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:40.968075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:40.968075Z digest=sha256:56e52792d8d1b576393d27870a4d1641a8f9738297178ccd955e07142e32d44c

Observation a97207d3-7a08-4b04-b7be-0cd06df75772 · outbound

This paper cites InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:41.079838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:41.079838Z digest=sha256:727abc7dbeea53637893d981ed01c3ed286616a6771c65859ae2f868d64578b9

Observation 08eef1c9-d3c6-4b8e-a355-ced9c4eaa905 · outbound

This paper cites Vamba: Understanding Hour-Long Videos with Hybrid Mamba-Transformers.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Vamba: Understanding Hour-Long Videos with Hybrid Mamba-Transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:41.182215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:41.182215Z digest=sha256:2e1d04c65f068081831672c4b8dbdc1579c86c9a097c0941531398cbbce21d2f

Observation 809ec41f-47d5-4412-a93c-ca8d346ecae4 · outbound

This paper cites Token-efficient long video understanding for multimodal llms.arXiv preprint arXiv:2503.04130, 2025.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Token-efficient long video understanding for multimodal llms.arXiv preprint arXiv:2503.04130, 2025

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:41.295575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:41.295575Z digest=sha256:50ece9684073d336582140bb2670124f4385d125ff98d6c5f5d50e83e9c41409

Observation 3ac6fee9-ec53-406f-b2b4-f7c2e83cff95 · outbound

This paper cites BIMBA: Selective-Scan Compression for Long-Range Video Question Answering.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation BIMBA: Selective-Scan Compression for Long-Range Video Question Answering

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:41.423639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:41.423639Z digest=sha256:535f239435de597351f2bb5bb930117752667c88595c56452a8a153f3c4eb652

Observation aece33d9-7349-46da-b737-66bdd2dbf49b · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:41.495843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:41.495843Z digest=sha256:9a3d64b25c349618b56b1acdffe19f4b5bcb975bf1278caf30e46d73dd369c74

Observation c716fa61-58fb-4624-89ce-fe750b5d2a98 · outbound

This paper cites VISTA: Enhancing Long-Duration and High-Resolution Video Understanding by Video Spatiotemporal Augmentation.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation VISTA: Enhancing Long-Duration and High-Resolution Video Understanding by Video Spatiotemporal Augmentation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:41.595618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:41.595618Z digest=sha256:18fb41937032d6e44b65eebda4a5c3cdd396425aaededcd3cf3419da5addeb0b

Observation 34216a65-fbda-4284-9c4c-1b6854591650 · outbound

This paper cites VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:41.669333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:41.669333Z digest=sha256:faea43e26556cc46b315a780b6d6206428181f29a4ddd159fa2abdcfd9f70105

Observation 157d2750-e5aa-4ea4-9f87-c7e1e62a4782 · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:41.743812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:41.743812Z digest=sha256:cdd1e287dda80547ca5f63f9b09fda49a5ac7ac4f0d0461d5c407d11c0c28a43

Observation 4f4b19fd-09e3-4e4d-95d7-a4aaf5e427f6 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:41.886033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:41.886033Z digest=sha256:c2164d150d7b0f3f0e75b26c5278b574b46f3419735596b4533d7fa334df3925

Observation 3e33eeb5-d633-4f8a-b2ed-99519e4ae733 · outbound

This paper cites Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:41.992087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:41.992087Z digest=sha256:b76111474fc079c9caaef08efa17a47f189f114c4acbb44bb4cb9aae021eebe2

Observation 5d257b92-d56b-4914-867c-72506d73fc7e · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation MLVU: Benchmarking Multi-task Long Video Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:42.108152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:42.108152Z digest=sha256:7c7bfa7cfec9d959069fae121442058ad266cc47cf80613229f733ac6b8fc160

Observation c0f84637-01cb-4874-a219-6fbf78ea032e · outbound

This paper cites Longvideobench: A benchmark for long- context interleaved video-language understanding.Advances in Neural Information Processing Systems, 37:28828–28857, 2024.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Longvideobench: A benchmark for long- context interleaved video-language understanding.Advances in Neural Information Processing Systems, 37:28828–28857, 2024

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:42.202827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:42.202827Z digest=sha256:dabb5409549d86b222219e83ed14722d948487b66785796f5792a79724a5446d

Observation 4f3d0a31-25bb-492d-a763-2c477de4a226 · outbound

This paper cites Hourvideo: 1-hour video- language understanding.Advances in Neural Information Processing Systems, 37:53168–53197, 2024.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Hourvideo: 1-hour video- language understanding.Advances in Neural Information Processing Systems, 37:53168–53197, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:46.268970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T15:34:42.274637Z digest=sha256:143bc3640da03471a8b7ce2050b39dd8cd6a219afff8d568c43b1323217e1b98

Observation 4cbf0ae1-1f36-4988-99b6-e62437f5ba40 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:42.341438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:42.341438Z digest=sha256:81fc9a07aced6f37a73217b7f027430f98044929fc2a3e404af129fbb37cc6c0

Observation 5f383376-799c-40e8-b0ff-10aae3869e41 · outbound

This paper cites Qwen2.5-VL Technical Report.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Qwen2.5-VL Technical Report

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:42.438272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:42.438272Z digest=sha256:643e3d750e50bf14e76cdf7d8838aaad9ad9352ba43889ff11f349a6cecbad02

Observation 3f9b5974-c0ef-464f-93fd-7d8253ab8b18 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation VideoChat: Chat-Centric Video Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:42.546538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:42.546538Z digest=sha256:c755a9aa26b1f4408b33e67abd2089a15f78f52f4cd7f908368fe63987a7093a

Observation 3fe5c5d6-7f0e-4d57-a531-0314a6b3bbd8 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:42.650236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:42.650236Z digest=sha256:495caf8e582fd2c275775bc1aeee2fcb27b8146aa74c472e08c23d2a4e733243

Observation f9db9ddf-06ec-471d-8a9e-933d2d1fbcb8 · outbound

This paper cites MANTIS: Interleaved Multi-Image Instruction Tuning.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation MANTIS: Interleaved Multi-Image Instruction Tuning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:42.725167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:42.725167Z digest=sha256:81a911ad2574a2333ea540b43dcb9ec7333fd263655dd7d057ad1e0206a7f96f

Observation cc29b80a-afa0-4295-88e0-6ee1ce4abb53 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:42.813436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:42.813436Z digest=sha256:785037c50fcc11d95435615d5b9c7c35843605bb4807ef35fb8b9ad7d292d007

Observation 911ec23a-c43c-4dbf-989f-a33908344c6f · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:42.886209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:42.886209Z digest=sha256:262d97d5bd58b1fba1ba3a6b37c370f8735fa098dd228cbc335b3fdbf0639608

Observation ce17acaf-da91-44b8-a33a-65f9f5f53e0c · outbound

This paper cites Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:42.981605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:42.981605Z digest=sha256:5196416bceb5cdd582527c863c0db1d1779942b9e2f0bd06833377a2e1e59d89

Observation b0cc6776-42c0-4274-841d-926965e00b77 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:43.096309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:43.096309Z digest=sha256:963ca4069a0b6801865298083417191006d276be9f6533efc04b11b33a1d5a4d

Observation 4cac2fae-867a-4ca3-81ec-80b4cd26c169 · outbound

This paper cites LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:43.173823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:43.173823Z digest=sha256:c199df0a8a0dada6433ceeaa56fbfa6b1311f15f98da5d951bb1bbb6f4c6a05a

Observation 5523b6de-9a63-4e83-8221-c14dd59cc6bc · outbound

This paper cites Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:43.251039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:43.251039Z digest=sha256:57bffe94044c8b8e08c7977efc2d7ba549fcaec990645ffa4b6c246e7be2d3b2

Observation a56cf553-118f-4410-91a7-10e529619f69 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:43.353762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:43.353762Z digest=sha256:255bc782287ecdf005521e436a10e88d4b4747ee62e9c64e1bdcd1c2394d9f1d

Observation 7d083842-2e43-4ed1-97ad-f4013fb07b65 · outbound

This paper cites Vript: A video is worth thousands of words.Advances in Neural Information Processing Systems, 37:57240–57261, 2024.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Vript: A video is worth thousands of words.Advances in Neural Information Processing Systems, 37:57240–57261, 2024

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:43.440428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:43.440428Z digest=sha256:60f563994b993faa312f7585a53de8d2d719359aae4e16648a4a943c750b9518

Observation 782cb605-d43c-49b8-be75-4e6046766571 · outbound

This paper cites Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:43.506440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:43.506440Z digest=sha256:71db784e4924af16555d751a53b72d4382551e640d6a1ff28ddfbb5879546473

Observation 085e1bbd-1e7e-441c-bca8-6507ba764d3f · outbound

This paper cites Video question answering via gradually refined attention over appearance and motion.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Video question answering via gradually refined attention over appearance and motion

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:43.574670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:43.574670Z digest=sha256:3ba74d9703518e43f46fd6b8cc96c457e7f0a9a2c461ba7dcca88031ed401bb4

Observation 0d2e9a99-d3cd-4e3a-bbe4-0395d87eacb7 · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:43.650831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:43.650831Z digest=sha256:0d2f5ee51906c1f8e993c323e03d84e8e346020b721c5423cf75e1955c3b509b

Observation 647e48b4-0e88-4c5d-82ce-b6601f34a514 · outbound

This paper cites TempCompass: Do Video LLMs Really Understand Videos?.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation TempCompass: Do Video LLMs Really Understand Videos?

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:43.723261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:43.723261Z digest=sha256:c455c4b1e5d6166c7287ccd5406aee45efeb6a426c03d0369dd89c290714ccfa

Observation 29aaa453-09ca-4625-84b3-d9457014c874 · outbound

This paper cites VideoVista: A Versatile Benchmark for Video Understanding and Reasoning.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:43.788047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:43.788047Z digest=sha256:b3dd22a4e3b2470b09b5ed140af548c88cbffac7f75f8e789d9811507df3a4ff

Observation cee2d8f7-c7d7-4872-8ffd-955bf72fc76f · outbound

This paper cites Measuring short-form factuality in large language models.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Measuring short-form factuality in large language models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:43.846476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:43.846476Z digest=sha256:2d709e1c261f7aaee0eeb5d18d8263e560b9e4906ac543bb18fd7d255599021e

Observation 6ab987bd-e7c2-4885-ba42-fd30ab5413d9 · outbound

This paper cites Video SimpleQA: Towards Factuality Evaluation in Large Video Language Models.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Video SimpleQA: Towards Factuality Evaluation in Large Video Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:43.911417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:43.911417Z digest=sha256:9e58aa0dea2ed725740a4e31b10c78cbda0faa9ac6a6c8815ffe78bf59db9bdb

Observation e73cd584-1302-45ee-9b4f-ad87622b26ea · outbound

This paper cites A Survey on LLM-as-a-Judge.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation A Survey on LLM-as-a-Judge

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:44.002274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:44.002274Z digest=sha256:741a33dd442008711fbdaa720c267619385165478fb17666eef6c7d8ac2f180c

Observation 490bec7a-661d-4ac6-85fa-759897b69dbd · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:44.072681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:44.072681Z digest=sha256:e8280c2eebc61db33ddc87ba6b712fe43f7f4db95bf5cb4003d6f2a5672ace7f

Observation c76730be-8774-46a2-a7c2-37b02a43f9a7 · outbound

This paper cites Gpt-4o.https://openai.com/index/hello-gpt-4o/, 2024.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Gpt-4o.https://openai.com/index/hello-gpt-4o/, 2024

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:46.047734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T15:34:44.129970Z digest=sha256:94033c47b8516107fd67f408148bc96e8af4aa3c4630118b9b1ecc3d82127da1

Observation 4d87df07-319c-4926-b206-0ca0809dfd24 · outbound

This paper cites Gpt-4o mini: Advancing cost-efficient intelligence, July 2024.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Gpt-4o mini: Advancing cost-efficient intelligence, July 2024

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:45.877992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T15:34:44.192407Z digest=sha256:6f0c3c24f0d3579b0c7c940a89a591e810dd3f3818adb3a843652fb34895837f

Observation d0f93d1e-fae4-425a-bff1-0880eecedd1f · outbound

This paper cites Introducing gpt-4.1 in the api, April 2025.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Introducing gpt-4.1 in the api, April 2025

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:45.790494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T15:34:44.250379Z digest=sha256:06364ec135af23f60c841a0800d4c2c354640a36416a867c9a37d748160c54e8

Observation f669b362-f506-4ab3-97ad-60abaef86525 · outbound

This paper cites Gemini 2.5: Our most intelligent ai model, March.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Gemini 2.5: Our most intelligent ai model, March

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:45.725060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T15:34:44.319389Z digest=sha256:a90e603622e3314ecdb4a608169e3ffb14db689e45a986480de22246e6c8726d

Observation 1bebcfb6-ddfd-44f2-a173-13b10dc86f4f · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:44.468889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:44.468889Z digest=sha256:56f3eb44298fcd24cd1a97c8c701953fdaede4f962e68ad5ca91ceebcf05361a

Observation 5d0872ad-c2e1-40e5-a047-d42e50b38f36 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:44.518133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:44.518133Z digest=sha256:7470ac99e6cc7bfa5a598c18de2add2e5163e38c632da29858304800bd024abd

Observation e2f50624-c532-4090-9107-576c2a735bad · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:44.591263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:44.591263Z digest=sha256:0224f743898c32fc8a02568ca66da5908e1f1d2f908b76c35de345b2c37428ae

Observation fb22c9ce-78de-42ef-8966-21a2949b6f12 · outbound

This paper cites Phi-4 Technical Report.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Phi-4 Technical Report

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:44.670107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:44.670107Z digest=sha256:2cac31788105f555f146f2b164e3a53fd238accd45547d4d5faca0d29c27305f

Observation 25113abc-91cd-44f7-93f6-2d069f44df5a · outbound

This paper cites Longllava: Scaling multi-modal llms to 1000 images efficiently via a hybrid architecture.arXiv preprint arXiv:2409.02889, 2024.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Longllava: Scaling multi-modal llms to 1000 images efficiently via a hybrid architecture.arXiv preprint arXiv:2409.02889, 2024

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:44.746446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:44.746446Z digest=sha256:4cc5b5ad52be36eb584b4a2b1d487e0951e9d1e54f510f5dcd80ff5c6f466edf

Observation f16b9d63-0c2a-4bd7-8b01-f283f01c91d8 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:44.818600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:44.818600Z digest=sha256:21df955862467ff30d9b172c2acea56decafde15dc8f526c9d0ad5a6d0fcb284

Observation ff54a88f-14c1-4d97-958e-95fd6b6afaba · outbound

This paper cites Keep" if the question can be answered by someone who has watched the video, even if the answer requires reasoning or summarizing visual or auditory evidence. -.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Keep" if the question can be answered by someone who has watched the video, even if the answer requires reasoning or summarizing visual or auditory evidence. -

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:45.565691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T15:34:44.884301Z digest=sha256:0f57f631e10ae46114dc743f11ad355504b6d173cd736c5cd92847525571560a

Observation 0a5fdf07-81df-48d0-b405-ba843d7a7362 · outbound

This paper cites Accessed: 2025-05-08.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Accessed: 2025-05-08

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:45.668732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T15:34:44.396406Z digest=sha256:da1ff72e5b4f32baeb2e60a5b8b8766900be87b44f6d4203b696c6459bb5df7b

Pith citing papers

Observation a0e8be73-74c8-40ed-ac7e-74a5cf6f213b · inbound

Vid-SME: Membership Inference Attacks against Large Video Understanding Models cites this paper.

Vid-SME: Membership Inference Attacks against Large Video Understanding Models VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:02.846860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:02.846860Z digest=sha256:6d3d5a83b80f497eb1b9afef8f0e45a9d5fcc75d26e38e764ba71da2d45f08bd

Observation c7fc560d-b42f-45f3-a437-4be1ade2a1d1 · inbound

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification cites this paper.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:11.304314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:11.304314Z digest=sha256:8c174bc89a52a536d29b4218ccaf72792c2752c79fc809a8aa71ae7f33a8981d

Observation 39eb555c-eb3a-4d13-b80d-f9c4fac7983e · inbound

Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding cites this paper.

Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation

Reference 98

Resolution
verified exact
arxiv_id, observed 2026-05-16T04:21:29.796098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-16T04:21:29.526008Z digest=sha256:f5fde1bc756a33a0f9a50b210ece4bfba08da721efc6e3b84ce2387925f22e94

Observation 47adbc6b-9538-4066-9e0a-a2ccd64bb11c · inbound

TrajTok: Learning Trajectory Tokens enables better Video Understanding cites this paper.

TrajTok: Learning Trajectory Tokens enables better Video Understanding VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:16:31.811999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T19:11:52.694778Z digest=sha256:7308bb9431b5af870da8d393a88c3e4c3b96c26709114ae3bc90d1e548f268b3

Observation 920533a0-f3fe-4684-a31e-1a4b0b8223d9 · inbound

TrajTok: Learning Trajectory Tokens enables better Video Understanding cites this paper.

TrajTok: Learning Trajectory Tokens enables better Video Understanding VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T20:38:40.952686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:38:40.952686Z digest=sha256:6ea68dec8afb21f85ab8f7f0856f92fe0c2a471108d9f239e8f6b8d09cb8d4c0

Observation 13235850-057d-420d-b4cf-99c338778fd2 · inbound

Video-Oasis: Rethinking Evaluation of Video Understanding cites this paper.

Video-Oasis: Rethinking Evaluation of Video Understanding VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-13T15:38:41.390945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T15:38:41.390945Z digest=sha256:2b2b8dc3eb16eead13f4a37934b1120b7e646baf72f960970b2813cfe442503a

Observation 4b5eeacb-ba77-4cfc-97f9-0dba3b9c2840 · inbound

WorldReasonBench: Human-Aligned Stress Testing of Video Generators as Future World-State Predictors cites this paper.

WorldReasonBench: Human-Aligned Stress Testing of Video Generators as Future World-State Predictors VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:11:25.138117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-12T03:39:01.254916Z digest=sha256:9ad1cbfd51fb61727ede481cf79763334edb1e12972256154e911809c53af03c

Observation 280e7169-00b3-4883-bce5-1fc4884e04b7 · inbound

Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity cites this paper.

Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:07:17.427951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-02T18:57:46.841456Z digest=sha256:a633a8b12a7178aa11aa4617a0ea603b448ed251997b8c33d63e7e48612f9dce

Observation 84ee4e48-c29d-4129-aa2e-0182b72b4e4e · inbound

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding cites this paper.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:46.835407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:46.835407Z digest=sha256:aeac3dacfc738c7fbcec0fd27064dcb2d050f62c4d9b883151fc134a00076d8b

Observation 40587f60-b205-4071-bfe4-bd9aa258a220 · inbound

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model cites this paper.

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation

Reference 155

Resolution
unresolved
no resolver link, observed 2026-07-31T06:20:14.256778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:20:14.256778Z digest=sha256:3bcd8301acc99d3d047163939af0743ee0767d17fa5c5cad9628fd37f870a985