Pith. sign in

Paper Citation Record · LEDGER

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding?

As of 9 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 8 inbound Pith citation observations for arXiv:2505.14321.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.14321 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:41:09.637593Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:24:55.001360Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T00:07:28.700018Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact1
  • verified fuzzy10
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4c33aa65-0e2e-42db-9161-a9511e63d799 · outbound

This paper cites Longvideobench: A benchmark for long-context interleaved video-language understanding, 2024.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? Longvideobench: A benchmark for long-context interleaved video-language understanding, 2024

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:06.956393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:06.956393Z digest=sha256:146d474948739add6e4671b67451176406417e0fad2e82172bf96565c7d75167

Observation 50720574-c593-4f33-843e-abb42ec207a6 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long-form video language understanding.NeurIPS, 2024.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? Egoschema: A diagnostic benchmark for very long-form video language understanding.NeurIPS, 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:41:12.356042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:41:07.010477Z digest=sha256:d0d81c2e7bbee4e272c8124ca3a898d18fc076ccd72d16e53e48c3a07825fcf1

Observation e5c37511-b97e-44fe-8da5-08490f695976 · outbound

This paper cites NExT-QA: Next phase of question-answering to explaining temporal actions.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? NExT-QA: Next phase of question-answering to explaining temporal actions

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:41:12.088606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:41:07.083589Z digest=sha256:96bdfae176e336b51b8574ad0c423b253bd90104120293c8e398b59a9bafe5d3

Observation d333cd39-fecf-4ff0-bd47-6f21b014e69d · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:07.210366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:07.210366Z digest=sha256:2e815d2d770210dcc6fb1b163dd5778e09e431365653a139edf93cba3469bbb8

Observation 37c1e67a-f6d3-49aa-bdf7-7ade136da3ed · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? MLVU: Benchmarking Multi-task Long Video Understanding

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:07.306705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:07.306705Z digest=sha256:80ed452bfd14a3ed06b9be336907a63c81c4cd8b7c872ef222ea7973502501d0

Observation 1747ad97-fddd-4dfb-b21d-bbf26f516b93 · outbound

This paper cites LVBench: An Extreme Long Video Understanding Benchmark.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? LVBench: An Extreme Long Video Understanding Benchmark

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:07.407582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:07.407582Z digest=sha256:356eaa4ee2dddbf0774863d47326ad3e4cda1f288607c671a1eb0b9271e0ac35

Observation f3f18659-5bfb-4b0e-a23b-f1fb93941091 · outbound

This paper cites Perception test: A diagnostic benchmark for multimodal video models.Advances in Neural Information Processing Systems, 36:42748– 42761, 2023.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? Perception test: A diagnostic benchmark for multimodal video models.Advances in Neural Information Processing Systems, 36:42748– 42761, 2023

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:41:11.817503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:41:07.469778Z digest=sha256:f69eba8545caba1cbc2a6a4789dc421165d3830336d4550b4deb35eed943a9f2

Observation ff29bc6e-46c5-44b8-b2d6-8873122790c4 · outbound

This paper cites Palm: Scaling language modeling with pathways.JMLR, 2023.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? Palm: Scaling language modeling with pathways.JMLR, 2023

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:41:11.568082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:41:07.584674Z digest=sha256:f25ee7387105976c6a519dac3583744ec31b26b46e4ce68d4733ad704b3dc521

Observation eedf7de9-9ee6-4b19-a0b1-7a1e74ef8153 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:07.658408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:07.658408Z digest=sha256:d82f54829d456810f1feeb72491d9d27f07785eed905c3b252d2113df511315a

Observation 9860e96a-faf4-4a6f-b7e1-5c126b77a85f · outbound

This paper cites GPT-4 Technical Report.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? GPT-4 Technical Report

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:07.716508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:07.716508Z digest=sha256:8615f9e022ea9536645371b459440b4891cc66f3c9d36fea22205dc3164849b6

Observation 05209c6a-ebe1-463a-a8dd-5a27d1eb520d · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? Improved Baselines with Visual Instruction Tuning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:07.780818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:07.780818Z digest=sha256:dbc30052af34750e0ab7eb43654b07974d2cb49177dd248694bb412aa882334c

Observation 1f978819-19ad-46ba-927d-11dabe00f256 · outbound

This paper cites LLaV A- NeXT: Improved reasoning, ocr, and world knowledge, 2024.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? LLaV A- NeXT: Improved reasoning, ocr, and world knowledge, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:41:11.296636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:41:07.876391Z digest=sha256:1b876bd1ecc058e8d95d57c9504fcaf3e177c529efb04029e5b57fc158d3c77f

Observation 1db472a9-81e2-4a50-8841-9cefcd3006ee · outbound

This paper cites Fewer Tokens and Fewer Videos: Extending Video Understanding Abilities in Large Vision-Language Models.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? Fewer Tokens and Fewer Videos: Extending Video Understanding Abilities in Large Vision-Language Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:41:09.987240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:41:07.984130Z digest=sha256:5d807376ef5e087cff79e96773ab623afeb9e1a4fa1cfe058d86f057b17d7053

Observation d3ce082d-2d69-45ce-a096-ac5d33a01150 · outbound

This paper cites Video-ChatGPT: Towards detailed video understanding via large vision and language models.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? Video-ChatGPT: Towards detailed video understanding via large vision and language models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:41:11.091970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:41:08.081997Z digest=sha256:5e1775128e461a925756406f970b06452844910dc7ff0fb3efadb32729a67457

Observation d045ce46-38f7-49f5-a98b-3709d6d129a6 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? Learning transferable visual models from natural language supervision

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:08.147677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:08.147677Z digest=sha256:d3ce5f2910f2e2df559aaa4ccb035a561dfe81bb4b808474083b02ed6a08031a

Observation 77a759a2-488b-4a4f-812a-0eb18daddc67 · outbound

This paper cites LLaV A-NeXT: A strong zero-shot video understanding model, 2024.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? LLaV A-NeXT: A strong zero-shot video understanding model, 2024

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:41:10.817989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:41:08.205977Z digest=sha256:ee40adf33b3ef18f013b4c4d30bf8e8f6e51363d1dc1d58d63f0d78cccdb859e

Observation ae24c0ee-71d4-42d7-b8f4-2ced1bbe7b05 · outbound

This paper cites VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:08.331967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:08.331967Z digest=sha256:53e11d1cf5b2f7a9d576a930a493aacb5d6dcbc1704fa2341fac22f71026cec1

Observation 33c74869-a7ea-414b-91da-4bde70336a44 · outbound

This paper cites InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:08.446932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:08.446932Z digest=sha256:216d15575a5a09422e1ca7d3c4a665bd3f91e2d0de2de229cda55e5356eab0fc

Observation 73db5c87-2b89-4b6e-abd9-9da8866b837f · outbound

This paper cites Apollo: An Exploration of Video Understanding in Large Multimodal Models.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:08.568366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:08.568366Z digest=sha256:0468d37f774724681a0bb2505acfa4db918dced44b707d513047a1a2c2fd61d5

Observation 15403d8c-9c3f-4dc4-87cf-6aab6ed4073b · outbound

This paper cites Oryx MLLM: On-demand spatial-temporal understanding at arbitrary resolution.ICLR, 2025.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? Oryx MLLM: On-demand spatial-temporal understanding at arbitrary resolution.ICLR, 2025

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:41:10.631715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:41:08.690765Z digest=sha256:55fb387ecdff3ad4be01912d300c89277c933897763c34b51dec760b5efff2d9

Observation 1ce74f51-9a08-4df3-8d8f-bffae76d590a · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:08.771392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:08.771392Z digest=sha256:3c6b00796eeae130fdda436f7aade2c675be03b19d6cfe42575c881f6317ce5b

Observation 18674d71-7b64-49eb-a2e6-63ad3c5caffa · outbound

This paper cites Video question answering via gradually refined attention over appearance and motion.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? Video question answering via gradually refined attention over appearance and motion

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:41:10.460870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:41:08.863605Z digest=sha256:4a5060f04f33ea7dfb43722f627325f5ba2c5f6ab3eb61b3b6271466751b4058

Observation 7b72c423-7680-4d6e-a3c0-1b7b844ae2b1 · outbound

This paper cites ActivityNet-QA: A dataset for understanding complex web videos via question answering.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? ActivityNet-QA: A dataset for understanding complex web videos via question answering

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:41:10.295260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:41:08.944154Z digest=sha256:f2bb216a260b1af86ec0a87678819138c04f01ff60d43e23c8895c8bc1ceb113

Observation 8f4aaf61-ea0e-4a96-bc14-66fe79aab981 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:09.035873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:09.035873Z digest=sha256:6fce1d4640a1442b0bde9a56c8f36a52b3a4c89c4bdb95b63e9654d97b14fa62

Observation 13727959-433c-4784-a7c7-20bf3e898152 · outbound

This paper cites LIME: Less Is More for MLLM Evaluation.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? LIME: Less Is More for MLLM Evaluation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:09.118503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:09.118503Z digest=sha256:b6f77c6b3d8b8f45cd39205d33f6fe8f692fbe58291dfb0521dc2cf57d26ceb2

Observation 7a46c902-f8b5-4033-88b7-c49360ce7c9a · outbound

This paper cites SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:09.214825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:09.214825Z digest=sha256:8653f71ea01ce5b53b6a8facaf4290f0ea8f9f43b291e0413e17909416e65b61

Observation 8d3ba9d7-f312-453b-8ac2-1aca202aa967 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:09.310216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:09.310216Z digest=sha256:b964768059dd784b448a1f7dd788343e489e550489bf23eb68654a874cbefb0e

Observation f1591ffc-cd7f-4534-bac7-83080ef6bccc · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? LLaVA-OneVision: Easy Visual Task Transfer

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:09.367899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:09.367899Z digest=sha256:580d2d805a7b6ed5f07d8c1bc2b73870ada7fe8125a293aa5144db78e360cf73

Observation d0d88534-9569-4b16-a312-cc1126394219 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:09.462250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:09.462250Z digest=sha256:4726be8afe055cd260afb8fc3eaacec6794e11b90829fdce7dd96ec9f4607234

Observation cd5c3eea-f837-4aaa-9500-c77784b901eb · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:09.539358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:09.539358Z digest=sha256:e9ebc75e38a7be57ef672c7851663eaa5be2a76ee34fcef9db876b9bbb4b679b

Observation 9223dad6-8fbc-4ace-ba0a-386a2e5366f4 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:09.637593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:09.637593Z digest=sha256:7defc9bfd929536096ce159497d74a685572fed45f3fdcf0935f1c50658be363

Pith citing papers

Observation 18307915-d59b-4e62-944f-aa7c483f1624 · inbound

NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding cites this paper.

NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T13:01:57.956043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:01:57.956043Z digest=sha256:c22be1ebfac385b2b10e5b7887276a98ff1d86fa7d90893796c4ec92da8f67e4

Observation 9afbff76-9efc-4df2-9eb1-4cf3663958cc · inbound

NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding cites this paper.

NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T06:36:33.806787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:36:33.806787Z digest=sha256:3bd6c4002ecdc0eb2f50d2b6b90ef2019ac7859455b840bee01642f2c2944804

Observation 4af70d98-2788-4b20-85b2-abe90fc7600a · inbound

RefereeBench: Are Video MLLMs Ready to be Multi-Sport Referees cites this paper.

RefereeBench: Are Video MLLMs Ready to be Multi-Sport Referees Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding?

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:48:02.252494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T08:38:27.081358Z digest=sha256:f970093cf8ef53da5c8e8dfc27c3cbec97a68a9d9a25ae148fad955a6b66e699

Observation 8b721bc2-2993-4738-92b6-4fd426d49e1c · inbound

VISTA: Video Interaction Spatio-Temporal Analysis Benchmark cites this paper.

VISTA: Video Interaction Spatio-Temporal Analysis Benchmark Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding?

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:41:22.517885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T15:18:00.436343Z digest=sha256:488ec50d517a014b41f93c54141988f7e9635e449b5961ee233e287e7281bda5

Observation 537c814d-9e7f-4b8f-819c-53b57e0df066 · inbound

VISTA: Video Interaction Spatio-Temporal Analysis Benchmark cites this paper.

VISTA: Video Interaction Spatio-Temporal Analysis Benchmark Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding?

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-01T00:25:09.422359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T00:24:25.203292Z digest=sha256:3e9958805dff898881cb1c9f98021b323c395ab5f164318be486d800645dfabe

Observation efdfaa89-8f69-4236-a9c6-9184a07e4464 · inbound

An Attribute-Based Measure of Video Complexity cites this paper.

An Attribute-Based Measure of Video Complexity Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding?

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:02:34.110328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T19:00:54.718177Z digest=sha256:e17e435e856ec875c15782fd065e0dc10e741fabe172b22d488141b73439927f

Observation 52eb70fb-bb63-4e0c-b641-84a3ea0666ea · inbound

Do Video Foundation Models Understand Intuitive Physics? A Layerwise Probing Analysis cites this paper.

Do Video Foundation Models Understand Intuitive Physics? A Layerwise Probing Analysis Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding?

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T00:07:28.701935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T17:25:15.493925Z digest=sha256:fc009f0f109dbc7c1459081f4204a543299cf446992f34150b1b1cc3014587c0

Observation cd04825c-b05d-4ba5-9fad-add7deb39880 · inbound

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping cites this paper.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding?

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:55.001360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:55.001360Z digest=sha256:15763174371fca4d2c81399cd58ff717777e0eaaa8f67d125fc8cddea05ddcc2