Pith. sign in

Paper Citation Record · LEDGER

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding

As of 13 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 1 inbound Pith citation observation for arXiv:2507.15028.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.15028 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:48:38.262186Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T09:12:09.932226Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

58 of 58 outbound references displayed

  • verified exact0
  • verified fuzzy38
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5df95d99-6457-4549-98e8-bc2296fb4eaf · outbound

This paper cites Vqa: Visual question answering.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Vqa: Visual question answering

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:48.049637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:48:30.939545Z digest=sha256:fa73635ac3cbf78fe854432e8dc0f6c97fe8e54dcf73d525b9f834a78590266e

Observation f91d00e8-c6f6-491f-b1a5-7bdc970597a8 · outbound

This paper cites TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:31.035101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:31.035101Z digest=sha256:c22d0ba9763569249fcbcf0d1f12f3704093db44e5552b17aa400fa1bd03c338

Observation 455bff05-8a2e-4bca-bcb1-2426ad781cbc · outbound

This paper cites Collecting highly paral- lel data for paraphrase evaluation.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Collecting highly paral- lel data for paraphrase evaluation

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:47.771747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:48:31.146121Z digest=sha256:d0a3db6eeff1ccbe92d011fd1bb2df89dd70815b97b40290abdd19435668b75e

Observation 5d1d99ad-54ad-465f-83df-4725f87d8d78 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:31.309034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:31.309034Z digest=sha256:ae99c8b3f038ec31d24eb2105b80325be5c3ac3cf9d25d6a380f7263411bdbc7

Observation 7811dae3-8c3b-4a69-ac98-5634d2a7ae5b · outbound

This paper cites Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:31.423977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:31.423977Z digest=sha256:5476026ad252bad4c664e187aaaf82db1b4ce82e0c39b74bf965f5c29c7340f1

Observation 419a10bf-2fc6-4e6e-a497-af875e6cbe22 · outbound

This paper cites MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:31.569319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:31.569319Z digest=sha256:9a9cd49879a35ba5d3a2b8f9d5c07674177e6528fad092d4d987c4fe6c8ae123

Observation 20cae88d-8699-4879-bab5-a606e491c550 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:31.692811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:31.692811Z digest=sha256:ca027f04225231299d99173da1782f1f0eeeb50d8f77c1b253f07ccf4a04eb26

Observation 6601ce9b-8936-4466-a2e9-2e1660cd6698 · outbound

This paper cites Agqa: A benchmark for compositional spatio-temporal reasoning.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Agqa: A benchmark for compositional spatio-temporal reasoning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:47.570794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:48:31.798160Z digest=sha256:4d1fe8dc94a3c9c1e3ef19e5ecaa1f2364abf122a35599db24167ad3fa15cd79

Observation c5bb15e9-d108-40c0-95af-fc0fd2632a22 · outbound

This paper cites Similarity and fea- tures of natural textures.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Similarity and fea- tures of natural textures

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:47.327426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:48:31.967725Z digest=sha256:df84134217c1b89eb433dd6d971950e4044ab50c3f8f6dfd6751a2230824fd52

Observation fef48a7e-f210-40f0-a400-b5f99936a1a2 · outbound

This paper cites Natural adversarial examples.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Natural adversarial examples

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:47.012478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:48:32.140486Z digest=sha256:038cb1bb91dd1d520b5e29e4eb89c48277378d5f96dd8f6f08f03a238be8735c

Observation db349333-1200-42ed-9571-7d04be0ba812 · outbound

This paper cites Video-mmmu: Evaluating knowledge acquisition from multi-discipline pro- fessional videos, 2025.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Video-mmmu: Evaluating knowledge acquisition from multi-discipline pro- fessional videos, 2025

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:46.740041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:48:32.242853Z digest=sha256:e23af2c0b1a556054bfe5dc6385b95e4a457f1ec1fbd8361977fb73d836d5e7b

Observation 09e9c5cb-554a-4e44-a014-9f9206135d35 · outbound

This paper cites Tgif-qa: Toward spatio-temporal reasoning in visual question answering.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Tgif-qa: Toward spatio-temporal reasoning in visual question answering

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:32.351024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:32.351024Z digest=sha256:a17cb1701fc0114492a3dbed3fe023832b1fc22eadc2b12eea97845ee987c2ae

Observation ddf0b73b-170f-460c-a1b1-b11ec693ce2a · outbound

This paper cites Robust modeling in cognitive science.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Robust modeling in cognitive science

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:46.500872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:48:32.475759Z digest=sha256:b1403f71f3fa5414bb6107f5bfa52d4dc7f6154856df64560957837e30d7c6eb

Observation 3a2e6533-99d6-4ee3-b0d6-3e5259281859 · outbound

This paper cites TVQA: Localized, Compositional Video Question Answering.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding TVQA: Localized, Compositional Video Question Answering

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:32.606188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:32.606188Z digest=sha256:15f0b563ab627c2d057e7e305c038c5f4168e8ddd74d8b0e1e490c8615a489a0

Observation 1ac2132a-7884-43c0-b223-c3ae4ba09978 · outbound

This paper cites Mvbench: A comprehensive multi- modal video understanding benchmark, 2023.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Mvbench: A comprehensive multi- modal video understanding benchmark, 2023

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:46.290809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:48:32.768765Z digest=sha256:00a0a38267525943d04e580544ff9dbf91fbc6d0eca166919c315fe1713d6ca3

Observation a15c7e92-a2ef-4d72-87a3-cf47fee2a286 · outbound

This paper cites Grounded language-image pre-training.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Grounded language-image pre-training

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:32.871605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:32.871605Z digest=sha256:c5d9a4887cbd41997cf0720d3bf2203e0b50196f1fff643d9fbe4ac28ac956cb

Observation 7196bd3a-3cfa-4d87-8e80-86f386fdbdef · outbound

This paper cites TempCompass: Do Video LLMs Really Understand Videos?.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding TempCompass: Do Video LLMs Really Understand Videos?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:33.000709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:33.000709Z digest=sha256:75da069230000f2b0bcad1ca5af40f39f57252cab4b48c283ac085871fa94265

Observation ef22287e-6b45-409f-83b5-5da365626f32 · outbound

This paper cites Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:33.180292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:33.180292Z digest=sha256:631ac5329610f367deffcf11331fed8cc81b8cfff2b19aa66988c3c75db44276

Observation c23e29bd-49d4-4ee2-a256-716b3e5727d9 · outbound

This paper cites Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:33.294475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:33.294475Z digest=sha256:65e710ccbf0d6ff95756a9940ddf4cb71c819fc9dc1508a70766b483b66ab6bf

Observation 37c6c68c-dfa0-458c-aa6f-d438fb959ff9 · outbound

This paper cites Ola: Pushing the Frontiers of Omni-Modal Language Model.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:33.446078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:33.446078Z digest=sha256:34f5f4cc335b665ec85b8a3f0cb850d79671a216b79b9c2b77f44d4b793921b7

Observation 765b1288-3e8a-46f3-a236-2e205096e505 · outbound

This paper cites Video detail caption, 2024.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Video detail caption, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:46.002685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:48:33.588967Z digest=sha256:819fe8faa6ae7bb1bd2c9df8763d4a9747b961dfd67baa9243339a91f8f2e98a

Observation d89edf69-5883-4e24-b7c9-6462a3481db4 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:45.768516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:48:33.712835Z digest=sha256:28ba5787573fbc2e3876cdde32106d3d8df7315b51d3b438e9eeb29e822a80b1

Observation a8110d16-6f0f-469c-847f-5dddec64fb52 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long- form video language understanding.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Egoschema: A diagnostic benchmark for very long- form video language understanding

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:45.555971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:48:33.843680Z digest=sha256:242e704ecaaec3ff8e7f3e33e2b6946b51e1b94e2b03600c52e92ca398536ae9

Observation e2f1aa5c-f826-4ed0-bcbd-fdbc2fef3663 · outbound

This paper cites Identifying the perceptual dimensions of visual complexity of scenes.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Identifying the perceptual dimensions of visual complexity of scenes

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:45.268685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:48:33.963318Z digest=sha256:7ea9b1c4ede08e62aa198661bcf8ff26d4baa5e1f84324bfbeb3e90e23ba5f25

Observation 48f1ad09-e144-43d9-8af3-4afa5c535bfe · outbound

This paper cites Hello gpt-4o.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Hello gpt-4o

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:44.992721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:48:34.090197Z digest=sha256:d52605efb1632e60dd8a01579f64b8f93e8bd45a8d723170c5a05c4e70492b2f

Observation 8b509cc9-209d-490f-bf22-4e17d8a20f18 · outbound

This paper cites Robustness analysis of video- language models against visual and language perturbations.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Robustness analysis of video- language models against visual and language perturbations

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:44.690772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:48:34.215109Z digest=sha256:57ff792d647b636f745496d644bc288ec41d0b219f154b6d78e8d65409f9af1f

Observation 8624ad6a-af68-4a15-b742-e3fedcfbcd4c · outbound

This paper cites Visual cot: Advancing multi-modal language models with a com- prehensive dataset and benchmark for chain-of-thought rea- soning.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Visual cot: Advancing multi-modal language models with a com- prehensive dataset and benchmark for chain-of-thought rea- soning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:44.432636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:48:34.327788Z digest=sha256:de1f9b01336fc787fd07f6bb9b54db8713a63481635517b62591856a2f718b85

Observation 3c072a10-ee74-403c-baa0-46613aa21e3a · outbound

This paper cites Complex narratives.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Complex narratives

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:44.211901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:48:34.436761Z digest=sha256:a09bfb80c71be1b12485dd88ac9037f78c7ab568698630db703b3c74edbb5c91

Observation 8b651d54-b4a8-49bb-baa6-c740cfbb7d51 · outbound

This paper cites A standardized set of 260 pictures: norms for name agreement, image agree- ment, familiarity, and visual complexity.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding A standardized set of 260 pictures: norms for name agreement, image agree- ment, familiarity, and visual complexity

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:43.994154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:48:34.545309Z digest=sha256:681775d36cb26ca14eb77285687f61cdaf8b7fedc14adcc98c99fff1cc40433a

Observation 83a81e53-1273-43de-940c-88950fadb301 · outbound

This paper cites Visual Agents as Fast and Slow Thinkers.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Visual Agents as Fast and Slow Thinkers

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:34.682053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:34.682053Z digest=sha256:f15bba31776c72f8e533c66c43f7b62ae2412b00075b66ac855c828a6de8167f

Observation 4be120cf-78ac-448b-a092-a75e3fc7285c · outbound

This paper cites Curious objects: How vi- sual complexity guides attention and engagement.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Curious objects: How vi- sual complexity guides attention and engagement

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:43.757236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:48:34.828072Z digest=sha256:b118d6bd301369798cab88e7ebe11ba3e2eb7b8e2df831c5daca039b929783e5

Observation c6d3932e-6194-486a-946a-42cb352693bb · outbound

This paper cites Cognitive load during problem solving: Ef- fects on learning.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Cognitive load during problem solving: Ef- fects on learning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:43.510219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:48:34.956258Z digest=sha256:ebdba70e8275546bf10392769f79afdb19b6cd5d617f93920d7ee791622be941

Observation 80026ae7-918d-488a-a749-558296988b73 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Gemini: A Family of Highly Capable Multimodal Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:35.106376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:35.106376Z digest=sha256:94eb9dcd164736cf5755574d4eff9ee019a46ffce9ed5f6b8b091aee971ae58a

Observation 42b5846c-c080-4b1c-b886-c596ca59958e · outbound

This paper cites Qwen2.5-vl, 2025.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Qwen2.5-vl, 2025

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:43.298939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:48:35.204531Z digest=sha256:2f8898a48f70c5885487a8a6c9e515c23722ff5e113dc5ff07a6f6edac6c5584

Observation fc779f9f-e966-485d-b1a5-df5446d54495 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models, 2023.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Chain-of-thought prompting elicits reasoning in large language models, 2023

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:43.068789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:48:35.335517Z digest=sha256:2d9cf8c2b44b8266279536a0b3c7ec844baca0eb4e9b678ba1d39889dfe9505c

Observation 544bbc8a-a175-4be3-93a3-8ee3c3b255c1 · outbound

This paper cites Star: A benchmark for situated reasoning in real-world videos.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Star: A benchmark for situated reasoning in real-world videos

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:42.832074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:48:35.438170Z digest=sha256:4842ebda8fce26816b52836d853dafbcb88b19abb683b11cfce9e3b221e686c2

Observation d1b49f1d-6949-4034-b8e3-6b30f8aca8ca · outbound

This paper cites Longvideobench: A benchmark for long-context interleaved video-language understanding, 2024.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Longvideobench: A benchmark for long-context interleaved video-language understanding, 2024

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:42.558730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:48:35.549580Z digest=sha256:35c9cb36f5155e133b051397dffc60fafba5e63e794237ad77b0715e568da3b3

Observation b79f48c1-7e3f-451e-8075-04db7d05888d · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Next-qa: Next phase of question-answering to explaining temporal actions

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:42.321877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:48:35.677619Z digest=sha256:98ba8698f06d777c953b25027ce19309c9ba841ff2cce0400ac624e7459b51ce

Observation be4cf59e-3032-4c5d-bc8b-fb4d9ffe5508 · outbound

This paper cites FunQA: Towards Surprising Video Comprehension.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding FunQA: Towards Surprising Video Comprehension

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:35.898141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:35.898141Z digest=sha256:e33291580b5fb44983a04685e14579ae7d2a6917b9e95457501a5e6b47df00e0

Observation 8615b8f0-4b17-4cfd-8b16-7733e1e6d649 · outbound

This paper cites Video question answer- ing via gradually refined attention over appearance and mo- tion.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Video question answer- ing via gradually refined attention over appearance and mo- tion

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:42.030420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:48:36.010706Z digest=sha256:43bda4d322d0ef129ffd1aeed2fff97ce232bab7e6b1444388b7f4f8c0ddf8e4

Observation e765a4ed-3fc3-4c43-a5a5-09eb60da7363 · outbound

This paper cites Sutd-trafficqa: A question answering benchmark and an efficient network for video rea- soning over traffic events.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Sutd-trafficqa: A question answering benchmark and an efficient network for video rea- soning over traffic events

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:41.727394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:48:36.149550Z digest=sha256:217ef1940a8d362ea37a43ad6e21ac5efa7ada75eec7389a9093f196c2d12a23

Observation 698ac212-0b18-4618-aebe-1225e15ab204 · outbound

This paper cites CLEVRER: CoLlision Events for Video REpresentation and Reasoning.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding CLEVRER: CoLlision Events for Video REpresentation and Reasoning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:36.257872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:36.257872Z digest=sha256:3c3d06f5cfd70a40368de09600d69e5b593b9cecb9a0dc9f7822515c09a34b1f

Observation b02e17e6-0383-4c89-b9c7-9ef1a00d10c9 · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:41.517617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:48:36.361773Z digest=sha256:03fc904fa2acb7c881e09540f87072dea6ac7a042fee0adb2e5354f5c621349b

Observation b49b54d2-abcb-4a86-aee5-8c7353405cd7 · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:41.286050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:48:36.499719Z digest=sha256:0820b763137367adfbdf943a8eb9a227120581ac4ecefdc90e63fa3bc477f2f5

Observation 29c5a716-8008-4753-9afe-d0c154f4e367 · outbound

This paper cites Social-iq: A question answer- ing benchmark for artificial social intelligence.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Social-iq: A question answer- ing benchmark for artificial social intelligence

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:41.022509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:48:36.656128Z digest=sha256:31251939d440c064fc5de884ea9992d2bd5c005076845079ba7845338123e782

Observation 263fd12e-d9a0-49d4-8027-6217a5f54e3e · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:36.756457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:36.756457Z digest=sha256:a5950bd5a4cc7fa5f6e87e41270d6253785c1a697dc63d6b8664f57b4fd24d1c

Observation bbe8c10f-0b64-438a-b966-196f252f44a8 · outbound

This paper cites B- avibench: Towards evaluating the robustness of large vision- language model on black-box adversarial visual-instructions,.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding B- avibench: Towards evaluating the robustness of large vision- language model on black-box adversarial visual-instructions,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:40.773699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:48:36.911831Z digest=sha256:d2a796eac0070342a260b1e8105e189aa5143907e35a1f8b501f8b214c7a205a

Observation 7d7965d5-2b29-4fac-8610-7ec7b3bbff7a · outbound

This paper cites Lmms- eval: Reality check on the evaluation of large multimodal models, 2024.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Lmms- eval: Reality check on the evaluation of large multimodal models, 2024

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:37.032657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:37.032657Z digest=sha256:634428b439e7849031bfdb46d3a386397d7d11eb5eb2a2651a7212598092b6d3

Observation fe22ba31-6ef0-422e-a89e-4b5188af112e · outbound

This paper cites Long Context Transfer from Language to Vision.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Long Context Transfer from Language to Vision

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:37.124321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:37.124321Z digest=sha256:0d53f13ad50488ed5290c40f2109e1420b1c6122826df7f3c7c2f34f163d518c

Observation 0d4b1904-9a86-4848-a923-e83820d76d7e · outbound

This paper cites Llava- next: A strong zero-shot video understanding model, 2024.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Llava- next: A strong zero-shot video understanding model, 2024

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:40.466201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:48:37.232691Z digest=sha256:7f0aabbb679442b08870df411dcd000ae4ccc4a016a0b41164cbc2cc8ce2ccc0

Observation 8364d8e5-e494-420c-88ef-7445affc3200 · outbound

This paper cites Video instruction tuning with synthetic data, 2024.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Video instruction tuning with synthetic data, 2024

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:40.216826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:48:37.388074Z digest=sha256:89b8d6ba1895660849454efd2ba2547f435d433939834d39c0ce24b6de42b74b

Observation 91e9c0ea-54b4-49e1-91ac-05896924ac87 · outbound

This paper cites Worldqa: Multimodal world knowledge in videos through long-chain reasoning, 2024.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Worldqa: Multimodal world knowledge in videos through long-chain reasoning, 2024

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:39.995188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:48:37.519206Z digest=sha256:1d72c04ebbf87d3c59dffeaa872841f19b7933069d9640c926d9f313e43cd3a6

Observation 2571e13e-793f-4fc5-bb6d-3ebd87c5ccdd · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding MLVU: Benchmarking Multi-task Long Video Understanding

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:37.637372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:37.637372Z digest=sha256:6337a8d1d8bcb6ffd7243c6b121a4f873c0b618eeb053608901a94ed879eb071

Observation ddccc800-85a9-4126-838a-ac6ec14d2461 · outbound

This paper cites Hierarchical video content description and summarization using unified semantic and visual similarity.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Hierarchical video content description and summarization using unified semantic and visual similarity

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:39.748619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:48:37.755434Z digest=sha256:d2af36ef59882bc3de1b5b45b47c529fd6cc098ad16e61010f97499c1c063cb6

Observation 12f08e3e-6f6d-4d24-8a18-9afa369badf2 · outbound

This paper cites In total, the annotation process cost 8227.32 human hours.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding In total, the annotation process cost 8227.32 human hours

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:39.494867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:48:37.878023Z digest=sha256:a8d648c278bac98c252d42f6be27ebcb24941e79f0ec225797af88a51ea8a282

Observation 5e1f3085-8a05-494a-8fe5-7f5260b953d0 · outbound

This paper cites • Aparaphrased correct be the set of videos where the para- phrased open-ended question is answered correctly.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding • Aparaphrased correct be the set of videos where the para- phrased open-ended question is answered correctly

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:39.209624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:48:37.984186Z digest=sha256:5557930c233ed97ef1ee273728f86e5f9373bfde0bb9407ef8d729c25192409f

Observation 4661c413-e575-40cf-8cb5-b6b4260aa7d2 · outbound

This paper cites 3 shows the prompt for evaluating open-ended an- swers.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding 3 shows the prompt for evaluating open-ended an- swers

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:38.958699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:48:38.104212Z digest=sha256:c42c0b64780ddb141683f04a3d0361ce193ba689ee8eeb6b0711fdaefcc3aed2

Observation f2b4833d-0cf6-4494-a4ee-ad5b033b6135 · outbound

This paper cites element” and “event.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding element” and “event

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:38.751950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:48:38.262186Z digest=sha256:85ea9d44c8c0bed4cdb1553fe96883b1d54f65591127f1eb8672c3476f52e752

Pith citing papers

Observation 029bc2f4-af1c-4e0d-9ee6-6c811843f413 · inbound

Video Reasoning without Training cites this paper.

Video Reasoning without Training Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T09:12:09.932226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:12:09.932226Z digest=sha256:a2119dc3386e5a01d02b1ef9585b541bd87ccc34a710676bf024575230f5fbb3