Pith. sign in

Paper Citation Record · LEDGER

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought

As of 19 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 2 inbound Pith citation observations for arXiv:2506.08817.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08817 v3

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:05:52.524432Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T17:05:08.124096Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T21:26:14.022696Z

Reference resolution

58 of 58 outbound references displayed

  • verified exact1
  • verified fuzzy3
  • unresolved54
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5bb3a05f-3d51-4414-aa4a-eae77d525432 · outbound

This paper cites Qwen2.5-VL Technical Report.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:48.635476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:48.635476Z digest=sha256:4e4ae8b0cb9df72bc1a679c65146b158b5e8b4528cf3f7e39b7dccb0ac1732aa

Observation 0c20c3de-2ff6-4232-adf4-ecdb902442a0 · outbound

This paper cites an unresolved cited work.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:48.718880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:48.718880Z digest=sha256:2376dbd8d36b9b1d6d52da621bc7691e3abfa39bf0b62db0c4b3cd09dd5374f4

Observation dc856c4e-d01d-4bd3-b423-99e84f180559 · outbound

This paper cites an unresolved cited work.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:48.783213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:48.783213Z digest=sha256:533bf86ff5d2a9d1e38d431306e3e1a3a737e65ff3ddd5d8310c977c17e7227a

Observation ad00dba6-a2b0-420e-8401-97d03d124e81 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Training Verifiers to Solve Math Word Problems

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:48.907545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:48.907545Z digest=sha256:42a5de5f10974ddb386e317390d255cd9e56230eadb714159841fb709508b3ce

Observation 2d51afe6-d7e0-45e8-9e31-499e2dac66c6 · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:48.972426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:48.972426Z digest=sha256:d89cddde473f593336c89a558b7b1e1ea18b2fe538cb98ff870d2b448e6ce11b

Observation e8f286c9-559b-43cd-9f87-03f5b28a30cc · outbound

This paper cites an unresolved cited work.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:05:54.623729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:05:49.032774Z digest=sha256:4bdbc2c9d05617e6412a8eb3a3ab37034e0fa8ea967bee57a31d5c3696fa1f10

Observation 45845968-340f-41a2-83e0-240c62ece8e1 · outbound

This paper cites something something.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought something something

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:49.127473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:49.127473Z digest=sha256:8ebb61392932d55a5efbfa2f4dfb3e2d199de4465bc8dc00a0afe10c6a92fce2

Observation be30a964-10a9-4205-8fd0-0c8eb3993f5a · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:49.226801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:49.226801Z digest=sha256:958596084514292f96655b6149e0d68dc53bc65ec84a681d1b3ac0dadc98b1d7

Observation 34742e67-040f-4a7a-9e70-068b551af868 · outbound

This paper cites an unresolved cited work.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:05:54.503344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:05:49.272926Z digest=sha256:8ddeb41f3ce5b2797c04f4fe712021cda21756abc566c80f14024652740ae1fb

Observation 7289eca3-1ae7-408e-9086-e68a09dd1326 · outbound

This paper cites an unresolved cited work.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:05:54.412188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:05:49.337060Z digest=sha256:cbc303c541944afb669e3d979c1ea608b8d6e4db7ae2994ed5ac0588538b661f

Observation 8dc862f6-e5e3-43bd-821d-b8a7a9c1b479 · outbound

This paper cites an unresolved cited work.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:05:54.336836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:05:49.405079Z digest=sha256:d95f02dcf3165380017815e6e9c5cbe29d970232beb4643ecdd425e686df9895

Observation e6b18c16-faef-4a5a-b53b-8358cfceb7fa · outbound

This paper cites an unresolved cited work.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:05:54.249080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:05:49.482878Z digest=sha256:6c3516e45f77cf570a339289c76dd65ab5be2a09e11d876c9e45cc7db9bf0a07

Observation c35d9176-81e6-4102-8e60-d4847392f46f · outbound

This paper cites an unresolved cited work.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:05:54.197627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:05:49.536320Z digest=sha256:700f3891194d15c979eac26f4fcf2509c5e2920914c8c8246633291f14e6e2e3

Observation f7ca53aa-bc24-46c9-91d9-710152be3999 · outbound

This paper cites an unresolved cited work.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:49.556596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:49.556596Z digest=sha256:878eb855d988b6d75b698f0f4376e097cc40c40178ce9b1f460f3dc419a5f940

Observation a9fb629d-3629-47a8-91f8-94e5bb752c8c · outbound

This paper cites Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:49.657151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:49.657151Z digest=sha256:a076660e3d5e11834922f3594cc68b5fa8ed4c3e40d9a9dab6fe0bd689bfd573

Observation 870a9dec-b907-4825-8931-9b5e0a04c1fe · outbound

This paper cites an unresolved cited work.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:49.733931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:49.733931Z digest=sha256:7ddd89363e6714d3cae8d713ee1795696437ff11b884a4ca2340b06209e2e3f5

Observation 37db3b06-e5cb-4610-a43d-2823f7bed0c4 · outbound

This paper cites GPT-4o System Card.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought GPT-4o System Card

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:49.800282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:49.800282Z digest=sha256:4ca66af160273da4ee106ab0062397f8b25a56904b0a0fbd23eb5663c4a38377

Observation 6bfa53ee-8221-4807-8584-9be349890b09 · outbound

This paper cites an unresolved cited work.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:05:54.107770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:05:49.844975Z digest=sha256:d22d7b0f8c747ebcfd67e446c092c335c3d1156efc1a81336d1982801f370414

Observation 9244296b-2b9f-4ed3-876f-d9d4dd91443a · outbound

This paper cites The Kinetics Human Action Video Dataset.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought The Kinetics Human Action Video Dataset

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:49.922593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:49.922593Z digest=sha256:55af5be4d45c4a13ea18982e97267e56e97301dd9ab8325d636a2f1f4721a1df

Observation a3a44d40-732c-40c1-bdc7-c74c8059104a · outbound

This paper cites an unresolved cited work.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:49.998300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:49.998300Z digest=sha256:57b598d19f5e6de3f7ab65fb52555f6a68dec7999741353ee0c33524ce678e2f

Observation d0f136bc-4194-4ab5-9013-016a1e8bf1f0 · outbound

This paper cites an unresolved cited work.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:50.106312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:50.106312Z digest=sha256:869d976294f4494578829404309d611418ca94c570a8a626a6dd54f291936e0a

Observation eb9729d3-c2fb-440f-8cb3-93a959fc0ed8 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought LLaVA-OneVision: Easy Visual Task Transfer

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:50.302426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:50.302426Z digest=sha256:f0878da9cdda627e2e739f2c3a0cfe7e26dfba8582c65d7a257a5683c74f2075

Observation 85e6105e-b847-47d6-90f3-150fec9a4228 · outbound

This paper cites VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:50.381964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:50.381964Z digest=sha256:9cd5a607195e0e40ac8352ba791029ce1e8baf1398f6c997ad9d5b9aff556457

Observation 011268a2-bc2b-48f8-b15a-1da400af515b · outbound

This paper cites an unresolved cited work.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:50.434498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:50.434498Z digest=sha256:4f89e561547e61d06764913da1dbc6d572602945ce910a5d6f4c8024c502f379

Observation ce02b562-e23d-4321-8962-9fb6545b3e90 · outbound

This paper cites an unresolved cited work.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:05:53.914958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:05:50.524811Z digest=sha256:e94aa5d42ae1daa2ca5fcafef425684933e9c3863210752d94918af6b9526532

Observation 07fb5ba5-d66d-4b7f-9f0d-f28c1cc6b393 · outbound

This paper cites an unresolved cited work.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:05:53.838608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:05:50.570583Z digest=sha256:8d5efb88f99a5dca8dbbfac47ad934e5e9c28764dfc3b12229fbbb1b606e82a6

Observation 5f02ac96-0acf-40be-bbac-f1541978648b · outbound

This paper cites A Hierarchical Reinforcement Learning Framework for Multi-UAV Combat Using Leader-Follower Strategy.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought A Hierarchical Reinforcement Learning Framework for Multi-UAV Combat Using Leader-Follower Strategy

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:05:52.891078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:05:50.617505Z digest=sha256:2b539383962018f740ccfe18d03e6031497a7999fd08ce5577592916ad9eccc0

Observation 52826fff-791f-4ee0-b6e3-2dc7c6cc644e · outbound

This paper cites an unresolved cited work.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:05:53.733960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:05:50.695041Z digest=sha256:280aafa45b35672f9c319f2ccaf19acf71786f87221686112d2fc72a74c2bf05

Observation 545ae7e5-360c-463b-94b1-9b9bbdeb0933 · outbound

This paper cites an unresolved cited work.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:05:53.623654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:05:50.749177Z digest=sha256:691a5b9363e88803ac50bdd03a9114a7b16701cc2a02ab3faa67b6c38085e05f

Observation 9b435fac-ba84-4bce-a64c-509e62ab856b · outbound

This paper cites Audio-Visual LLM for Video Understanding.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Audio-Visual LLM for Video Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:50.794819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:50.794819Z digest=sha256:5e187545403af33bc95d063cbbc490ba81ecb6659809d6b0a6972b930a3c7c73

Observation f37a610c-9f3a-4389-ab15-0faeda8ee81e · outbound

This paper cites an unresolved cited work.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:05:53.475500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:05:50.823238Z digest=sha256:e113afda96e57675315c38c87cd70fa0d69fbea688476f0a0224dfbcf2a61745

Observation 3b56d1d7-9785-4bd2-8a21-d6024728f220 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:50.986480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:50.986480Z digest=sha256:6538657ce5fc6ffc60385c7b2e4903022f8bb7cb3e0135c8ba01badf40d2522a

Observation 670e4871-610a-48f4-beaa-f688a8d2c2a2 · outbound

This paper cites CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:51.040823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:51.040823Z digest=sha256:421eaa5789440b5f3b8e4117eee42f5335b06577a65951efbfc9b589ac09dda9

Observation 8268e98b-cb9a-496f-be14-873e0afb1ee1 · outbound

This paper cites AffordGrasp: In-Context Affordance Reasoning for Open-Vocabulary Task-Oriented Grasping in Clutter.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought AffordGrasp: In-Context Affordance Reasoning for Open-Vocabulary Task-Oriented Grasping in Clutter

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:51.113645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:51.113645Z digest=sha256:710d3d342c1c2881d809ed97509b5f9422fdc3c1442f59b95b99edba456f7e79

Observation 3c9ea9b9-9b8b-455f-9416-0a93b2cecd2f · outbound

This paper cites an unresolved cited work.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:51.156506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:51.156506Z digest=sha256:20414509a4cd46d4a8d36cd7b4d8e816d16e3a9518fee932006d0c2f7d3d0b41

Observation 154ac2fd-1843-47dc-a81a-9cbd42e0e491 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:51.174239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:51.174239Z digest=sha256:e1c668d180d30dc3415c6c224d3144f2190b646636e8d59ce5fdc2f0af81cbba

Observation dfcfbc84-64a7-471a-b4e7-bcf3b59f2c27 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:51.222586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:51.222586Z digest=sha256:87e556f535d528f719c38f504c71702d87506c94c2ae3fae82d3143a5e472013

Observation 06d262f9-5160-4354-bb14-4ce73c8dbb54 · outbound

This paper cites ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:51.278905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:51.278905Z digest=sha256:18125c6ca00f191a0d0cc79850ccbde1e68f592d45ca7f8a34612c2c5c5479f9

Observation be460529-5b58-4610-be7f-b68b835bcbec · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:51.321299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:51.321299Z digest=sha256:8b9b2cf384492339e5c7a9fd4540a9dc5cdf69b3fca2efc5c6b9fd5afb91447e

Observation dc595d87-8263-4001-ad95-c20fdf1fbb47 · outbound

This paper cites InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:51.389244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:51.389244Z digest=sha256:8bae91c81632076eb4482537d72d5fb1388f94e0df0297add8a0f514127838fc

Observation fad9e6b4-fe1e-405e-810d-9ec4116d083e · outbound

This paper cites an unresolved cited work.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:51.452700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:51.452700Z digest=sha256:be778662bfba1b9e1c85deed597d6d74d2fb5ce61620646aa39c8f73d445a83e

Observation 7e64901f-1273-4103-ad89-2459ca0a4ad2 · outbound

This paper cites an unresolved cited work.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:51.531489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:51.531489Z digest=sha256:579da572ea81cccf6415bd035bc57c4c39d498b816a8cb98566e06497a625820

Observation 383f9987-eab1-40c9-98ae-e2b977407e5d · outbound

This paper cites an unresolved cited work.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:51.727388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:51.727388Z digest=sha256:23ff8d57c5c3ecdb06ba5e1cafc4dba33d243b27cad850373794608e0d9eb4b0

Observation cca739ce-27b4-4b82-bfb1-b54fc1421926 · outbound

This paper cites an unresolved cited work.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:51.766666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:51.766666Z digest=sha256:cc8f6a9320bad1e509d36b69a4ac9daf125fe4deba1e98acf10a01b85c0a16d3

Observation 25eb0420-15b6-4e50-a4d0-e66fe251e329 · outbound

This paper cites an unresolved cited work.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:05:53.208143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:05:51.808790Z digest=sha256:8f6755739929e7dea0b7128930c35e45d3ca3f3cbb1898e59e5e03f8b7c38b39

Observation 5edce4f5-8057-4003-b349-e0275c3d0714 · outbound

This paper cites Unhackable Temporal Rewarding for Scalable Video MLLMs.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unhackable Temporal Rewarding for Scalable Video MLLMs

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:51.881077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:51.881077Z digest=sha256:e28d633c6c5f06e5596bf7d787eccd584549be0ef9034e6aaece9c4d91bc8a71

Observation 5c434370-7d88-4e53-9b1d-3bc0a6e97398 · outbound

This paper cites an unresolved cited work.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:51.932240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:51.932240Z digest=sha256:1d97b1ea087485fc98fa3a0e9c1fff64159f542f85c94446cc1a02e891c1b924

Observation f74ad4db-bbe2-4757-bf7f-74ee999b15a5 · outbound

This paper cites Multi-Floor Zero-Shot Object Navigation Policy.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Multi-Floor Zero-Shot Object Navigation Policy

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:52.050076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:52.050076Z digest=sha256:08b76db8050ce25987e99564ebc219ec79e53470a7782b2ba359e3c3322f027e

Observation 3aa44a55-2906-47d5-8c55-b4bd0248f676 · outbound

This paper cites TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:52.150407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:52.150407Z digest=sha256:f992e73fff6b18a6215a70117dda5f6d36604232210ae3ffba845fdef313472d

Observation f030983c-b298-4f1d-aa5a-bf5106fa7115 · outbound

This paper cites an unresolved cited work.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:05:53.121846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:05:52.262230Z digest=sha256:acb1876d24fca23f6ae3bc83bb5bcdb4f0acf0d764b7819ce614718e12078df9

Observation c48b353d-8920-456e-bd51-8f4dbf33385e · outbound

This paper cites MapNav: A Novel Memory Representation via Annotated Semantic Maps for Vision-and-Language Navigation.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought MapNav: A Novel Memory Representation via Annotated Semantic Maps for Vision-and-Language Navigation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:51.969310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:51.969310Z digest=sha256:2dba4c3fbc237199c28ba8a88d47cdaea235a8037c9d6fe1f6f3aee3f8d6aca3

Observation 13dd506a-9ce8-4054-af5f-43844a6ab1c3 · outbound

This paper cites Multimodal Chain-of-Thought Reasoning in Language Models.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Multimodal Chain-of-Thought Reasoning in Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:52.403000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:52.403000Z digest=sha256:333e609fc26d7c192ba7dfd6988ab350a62359705142d09ee3b5ca8bca8a54f9

Observation ba8c163e-b26f-4ddd-8163-f842de1875b1 · outbound

This paper cites an unresolved cited work.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:52.452676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:52.452676Z digest=sha256:5a147c6f534554f9fd99c695e8c0d8d0c8bf6c013ef5749bd7dfa2221a36e422

Observation f7c99c30-876e-410f-a74d-9c6175633fc5 · outbound

This paper cites Automatic Chain of Thought Prompting in Large Language Models.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Automatic Chain of Thought Prompting in Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:52.320226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:52.320226Z digest=sha256:449648fe9fc33dcc1319f2edb2cb92089301d44544d14cc64eb4addaa0a2add8

Observation fd1f3ac7-99db-4b8d-8c00-7f2e08109049 · outbound

This paper cites InProceedings of the IEEE International Conference on Computer Vision.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought InProceedings of the IEEE International Conference on Computer Vision

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:05:54.009897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:05:50.187477Z digest=sha256:f57c62ff27a4469648e1f654284c2b9eed3a991ba59982842d802a6d31a40f30

Observation 03b014c6-c455-4380-aa62-984570a8de8e · outbound

This paper cites InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:52.524432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:52.524432Z digest=sha256:e9eec1a4dc73993c385ede6c9972e4b45f95e98b3277218f3a04f32e9a8a052b

Observation 9bec9162-b2a2-4d5a-b069-e8adac345cfa · outbound

This paper cites In European Conference on Computer Vision.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought In European Conference on Computer Vision

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:05:53.292380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:05:51.609150Z digest=sha256:05ec415b44d5bfd038b7a26858496bfc433043501c13be209089c0925798cfaf

Observation 50aa44b2-301c-4222-b0e4-d1aa835901ab · outbound

This paper cites InIEEE International Conference on Multimedia and Expo (ICME).

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought InIEEE International Conference on Multimedia and Expo (ICME)

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:05:53.399864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:05:50.898675Z digest=sha256:b6236b1430a736113e7debe5d9c8a476d3d1e60cdea5b6ad2abdaae2329cef2e

Pith citing papers

Observation 3c1f2cfd-c49a-4438-a077-9db5e8271014 · inbound

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking cites this paper.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:50:17.350435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:9108078d4ae9af1fc31a423f008eb4538da14266fb11a362d1e7c36d9ab7e250

Observation 24ee94e1-8303-4108-aa2e-2f058298ad39 · inbound

OneVLA: A Unified Framework for Embodied Tasks cites this paper.

OneVLA: A Unified Framework for Embodied Tasks Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:26:14.024501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T17:05:08.124096Z digest=sha256:a2a984ebd9e59df9cd21bf4a56881092ed5355d572ec5ed095299a70eefc72fa