Pith. sign in

Paper Citation Record · LEDGER

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought

As of 7 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 2 inbound Pith citation observations for arXiv:2506.08817.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08817 v3

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:05:52.524432Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T17:05:08.124096Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T21:26:14.022696Z

Reference resolution

58 of 58 outbound references displayed

  • verified exact1
  • verified fuzzy3
  • unresolved54
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5bb3a05f-3d51-4414-aa4a-eae77d525432 · outbound

This paper cites Qwen2.5-VL Technical Report.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:48.635476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:48.635476Z digest=sha256:b5cf36d51731f81d52597322f37f6c9360e0630a71cbf6d0247cef4808bb8de6

Observation 0c20c3de-2ff6-4232-adf4-ecdb902442a0 · outbound

This paper cites an unresolved cited work.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:48.718880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:48.718880Z digest=sha256:a7e1029653246f8dd2645625f74ada40cbef179ec4fb13159c753c4bf84a5a5a

Observation dc856c4e-d01d-4bd3-b423-99e84f180559 · outbound

This paper cites an unresolved cited work.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:48.783213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:48.783213Z digest=sha256:daf4fdf38b618c57637162f78b42322910603775b46b7a3e46fa5c38f9f4d632

Observation ad00dba6-a2b0-420e-8401-97d03d124e81 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Training Verifiers to Solve Math Word Problems

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:48.907545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:48.907545Z digest=sha256:598c658b82ffbb50b0c4e7afe6d800b807a37b1e5c84fb9feda44908f965c8b5

Observation 2d51afe6-d7e0-45e8-9e31-499e2dac66c6 · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:48.972426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:48.972426Z digest=sha256:8a3ed7c265aea7df3e38c125c1073782fe9ac92200ac32f06dd6c188924003d5

Observation e8f286c9-559b-43cd-9f87-03f5b28a30cc · outbound

This paper cites an unresolved cited work.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:05:54.623729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:05:49.032774Z digest=sha256:ee24a2df832ff142392f7df4a322838af3224b9285faf5f0a89b7dc2d20be8a4

Observation 45845968-340f-41a2-83e0-240c62ece8e1 · outbound

This paper cites something something.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought something something

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:49.127473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:49.127473Z digest=sha256:c44199449d9a7fa21ad2b480d2de657e787c5d356ab1ba3822727dbd57ac122d

Observation be30a964-10a9-4205-8fd0-0c8eb3993f5a · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:49.226801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:49.226801Z digest=sha256:04547a0e0ad17fff5e9bce260b04215f515d4b2212b5432d294f104af3d22a7e

Observation 34742e67-040f-4a7a-9e70-068b551af868 · outbound

This paper cites an unresolved cited work.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:05:54.503344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:05:49.272926Z digest=sha256:545d2d47a7b0edca33ec6c42d52cf178377bb1e27da98be180c1eccf7da2599b

Observation 7289eca3-1ae7-408e-9086-e68a09dd1326 · outbound

This paper cites an unresolved cited work.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:05:54.412188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:05:49.337060Z digest=sha256:739780954a83bd7d5b3b208aa3b3f92613377e4392ef6d019b7c2394c75578ce

Observation 8dc862f6-e5e3-43bd-821d-b8a7a9c1b479 · outbound

This paper cites an unresolved cited work.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:05:54.336836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:05:49.405079Z digest=sha256:81eada61b6a01e717a9507c716ee9d4d938a3c5651c87dfe2add890b1ff68807

Observation e6b18c16-faef-4a5a-b53b-8358cfceb7fa · outbound

This paper cites an unresolved cited work.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:05:54.249080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:05:49.482878Z digest=sha256:e50a1deece0a9db8da0c12673a3f2a49a7414aad044f57d896e78f3595d2ab2e

Observation c35d9176-81e6-4102-8e60-d4847392f46f · outbound

This paper cites an unresolved cited work.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:05:54.197627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:05:49.536320Z digest=sha256:92a8765e1397c542aa708487c966431388cd036d88789b135413f284543cde03

Observation f7ca53aa-bc24-46c9-91d9-710152be3999 · outbound

This paper cites an unresolved cited work.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:49.556596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:49.556596Z digest=sha256:ae88e58f41e7bd0a87685cba077e756bc6aefc645e72698857c8ac2a19de5a5f

Observation a9fb629d-3629-47a8-91f8-94e5bb752c8c · outbound

This paper cites Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:49.657151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:49.657151Z digest=sha256:d5b1eb1bb5b631046c5cac3597415f41e11e4ad3c5b91865e31695d365210103

Observation 870a9dec-b907-4825-8931-9b5e0a04c1fe · outbound

This paper cites an unresolved cited work.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:49.733931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:49.733931Z digest=sha256:45ccc4f278ee451b0c12fc57e4e06ed222429a9a289d51ace6ccd5ee532aa3c6

Observation 37db3b06-e5cb-4610-a43d-2823f7bed0c4 · outbound

This paper cites GPT-4o System Card.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought GPT-4o System Card

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:49.800282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:49.800282Z digest=sha256:ba3b6c37c85ac2f9c8b5d6a5897f1eb4dfa8413fe34de2c14903e16b41126f79

Observation 6bfa53ee-8221-4807-8584-9be349890b09 · outbound

This paper cites an unresolved cited work.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:05:54.107770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:05:49.844975Z digest=sha256:782f0b4ba7f8ba1ccf2656698d35d754c5d7ad3cf341ac370806dee9197155bb

Observation 9244296b-2b9f-4ed3-876f-d9d4dd91443a · outbound

This paper cites The Kinetics Human Action Video Dataset.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought The Kinetics Human Action Video Dataset

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:49.922593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:49.922593Z digest=sha256:df05ec40e7a879a2e0a2e6cbfaf5ab9cb9b37da49101d22b17f3926b39208ce2

Observation a3a44d40-732c-40c1-bdc7-c74c8059104a · outbound

This paper cites an unresolved cited work.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:49.998300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:49.998300Z digest=sha256:374d34b1894323abf5129cf93483aecf03875ce953654c31050def9ac732334c

Observation d0f136bc-4194-4ab5-9013-016a1e8bf1f0 · outbound

This paper cites an unresolved cited work.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:50.106312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:50.106312Z digest=sha256:845865e1d0b75e8e2240154afb08294dab1c7c1a134a6d6b20ed3b55b7d7d920

Observation eb9729d3-c2fb-440f-8cb3-93a959fc0ed8 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought LLaVA-OneVision: Easy Visual Task Transfer

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:50.302426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:50.302426Z digest=sha256:bcda890983ddc929d323efb9f4146e45c424d988b613d659d245519154291066

Observation 85e6105e-b847-47d6-90f3-150fec9a4228 · outbound

This paper cites VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:50.381964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:50.381964Z digest=sha256:e161ed2dde4657505aae2db5ab8c1092755671ef27fff9efa7d5e728b6383e06

Observation 011268a2-bc2b-48f8-b15a-1da400af515b · outbound

This paper cites an unresolved cited work.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:50.434498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:50.434498Z digest=sha256:d49395c721f57219c67e64177d8456716900e5812464840dbaf64063ca6fd8b5

Observation ce02b562-e23d-4321-8962-9fb6545b3e90 · outbound

This paper cites an unresolved cited work.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:05:53.914958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:05:50.524811Z digest=sha256:88346236d3451d74dbc9a4426fdc889b8de53f968545a167f071de18aefc4d6f

Observation 07fb5ba5-d66d-4b7f-9f0d-f28c1cc6b393 · outbound

This paper cites an unresolved cited work.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:05:53.838608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:05:50.570583Z digest=sha256:9209fb3e691784bd2827a6038c4ded5a3650fe5ce1fafbd8ceec355b09b0e2c9

Observation 5f02ac96-0acf-40be-bbac-f1541978648b · outbound

This paper cites A Hierarchical Reinforcement Learning Framework for Multi-UAV Combat Using Leader-Follower Strategy.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought A Hierarchical Reinforcement Learning Framework for Multi-UAV Combat Using Leader-Follower Strategy

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:05:52.891078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:05:50.617505Z digest=sha256:689e20127f7c95ee1ae84869e5730e70cc0089cfa716037ef4cfd6551e3576b5

Observation 52826fff-791f-4ee0-b6e3-2dc7c6cc644e · outbound

This paper cites an unresolved cited work.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:05:53.733960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:05:50.695041Z digest=sha256:4a191fb33552fd621ab9f4e1616656be01c646ab29fa4edc9ed410f20ef93761

Observation 545ae7e5-360c-463b-94b1-9b9bbdeb0933 · outbound

This paper cites an unresolved cited work.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:05:53.623654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:05:50.749177Z digest=sha256:e2f3bdb2a3ff3ea81dd5c141c121f46b75b79e7a06e19ddc7d088bc3edeef245

Observation 9b435fac-ba84-4bce-a64c-509e62ab856b · outbound

This paper cites Audio-Visual LLM for Video Understanding.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Audio-Visual LLM for Video Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:50.794819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:50.794819Z digest=sha256:3c3fb7d7bfeb8d0024e491a550bb11f1ed5dba3f80bba1d1bea684607963d7ce

Observation f37a610c-9f3a-4389-ab15-0faeda8ee81e · outbound

This paper cites an unresolved cited work.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:05:53.475500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:05:50.823238Z digest=sha256:2426aa3cbd4f37d522b62bba2aed738855eb215192952b63f73d946d3393dc9a

Observation 3b56d1d7-9785-4bd2-8a21-d6024728f220 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:50.986480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:50.986480Z digest=sha256:c8afb2d77156b785ca312bbf2cf958af7d5ed3d348f065f93a993c13926719ae

Observation 670e4871-610a-48f4-beaa-f688a8d2c2a2 · outbound

This paper cites CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:51.040823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:51.040823Z digest=sha256:7f3ffed628c77649daebb70e58bf9c8da9167add63e219c3947890ca6f3d7b8f

Observation 8268e98b-cb9a-496f-be14-873e0afb1ee1 · outbound

This paper cites AffordGrasp: In-Context Affordance Reasoning for Open-Vocabulary Task-Oriented Grasping in Clutter.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought AffordGrasp: In-Context Affordance Reasoning for Open-Vocabulary Task-Oriented Grasping in Clutter

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:51.113645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:51.113645Z digest=sha256:b0bb059e9739f627dde5177e73e46380c90fcfdf3b478110c9f1887a9c92d0c9

Observation 3c9ea9b9-9b8b-455f-9416-0a93b2cecd2f · outbound

This paper cites an unresolved cited work.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:51.156506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:51.156506Z digest=sha256:3019d3eec7e6ed921cfc9ba7610337f8750f7d995600ff94c655708d421f93c7

Observation 154ac2fd-1843-47dc-a81a-9cbd42e0e491 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:51.174239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:51.174239Z digest=sha256:00a2fc019c0f60aadd053d3d83515c7780e666009de9a51669ab3fff8bc16271

Observation dfcfbc84-64a7-471a-b4e7-bcf3b59f2c27 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:51.222586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:51.222586Z digest=sha256:4679ac255af446998b8a00f0dffad9fc87523153928a1753417fa0fed891eae9

Observation 06d262f9-5160-4354-bb14-4ce73c8dbb54 · outbound

This paper cites ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:51.278905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:51.278905Z digest=sha256:fc3a4e9d5053ad9e9d173cd3d43f320afc2d3ca5300a215e971a814a7186d8bc

Observation be460529-5b58-4610-be7f-b68b835bcbec · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:51.321299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:51.321299Z digest=sha256:2a6ade8ee3114650d5eeceb9d3ddeeaa0a4412e6056894c122bed7f965a73325

Observation dc595d87-8263-4001-ad95-c20fdf1fbb47 · outbound

This paper cites InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:51.389244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:51.389244Z digest=sha256:e7d380c8ef721fdfd90dd4614108898159ed450bc59a9e2029e58968495b7f36

Observation fad9e6b4-fe1e-405e-810d-9ec4116d083e · outbound

This paper cites an unresolved cited work.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:51.452700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:51.452700Z digest=sha256:188ea43c8cf5d1be702e6e881f0143a6418f2687b0d8462a88bf85c546bba7db

Observation 7e64901f-1273-4103-ad89-2459ca0a4ad2 · outbound

This paper cites an unresolved cited work.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:51.531489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:51.531489Z digest=sha256:4bd9334ebd25598eb2b5ed8538667df9d449e3657b921d3ef243ec7f8c421aaf

Observation 383f9987-eab1-40c9-98ae-e2b977407e5d · outbound

This paper cites an unresolved cited work.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:51.727388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:51.727388Z digest=sha256:20bf6adeb0aa41e272276a5018f1195201a1b24ac9e59d99d1771f5c82694d2b

Observation cca739ce-27b4-4b82-bfb1-b54fc1421926 · outbound

This paper cites an unresolved cited work.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:51.766666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:51.766666Z digest=sha256:6a47a50789e94f0dbc7e6c8f8ec06742adee7cb31b319f5ec0d461942fe64b91

Observation 25eb0420-15b6-4e50-a4d0-e66fe251e329 · outbound

This paper cites an unresolved cited work.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:05:53.208143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:05:51.808790Z digest=sha256:96e08080394cb5160b280640ad5b61ca99259e292a5abe245cfb82d4ab0b74a1

Observation 5edce4f5-8057-4003-b349-e0275c3d0714 · outbound

This paper cites Unhackable Temporal Rewarding for Scalable Video MLLMs.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unhackable Temporal Rewarding for Scalable Video MLLMs

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:51.881077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:51.881077Z digest=sha256:dd9fffeb4656ef5712744d2a7c3878632ce1a8ab3bfe2f10d4bbe968c3e4964b

Observation 5c434370-7d88-4e53-9b1d-3bc0a6e97398 · outbound

This paper cites an unresolved cited work.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:51.932240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:51.932240Z digest=sha256:16478bd3990c9f5893f09c6200d5938d1e4f35bd7c6893dd2900df09a9546f0c

Observation f74ad4db-bbe2-4757-bf7f-74ee999b15a5 · outbound

This paper cites Multi-Floor Zero-Shot Object Navigation Policy.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Multi-Floor Zero-Shot Object Navigation Policy

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:52.050076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:52.050076Z digest=sha256:4d8b48a97c04d027435845c9d430c5a6a6f45aaf986919c8f2ca8cca0fb34265

Observation 3aa44a55-2906-47d5-8c55-b4bd0248f676 · outbound

This paper cites TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:52.150407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:52.150407Z digest=sha256:d387caafa0f6640279bd28ebb1edceca27d9a5e04895577bbc2937b7a84f932b

Observation f030983c-b298-4f1d-aa5a-bf5106fa7115 · outbound

This paper cites an unresolved cited work.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:05:53.121846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:05:52.262230Z digest=sha256:2f08e6972e112519dfbf9b5ea184ea4feab1dee57224fb19b1a4453172d532b3

Observation c48b353d-8920-456e-bd51-8f4dbf33385e · outbound

This paper cites MapNav: A Novel Memory Representation via Annotated Semantic Maps for Vision-and-Language Navigation.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought MapNav: A Novel Memory Representation via Annotated Semantic Maps for Vision-and-Language Navigation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:51.969310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:51.969310Z digest=sha256:b7b7187dd78f1a65181e0ae3e8d6f808d3a3279386d0a87ac5a2cf547390149f

Observation 13dd506a-9ce8-4054-af5f-43844a6ab1c3 · outbound

This paper cites Multimodal Chain-of-Thought Reasoning in Language Models.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Multimodal Chain-of-Thought Reasoning in Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:52.403000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:52.403000Z digest=sha256:0f6e8cb9d593c57cf6db0d1ceaf6e143bce438a51ab9cb8c526fbc2983394252

Observation ba8c163e-b26f-4ddd-8163-f842de1875b1 · outbound

This paper cites an unresolved cited work.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:52.452676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:52.452676Z digest=sha256:4b3c536997defcaced409089900fce259322700d5fd7cf4c68ceb8ceb1c5c678

Observation f7c99c30-876e-410f-a74d-9c6175633fc5 · outbound

This paper cites Automatic Chain of Thought Prompting in Large Language Models.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Automatic Chain of Thought Prompting in Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:52.320226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:52.320226Z digest=sha256:8371150613cafcdbdee8b9db4c473379a5045f4b4d993c0e67abad33df7ae001

Observation fd1f3ac7-99db-4b8d-8c00-7f2e08109049 · outbound

This paper cites InProceedings of the IEEE International Conference on Computer Vision.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought InProceedings of the IEEE International Conference on Computer Vision

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:05:54.009897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:05:50.187477Z digest=sha256:d8cf48282301ed0e4e0e8f41b9e0e1a175e87e39f07bf6a74255f099804af524

Observation 03b014c6-c455-4380-aa62-984570a8de8e · outbound

This paper cites InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:52.524432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:52.524432Z digest=sha256:d67fd13f6f0503e9d94afa9d3ca232e275ccc650132fb3465e20b51daa28cda0

Observation 9bec9162-b2a2-4d5a-b069-e8adac345cfa · outbound

This paper cites In European Conference on Computer Vision.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought In European Conference on Computer Vision

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:05:53.292380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:05:51.609150Z digest=sha256:cb896e55e509cdb0bd3789297fe932ecf1bd181a8048636ba0cc7935135771dd

Observation 50aa44b2-301c-4222-b0e4-d1aa835901ab · outbound

This paper cites InIEEE International Conference on Multimedia and Expo (ICME).

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought InIEEE International Conference on Multimedia and Expo (ICME)

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:05:53.399864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:05:50.898675Z digest=sha256:dbe4a816f52d15cb724e305f3285803fa65beb2ce359b0368dbd9ee63c28f279

Pith citing papers

Observation 3c1f2cfd-c49a-4438-a077-9db5e8271014 · inbound

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking cites this paper.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:50:17.350435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:be4b230acd3314c528b90d15f39d54386b9a51e2b358f5e56709afbabc7408c3

Observation 24ee94e1-8303-4108-aa2e-2f058298ad39 · inbound

OneVLA: A Unified Framework for Embodied Tasks cites this paper.

OneVLA: A Unified Framework for Embodied Tasks Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:26:14.024501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T17:05:08.124096Z digest=sha256:b71e9b96afccb95b7fbde40766919b7fef68d67cfe43e9019c3ddb77b8236fc4