Pith. sign in

Paper Citation Record · LEDGER

Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 22 inbound Pith citation observations for arXiv:2501.03230.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.03230 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 22 of 22 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:18:57.344627Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T08:57:47.799001Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2347e678-cf70-4e00-b626-809cef25365c · inbound

Uncertainty-o: One Model-agnostic Framework for Unveiling Uncertainty in Large Multimodal Models cites this paper.

Uncertainty-o: One Model-agnostic Framework for Unveiling Uncertainty in Large Multimodal Models Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:36:15.540187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:36:15.540187Z digest=sha256:8477bfe5b6419371debc610db0870b6c99201ad5134000ac6b7c30322927586c

Observation 9675d94b-f65d-4fbb-a217-82b0059f9647 · inbound

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos cites this paper.

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:22:55.838095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:22:55.838095Z digest=sha256:48e6aaf7a300217f6dcb1afe71bb22f5d18dde133cd03b9406a9166b88ca3673

Observation 90ecb6a8-ce9f-4510-9908-270fc34ff20f · inbound

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? cites this paper.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:32.061606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:32.061606Z digest=sha256:c9cd864c14b0a38cf5b0e0a5ce75acce00351c9104cf48d37479dfba52f25bfd

Observation 480c83b2-d71e-49d9-a101-d9b45129049b · inbound

Argus Inspection: Do Multimodal Large Language Models Possess the Eye of Panoptes? cites this paper.

Argus Inspection: Do Multimodal Large Language Models Possess the Eye of Panoptes? Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:18:57.344627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:18:57.344627Z digest=sha256:6e5e464f7eab464724b10a1561203bea30607a5bf8122e1a548ff9a7ea12b01d

Observation 2ce8fd4d-5cef-4121-8368-d09de60cdb0a · inbound

DAVID-XR1: Detecting AI-Generated Videos with Explainable Reasoning cites this paper.

DAVID-XR1: Detecting AI-Generated Videos with Explainable Reasoning Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:53.934003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:53.934003Z digest=sha256:d14022c00dbc846a4e42681b2e86fa0c1d44d6ec365c8a65b789d6a5cafca749

Observation 01f3cec6-94bc-4467-b9a8-00f90c6b63fb · inbound

CausalStep: A Benchmark for Explicit Stepwise Causal Reasoning in Videos cites this paper.

CausalStep: A Benchmark for Explicit Stepwise Causal Reasoning in Videos Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T15:14:17.320440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:14:17.320440Z digest=sha256:f53b90312d2e2e93a14ffff503dcac5c15a0d5711459203b91cd1bcf3b872c16

Observation 4359e34a-b13d-4732-aa9a-1f327d90bb74 · inbound

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey cites this paper.

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition

Reference 155

Resolution
unresolved
no resolver link, observed 2026-08-05T20:28:57.362974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:28:57.362974Z digest=sha256:f9be84341b5059d1036ed84be07388fe0a68f5ab9b3210321cc9e1ee5f81124f

Observation f57066a7-b0b7-4587-9c3a-b9919b65d70d · inbound

ProPy: Building Interactive Prompt Pyramids upon CLIP for Partially Relevant Video Retrieval cites this paper.

ProPy: Building Interactive Prompt Pyramids upon CLIP for Partially Relevant Video Retrieval Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T16:05:18.612817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:05:18.612817Z digest=sha256:8682be71d6c6f409b08c481a929923b83530f9fcce851e93cea49a1e6ad158b9

Observation 06ccf859-2cfb-4de3-bcfe-fe36e36c6fe1 · inbound

AdsQA: Towards Advertisement Video Understanding cites this paper.

AdsQA: Towards Advertisement Video Understanding Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T20:20:36.704406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:20:36.704406Z digest=sha256:816bf666cb1b9ed6697d374e9998fd7f307ff8b9fd7162be820a5891aafa2b39

Observation 0cab05a2-37db-4e22-9ab5-6e1e9e8df3d7 · inbound

LaRe: Latent Refocusing for Multimodal Reasoning cites this paper.

LaRe: Latent Refocusing for Multimodal Reasoning Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T00:15:15.354915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:15:15.354915Z digest=sha256:a5561edc58a25c2c95a67c18da64e30657effb985104c5b76c5523d8f35eaab9

Observation ffab02f4-3f31-4207-9b4f-1da7cea574b9 · inbound

SCP: Spatial Causal Prediction in Video cites this paper.

SCP: Spatial Causal Prediction in Video Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:50:11.257532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T16:47:44.523606Z digest=sha256:2af28511eb4a6d92c0208af737c2658f2855a67e48b8175490b4cfba2c43e074

Observation 0fdad807-aa9f-4112-b2a9-f42f8e0db055 · inbound

EndoCoT: Scaling Endogenous Chain-of-Thought Reasoning in Diffusion Models cites this paper.

EndoCoT: Scaling Endogenous Chain-of-Thought Reasoning in Diffusion Models Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T22:25:04.080629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T22:25:04.080629Z digest=sha256:c90fd78c4d11a57e14e686fa661637d14dde77ba7bc51f8af66ed373f76f113c

Observation 0d6f5247-243e-4e29-b227-e69e78685e10 · inbound

Reasoning Resides in Layers: Restoring Temporal Reasoning in Video-Language Models with Layer-Selective Merging cites this paper.

Reasoning Resides in Layers: Restoring Temporal Reasoning in Video-Language Models with Layer-Selective Merging Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:56:05.690356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:42:43.948462Z digest=sha256:201cb15646832aa58e52742b14ac28e553c7cd2b0c43280f8df8b796a8637beb

Observation 6d566b1f-c33e-478e-be43-1413a3faa1dd · inbound

Chain-of-Glimpse: Search-Guided Progressive Object-Grounded Reasoning for Video Understanding cites this paper.

Chain-of-Glimpse: Search-Guided Progressive Object-Grounded Reasoning for Video Understanding Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:05:22.155788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T12:03:09.408019Z digest=sha256:a81d6893cbb374915d098105ca37c7a066a593bf827fd120c682584a4a81f10d

Observation e74bd42d-a812-4456-a4f0-cd6f30bd7042 · inbound

Chain-of-Glimpse: Search-Guided Progressive Object-Grounded Reasoning for Video Understanding cites this paper.

Chain-of-Glimpse: Search-Guided Progressive Object-Grounded Reasoning for Video Understanding Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:37:41.741853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T17:34:10.344111Z digest=sha256:98c9b69512a12fdf99c21a8802fe6dafddbe4a743f72d8c20563debb61e065cf

Observation ab8ee26b-d76d-4004-a8de-38e451d00b96 · inbound

When Text Hijacks Vision: Benchmarking and Mitigating Text Overlay-Induced Hallucination in Vision Language Models cites this paper.

When Text Hijacks Vision: Benchmarking and Mitigating Text Overlay-Induced Hallucination in Vision Language Models Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:46:37.539454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T06:41:59.641410Z digest=sha256:222d7081c654c3693ebd844c3e0ee5067e17c6a04609cdf0b12de06b9f91db7b

Observation aecfa82f-bb1c-4e17-b1ea-9835a388b5e2 · inbound

Act2See: Emergent Active Visual Perception for Video Reasoning cites this paper.

Act2See: Emergent Active Visual Perception for Video Reasoning Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-09T05:45:22.936174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T19:34:53.683729Z digest=sha256:d58d4960ab3ff368e1db9dac2b1ae0b48deeaa43a674ae83d60248e9ec6ae14e

Observation 09ead882-9a5b-48be-b467-633f37849428 · inbound

ViSRA: A Video-based Spatial Reasoning Agent for Multi-modal Large Language Models cites this paper.

ViSRA: A Video-based Spatial Reasoning Agent for Multi-modal Large Language Models Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:46:35.770004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T04:00:23.681682Z digest=sha256:f23eac1bcf28365f6b1d1f8531d8c7dce24afe708ccc63dbc81a429e510dadcc

Observation bbb98392-7b8e-4250-bac0-197f37dd3781 · inbound

AVTrack: Audio-Visual Tracking in Human-centric Complex Scenes cites this paper.

AVTrack: Audio-Visual Tracking in Human-centric Complex Scenes Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition

Reference 74

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:46:20.158878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T15:00:41.632699Z digest=sha256:1aea0c55e42fff8e60b9033974feef2f9df088483f36890f04d623f845d6ac63

Observation db96b895-cc57-46af-b777-73f535fa5fc5 · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition

Reference 178

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.742811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:8e9e5fb51ce4a0ae00047d405057a5e8825c5d2de0b27ffa0e764dabaa579961

Observation d3fa565a-c03c-4bc6-82f0-7c6b65c5f10a · inbound

Counterfactual Reasoning for Fine-Grained Evidence Disentanglement in VideoQA cites this paper.

Counterfactual Reasoning for Fine-Grained Evidence Disentanglement in VideoQA Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:57:30.197957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T16:55:35.743040Z digest=sha256:628cbc823a547cbd70e279d2182d93b67ffec4237434f2043fdc382875d9ec07

Observation cce4ce6e-70f5-48c3-bd5d-1e146ed7bfbf · inbound

Plan-and-Verify Video Reward Reasoning with Spatio-Temporal Scene Graph Grounding cites this paper.

Plan-and-Verify Video Reward Reasoning with Spatio-Temporal Scene Graph Grounding Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T08:57:47.800558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T10:37:52.186246Z digest=sha256:6c8941a55f3a9f159c3347200545d80186ea31498d9e2b0b69433b0e14bebe91