Pith. sign in

Paper Citation Record · LEDGER

TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2410.23266.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.23266 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T19:06:46.226297Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T05:57:41.270398Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b73cabd9-3ec1-4fdf-ac0f-828b168ce800 · inbound

Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! cites this paper.

Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T19:06:46.226297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:06:46.226297Z digest=sha256:96677b3698cb4e35400ba42dd7c6274dd74c1905cffedbfa7d9f2834dcbe324f

Observation 32543ac6-b29e-42a2-954b-c83eec350e2b · inbound

Seed1.5-VL Technical Report cites this paper.

Seed1.5-VL Technical Report TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models

Reference 118

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:05.439460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:20870a12f31cbb98dbcd71f245730ae1f5053347c09e35e32e06dce70a8da888

Observation ebbce05a-4324-48b2-8795-b111606f7354 · inbound

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos cites this paper.

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:03.750220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:03:03.750220Z digest=sha256:31620cb704812283b6a40f86f7ec7f855ddc1055bb2a827f61bfafd8fb98216b

Observation 51d896f8-66df-4027-9c6e-e48beb66ab3b · inbound

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos cites this paper.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:48.261199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:48.261199Z digest=sha256:e0ca8c4e8fb26344dc9fe2440df60768411fedd0a181133b41a22fd4747a3df8

Observation 1084599e-7509-4f0e-a156-9c5c0b60a8f8 · inbound

VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos cites this paper.

VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T10:25:36.299124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:25:36.299124Z digest=sha256:243efbbebe657233c7441d67ebb139f4d50ac74a57ff3fcde89236a415c111d9

Observation dd1c37b8-704b-4224-888d-0992d7313db4 · inbound

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning cites this paper.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:33:51.178602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:f8a6b84d0bf714696710121350f19925a46c36a467323bb79885c4fa65751c71

Observation ffdbdba9-ab23-41ab-b009-7575a0a0e9f7 · inbound

ISO-Bench: Benchmarking Multimodal Causal Reasoning in Visual-Language Models through Procedural Plans cites this paper.

ISO-Bench: Benchmarking Multimodal Causal Reasoning in Visual-Language Models through Procedural Plans TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T11:08:10.355570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:08:10.355570Z digest=sha256:6a2a9b1e802998975826c00b225219633f1ea3a15ace73e63536619c10b2147f

Observation e6b5b6e6-29d1-4b52-93ca-c5210d58cc63 · inbound

SVCBench: A Streaming Video Counting Benchmark for Spatial-Temporal State Maintenance cites this paper.

SVCBench: A Streaming Video Counting Benchmark for Spatial-Temporal State Maintenance TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-14T22:09:56.270516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T22:09:56.270516Z digest=sha256:ee9a22d16c6f41484455faaf451f51ba54de908a51ba5e8a72e0655d3cce623b

Observation 6adaaf95-1685-46ff-8bf3-bcd8f52a8fdb · inbound

Reasoning Resides in Layers: Restoring Temporal Reasoning in Video-Language Models with Layer-Selective Merging cites this paper.

Reasoning Resides in Layers: Restoring Temporal Reasoning in Video-Language Models with Layer-Selective Merging TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:56:05.675745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T15:42:43.948462Z digest=sha256:bf5c1f772504875cb7edf3fe00780404d4f06c0b1509fbc5454b92caa165b0e5

Observation cd74b951-ddc4-4b90-84ca-f9892bb792f2 · inbound

DenseStep2M: A Scalable, Training-Free Pipeline for Dense Instructional Video Annotation cites this paper.

DenseStep2M: A Scalable, Training-Free Pipeline for Dense Instructional Video Annotation TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:46:26.115404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-07T13:50:23.835653Z digest=sha256:85fe309da78f26b746ce8ee2f099bc2776b3f332caf439194a68587ffc925640

Observation c71e0dfd-8391-4f1b-9d1b-0a901beee377 · inbound

VLMaxxing through FrameMogging Training-Free Anti-Recomputation for Video Vision-Language Models cites this paper.

VLMaxxing through FrameMogging Training-Free Anti-Recomputation for Video Vision-Language Models TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:11:13.608173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T01:30:15.463051Z digest=sha256:214795af5a775cfcb7af9c0bcd694cdb1d887054997c97cf997b8c6b31865198

Observation 43914f2a-3cf6-4734-a493-640c3ec42e7f · inbound

EvoGround: Self-Evolving Video Agents for Video Temporal Grounding cites this paper.

EvoGround: Self-Evolving Video Agents for Video Temporal Grounding TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:32:52.309602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-14T19:29:47.356665Z digest=sha256:6e708660397899901e0974a069acb33f618e70fc0e919b728b1557deecd38388

Observation 6bb766fe-00df-468e-b65c-38e645fc268b · inbound

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes cites this paper.

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models

Reference 205

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:57:41.272017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-27T12:59:51.091008Z digest=sha256:09ee76dacd3616ac90dbebf0f1622800c341b6822ad674ccf3b37191bf9a8209