Pith. sign in

Paper Citation Record · LEDGER

Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No!

As of 11 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 6 inbound Pith citation observations for arXiv:2501.10674.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.10674 v2

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T19:06:46.257397Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:41:35.841271Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

27 of 27 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 7dd80212-3df7-48f8-86d5-d033f731aa3d · outbound

This paper cites online" 'onlinestring :=.

Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T19:06:46.133639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:06:46.133639Z digest=sha256:1295023f3c45a83ec778b9bfe176529b22673f42f4bb087925e6830a832dae99

Observation ba5d8bb1-90b7-4a7c-a59b-bbf1ec772c79 · outbound

This paper cites write newline.

Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T19:06:46.138060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:06:46.138060Z digest=sha256:63cff0743881f6be944b6aec67f555e08567d40559c88c61b3f343da656a7dc7

Observation 2298c8bd-d5e7-41cc-afe2-5efbd23540fa · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T19:06:46.143202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:06:46.143202Z digest=sha256:61273aeb811edfa45f80a3ea2f96d18d3be1ed6d2fad79274709bbfcd399467b

Observation 43e812f6-e863-40dd-8404-0d52591f11a0 · outbound

This paper cites TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models.

Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T19:06:46.147800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:06:46.147800Z digest=sha256:084f3ca5af49369beb6d220ff6d3d2ffb85f1058428276c5fcaa899e95b0a616

Observation 46023000-6df2-4999-a19b-08e5156b50fc · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T19:06:46.152258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:06:46.152258Z digest=sha256:42532f707fb4a8699239a4b06bf22f5f7df28ede799aafb9821b6ef78f7ba1db

Observation 15d73432-2ef0-47f1-99e4-6a83f7777c63 · outbound

This paper cites an unresolved cited work.

Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T19:06:46.157210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:06:46.157210Z digest=sha256:c9bef8a4a7051b4a5f4240dcf6663d6b6f38e2cb48983fcedf5ad484646819ff

Observation f39631f1-dd69-4503-b0e7-984f504b983c · outbound

This paper cites Lost in Time: A New Temporal Benchmark for VideoLLMs.

Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! Lost in Time: A New Temporal Benchmark for VideoLLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T19:06:46.161997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:06:46.161997Z digest=sha256:994d9713fba6d78fe41d1ba9147aec3f77d9010904df88b69d27e6a03b39a834

Observation 78497f42-d7e3-4fe4-a8aa-34f3928f4c1b · outbound

This paper cites The Llama 3 Herd of Models.

Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! The Llama 3 Herd of Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T19:06:46.166569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:06:46.166569Z digest=sha256:278347b1e4c003aa80c018a4c6a3e4febb21678c477a5de7b2f3572e3473cab3

Observation a2c53ff8-b438-4eed-9498-3c184a07db16 · outbound

This paper cites Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning.

Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T19:06:46.171743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:06:46.171743Z digest=sha256:87dcdf08d678138717832e748c5a1c800ea00513963f027e2adb876650aafe47

Observation f214ae92-a8f1-4149-8d5c-b2e571e5122e · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T19:06:46.176421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:06:46.176421Z digest=sha256:f4e4c4ad301570d2f0de0cc566266650247c19fd34cd8371eef0629f87caefa8

Observation c60dc037-365f-481d-8ed1-333763189b6a · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T19:06:46.181220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:06:46.181220Z digest=sha256:5e6c86e7cf55cb0ec967fd759049695678ec4c0ed81a9c63d952634fa233bf47

Observation 11da3a5a-f5f3-4f84-a7e9-c6d5d000fa73 · outbound

This paper cites an unresolved cited work.

Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T19:06:46.190545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:06:46.190545Z digest=sha256:b6f43903f562375266c04c7f3f992d8cb8bc3433644153bae36e72fc3f74239d

Observation 7e1cb9d6-6066-465c-97fb-6a20d02cbeac · outbound

This paper cites A Survey on Evaluation of Multimodal Large Language Models.

Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! A Survey on Evaluation of Multimodal Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T19:06:46.194763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:06:46.194763Z digest=sha256:0cc89746ec5019bdb6a6b5848e556f9b2f27d6fc3efc2da2cfedcb8cc6af12e7

Observation eba48bad-0095-41da-9435-6070b2508854 · outbound

This paper cites GPT-4o System Card.

Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! GPT-4o System Card

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T19:06:46.199672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:06:46.199672Z digest=sha256:d1e36968d02d55d936a895dd1a93c25f9f75e36c33667fc3c88aa97bc53e45ba

Observation 84c0c77f-f77a-4e52-a589-1c80fa26f14d · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T19:06:46.203899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:06:46.203899Z digest=sha256:e592abb3193662406423d71e84f9df9fe5869e5829a551413c198adb5ff2731d

Observation 19f2362d-041d-4999-a00e-b8d154881caa · outbound

This paper cites A Survey on Benchmarks of Multimodal Large Language Models.

Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! A Survey on Benchmarks of Multimodal Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T19:06:46.207735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:06:46.207735Z digest=sha256:694680a03b66bebad05a1d980992e18935e693883ee2fd8d6a10023a9f8b15bd

Observation 8bc3818a-b01c-4abd-b3bf-96e07c1b8ff1 · outbound

This paper cites an unresolved cited work.

Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T19:06:46.211547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:06:46.211547Z digest=sha256:791b432e68607cb06834e93320e97723b2c2b3539586bce9d4d0d674673a6e04

Observation a90468de-f4ed-4474-a62d-bb595c051299 · outbound

This paper cites an unresolved cited work.

Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T19:06:46.215735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:06:46.215735Z digest=sha256:62e9c5b0b5944f191b18a355763af0b25249ce4d7f44ac92070422cd32c96802

Observation ef8efdad-a3f1-460d-bb10-a93245101439 · outbound

This paper cites MIBench: Evaluating Multimodal Large Language Models over Multiple Images.

Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! MIBench: Evaluating Multimodal Large Language Models over Multiple Images

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T19:06:46.219688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:06:46.219688Z digest=sha256:7483d7b34d002fabf2ba56a38bec62a8f2edb5bc863b49bf9d77fec0cf898c67

Observation b73cabd9-3ec1-4fdf-ac0f-828b168ce800 · outbound

This paper cites TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models.

Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T19:06:46.226297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:06:46.226297Z digest=sha256:538e11ee490bcf9900cc1be5e4298df9f4ca678ad1bc8a5bc02c1ac5749a8938

Observation 06148cbe-1720-43c6-a6d6-afef2290c80f · outbound

This paper cites Temporal Insight Enhancement: Mitigating Temporal Hallucination in Multimodal Large Language Models.

Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! Temporal Insight Enhancement: Mitigating Temporal Hallucination in Multimodal Large Language Models

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-10T19:06:46.507260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T19:06:46.230495Z digest=sha256:4d9effd602fc3c9fb830a32324df0c4729c9d76362b4d1f1d6e297b46f88b9ea

Observation 672ff67a-4355-480b-b342-e31fa8cf8e42 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! Gemini: A Family of Highly Capable Multimodal Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T19:06:46.234287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:06:46.234287Z digest=sha256:01d566c02a6dd65f332be55b10f2e91f6f33f9eb631770df8c243053cb37d1ae

Observation 71c16368-7823-4976-8384-197af39a3626 · outbound

This paper cites an unresolved cited work.

Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T19:06:46.238897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:06:46.238897Z digest=sha256:b11ce5290efc239ec8122f59540a881c246dd8f745b9d785215a33ea50956057

Observation 7c0e8e8b-75df-41a1-965b-74ead9564c02 · outbound

This paper cites an unresolved cited work.

Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T19:06:46.243385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:06:46.243385Z digest=sha256:d8f8c27af9075e8fddf4de70a882f231bc9566756fd275c4dfe508ee70598f82

Observation e2760568-7eca-4b4f-bd7a-96a93492d544 · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T19:06:46.247965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:06:46.247965Z digest=sha256:9a5a021ff78cfd1e52126769e923a4f1f1c91a9814dd63236f31142c42c1c65f

Observation d5d02360-6293-4877-a109-8163854aad09 · outbound

This paper cites A Survey on Multimodal Large Language Models.

Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! A Survey on Multimodal Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T19:06:46.252570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:06:46.252570Z digest=sha256:c4bde4e20d327b2996d3272e344b60cbabf1ef41ea01978039226eeaa99f2f2c

Observation f3d9fc16-207f-48c3-8879-667deed439f1 · outbound

This paper cites MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI.

Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T19:06:46.257397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:06:46.257397Z digest=sha256:932e7f83c7e16e85b693e26e4493d105b75eb00a8029dbf09b3c1cc89e1e99da

Pith citing papers

Observation d3718325-5de1-4fc0-92c6-5c98b9263dea · inbound

Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs cites this paper.

Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No!

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:07:15.464858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T11:03:59.222849Z digest=sha256:b75ac9cdc541693fb0522e145f56e1cfb243a614c955e53b731cca4fbee24532

Observation 6039192c-0f0b-445f-9fff-c3c2200c1846 · inbound

Pluri-perspectivism in Human-robot Co-creativity with Older Adults cites this paper.

Pluri-perspectivism in Human-robot Co-creativity with Older Adults Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No!

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T18:41:35.841271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:41:35.841271Z digest=sha256:8f2b35c086c9207dc97a13e972d602e003562dad1f3476a422c5bf134611350d

Observation 11c15949-1f24-437e-bb22-6a43f81c0ec0 · inbound

Lost in the Vibrations: Vision Language Models Fail the Dynamic Gauges Test cites this paper.

Lost in the Vibrations: Vision Language Models Fail the Dynamic Gauges Test Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No!

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:21:26.657532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T06:21:17.509796Z digest=sha256:167f8fb03164a94830f25ed80b6cbbce2feb9eb31a95dc4a42ade3d7d78950f1

Observation d7480da3-8d0c-4b46-9d8b-5d4996d25f45 · inbound

OralMLLM-Bench: Evaluating Cognitive Capabilities of Multimodal Large Language Models in Dental Practice cites this paper.

OralMLLM-Bench: Evaluating Cognitive Capabilities of Multimodal Large Language Models in Dental Practice Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No!

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:05:50.282084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T01:03:50.055081Z digest=sha256:2606a02644456e6ffb2f8aac28375c128bf4391d16b4a4cd084ec192d0a2ece9

Observation b9359344-7ff2-41b1-a7bf-0e96f2dc3622 · inbound

Seeing Time: Benchmarking Chronological Reasoning and Shortcut Biases in Vision-Language Models cites this paper.

Seeing Time: Benchmarking Chronological Reasoning and Shortcut Biases in Vision-Language Models Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No!

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:06:59.623227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T01:30:36.491137Z digest=sha256:4fae12f99f9ca259461ef1c8369226b28d9f6538900bfd153d1b994d57f0c211

Observation b98cd653-213b-47b3-b2ea-f2fa067e1798 · inbound

Video-MME-Logical: A Controlled Diagnostic Benchmark for Video Temporal-Logical Reasoning cites this paper.

Video-MME-Logical: A Controlled Diagnostic Benchmark for Video Temporal-Logical Reasoning Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No!

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T19:13:52.947250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-29T04:53:25.259840Z digest=sha256:dacfc19386253150bd37b5487ba5ea0b781c6e564ab5c56a1214c3c3b88b11d4