Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T19:06:46.257397Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 6 inbound Pith citation observations for arXiv:2501.10674.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T19:06:46.257397Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:41:35.841271Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
27 of 27 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 7dd80212-3df7-48f8-86d5-d033f731aa3d · outbound
Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! online" 'onlinestring :=
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba5d8bb1-90b7-4a7c-a59b-bbf1ec772c79 · outbound
Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! write newline
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2298c8bd-d5e7-41cc-afe2-5efbd23540fa · outbound
Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43e812f6-e863-40dd-8404-0d52591f11a0 · outbound
Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46023000-6df2-4999-a19b-08e5156b50fc · outbound
Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15d73432-2ef0-47f1-99e4-6a83f7777c63 · outbound
Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! Unresolved cited work
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f39631f1-dd69-4503-b0e7-984f504b983c · outbound
Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! Lost in Time: A New Temporal Benchmark for VideoLLMs
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78497f42-d7e3-4fe4-a8aa-34f3928f4c1b · outbound
Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! The Llama 3 Herd of Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2c53ff8-b438-4eed-9498-3c184a07db16 · outbound
Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f214ae92-a8f1-4149-8d5c-b2e571e5122e · outbound
Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c60dc037-365f-481d-8ed1-333763189b6a · outbound
Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11da3a5a-f5f3-4f84-a7e9-c6d5d000fa73 · outbound
Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! Unresolved cited work
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e1cb9d6-6066-465c-97fb-6a20d02cbeac · outbound
Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! A Survey on Evaluation of Multimodal Large Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eba48bad-0095-41da-9435-6070b2508854 · outbound
Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! GPT-4o System Card
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84c0c77f-f77a-4e52-a589-1c80fa26f14d · outbound
Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19f2362d-041d-4999-a00e-b8d154881caa · outbound
Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! A Survey on Benchmarks of Multimodal Large Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bc3818a-b01c-4abd-b3bf-96e07c1b8ff1 · outbound
Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! Unresolved cited work
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a90468de-f4ed-4474-a62d-bb595c051299 · outbound
Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! Unresolved cited work
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef8efdad-a3f1-460d-bb10-a93245101439 · outbound
Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! MIBench: Evaluating Multimodal Large Language Models over Multiple Images
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b73cabd9-3ec1-4fdf-ac0f-828b168ce800 · outbound
Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06148cbe-1720-43c6-a6d6-afef2290c80f · outbound
Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! Temporal Insight Enhancement: Mitigating Temporal Hallucination in Multimodal Large Language Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 672ff67a-4355-480b-b342-e31fa8cf8e42 · outbound
Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! Gemini: A Family of Highly Capable Multimodal Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71c16368-7823-4976-8384-197af39a3626 · outbound
Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! Unresolved cited work
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c0e8e8b-75df-41a1-965b-74ead9564c02 · outbound
Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! Unresolved cited work
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2760568-7eca-4b4f-bd7a-96a93492d544 · outbound
Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! LLaVA-CoT: Let Vision Language Models Reason Step-by-Step
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5d02360-6293-4877-a109-8163854aad09 · outbound
Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! A Survey on Multimodal Large Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3d9fc16-207f-48c3-8879-667deed439f1 · outbound
Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3718325-5de1-4fc0-92c6-5c98b9263dea · inbound
Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No!
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6039192c-0f0b-445f-9fff-c3c2200c1846 · inbound
Pluri-perspectivism in Human-robot Co-creativity with Older Adults Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No!
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11c15949-1f24-437e-bb22-6a43f81c0ec0 · inbound
Lost in the Vibrations: Vision Language Models Fail the Dynamic Gauges Test Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No!
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d7480da3-8d0c-4b46-9d8b-5d4996d25f45 · inbound
OralMLLM-Bench: Evaluating Cognitive Capabilities of Multimodal Large Language Models in Dental Practice Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No!
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b9359344-7ff2-41b1-a7bf-0e96f2dc3622 · inbound
Seeing Time: Benchmarking Chronological Reasoning and Shortcut Biases in Vision-Language Models Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No!
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b98cd653-213b-47b3-b2ea-f2fa067e1798 · inbound
Video-MME-Logical: A Controlled Diagnostic Benchmark for Video Temporal-Logical Reasoning Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No!
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.