Pith. sign in

Paper Citation Record · LEDGER

TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 41 inbound Pith citation observations for arXiv:2410.10818.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.10818 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 41 of 41 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:24:07.209929Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T23:39:03.234328Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d5b6241f-c506-4dde-888a-5ebddaa3d455 · inbound

LLaVA-Video: Video Instruction Tuning With Synthetic Data cites this paper.

LLaVA-Video: Video Instruction Tuning With Synthetic Data TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 216

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:20:33.170271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:d43e9e61f27dfdbe5449b7af4c873497b2a4ae39f1dc6fdd4a755f58f2bec471

Observation 8728683d-c1dd-4cae-8913-cb685f76a829 · inbound

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos cites this paper.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:32:41.302744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:c891b9f780d6c999d7157ab8e4483c36faed42cd3dd3d81db7ccf3a892a380df

Observation 4d7bc3d6-cca5-45e2-9860-084fb8cfc5d9 · inbound

Seed1.5-VL Technical Report cites this paper.

Seed1.5-VL Technical Report TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:05.681583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:fcdeb9b8edd338156ff686708d52c36bcd445ea5ceaf94ca2106ca74b7f324fb

Observation a6bc6d7e-b309-427f-aa28-fe817f779cc3 · inbound

RTime-QA: A Benchmark for Atomic Temporal Event Understanding in Large Multi-modal Models cites this paper.

RTime-QA: A Benchmark for Atomic Temporal Event Understanding in Large Multi-modal Models TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:07.209929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:24:07.209929Z digest=sha256:ab00dede70f7719ffd482dec8eda1b2b23e84a219a5905a17a19b9105ff90738

Observation 0068b704-aecd-46cc-a8c5-175b3944aa05 · inbound

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos cites this paper.

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:59.572640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:02:59.572640Z digest=sha256:fa44fd29234465395b697282cba8f8370ae25ead867150f496f06154c7619f48

Observation 358b7850-828e-4e44-905f-3a296da34640 · inbound

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding cites this paper.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:17:18.545824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:1d09f990f87675c698358600bb8cf65c97bcbdd09bb550e48cd035261a7440a1

Observation 8c97c53f-09c4-49cc-a2be-eebe6f273efc · inbound

Fostering Video Reasoning via Next-Event Prediction cites this paper.

Fostering Video Reasoning via Next-Event Prediction TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:40.431117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:40.431117Z digest=sha256:0e7ff06564237f6aa751ddbf696078d6dda08a6887985b1a52f2de54d8069f8f

Observation 305f6973-6c73-49c7-a805-d19b14918540 · inbound

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times cites this paper.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.121411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.121411Z digest=sha256:4630abfde5a40b1f905560c6e4eb7f996b13c4cf104def89a76a4eeadac7408a

Observation 0a22fff6-cf20-4c15-89bc-b46ffd20a18e · inbound

Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs cites this paper.

Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:07:15.380828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T11:03:59.222849Z digest=sha256:60128602e8fd68c89e1b8906b5cd7f74d5a9c2adaed202faf84fb09f798c8ba7

Observation b3e7157c-7b90-4817-989b-f5200612b389 · inbound

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning cites this paper.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:33:50.948202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:dfb0885c5c781e91ea1a079af05691515a709fc09a54847b7741d7163ae627ee

Observation e849daa4-bf69-475c-8e19-353675a6c28c · inbound

ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models cites this paper.

ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:13.972207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:03:13.972207Z digest=sha256:03ce750c9a0075b0bbd6cabbf4c2a8f3045f89ef6eaed78f1fa29105b69bf799

Observation bbe643d1-35b2-4d2e-9e1a-e0b64b67fc02 · inbound

GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? cites this paper.

GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T17:58:26.999928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:58:26.999928Z digest=sha256:4d51d9e8eaad3c3a5889e83a34194e15a4aed76ab2c4aa785d4188e9c1c62aeb

Observation f91d00e8-c6f6-491f-b1a5-7bdc970597a8 · inbound

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding cites this paper.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:31.035101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:31.035101Z digest=sha256:a855547acadce70eb2123ccf7f7ccd17f9937d000bb3a28f60894c22e03c7ba9

Observation ad492195-bd08-400d-a93c-a65b26054333 · inbound

CausalStep: A Benchmark for Explicit Stepwise Causal Reasoning in Videos cites this paper.

CausalStep: A Benchmark for Explicit Stepwise Causal Reasoning in Videos TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T15:14:17.287285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:14:17.287285Z digest=sha256:5f4c63af8bf5dc2bb8e6b1fccc44801a857fd339744718b11b32c4abb06a3030

Observation 15f9e0ce-387e-4e56-a998-6d9afa19ef88 · inbound

AdsQA: Towards Advertisement Video Understanding cites this paper.

AdsQA: Towards Advertisement Video Understanding TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T20:20:36.661579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:20:36.661579Z digest=sha256:5f8d844c78075881f44af32c27c3271d364ccec62f238f89b593a92db99a5fa8

Observation af06448f-fba1-4f4c-bcbd-94a2990c6610 · inbound

NeMo: Needle in a Montage for Video-Language Understanding cites this paper.

NeMo: Needle in a Montage for Video-Language Understanding TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T13:54:11.277795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:54:11.277795Z digest=sha256:f9732905cc421af8f6dc41a113c701502ae3add553f7ba623cc766f960ab74cf

Observation 6a8aaab0-b1f6-4fde-bec6-dbbb05df58b9 · inbound

LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling cites this paper.

LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-22T12:31:32.156486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T12:26:35.347190Z digest=sha256:ee1f9112377185461f5f8191f14cbb5ab1183401e7fd289238ac652d1dcd2a0e

Observation d12b6c69-c134-451c-910b-20cabb957091 · inbound

See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models cites this paper.

See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:13:52.103038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T02:12:55.170296Z digest=sha256:85cd168686c5f390438c8760dbd5731ab6ba80c50ad4076085f12ffa6543cd7e

Observation c84d5d53-66bb-46aa-8f21-429c8c812f36 · inbound

From Segments to Scenes: Temporal Understanding for Agentic Autonomous Driving via Vision-Language Models cites this paper.

From Segments to Scenes: Temporal Understanding for Agentic Autonomous Driving via Vision-Language Models TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T18:28:15.922879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:28:15.922879Z digest=sha256:bcff2fa9011c0e6b0df688afb9b6394e89c208010ae93539aa0a8e58cd64ea5c

Observation f6290141-24a4-455e-b8f9-c5e6fe3d58e8 · inbound

GTASA: Ground Truth Annotations for Spatiotemporal Analysis, Evaluation and Training of Video Models cites this paper.

GTASA: Ground Truth Annotations for Spatiotemporal Analysis, Evaluation and Training of Video Models TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:26:01.762020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T16:38:10.956726Z digest=sha256:76cbae10c96631fc8be6e9e9ada267d6db38ca1b67361b678a9e9bb9d93513ff

Observation 6968e55c-af11-43d4-910d-b458eb5ba282 · inbound

POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs cites this paper.

POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:41:04.015830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T15:23:08.671342Z digest=sha256:842a0aee20588131b06546b225e303c8786df38ed94cd8571104b41d20a221c2

Observation 4d934730-f911-4785-8ed2-1e50700138de · inbound

When Text Hijacks Vision: Benchmarking and Mitigating Text Overlay-Induced Hallucination in Vision Language Models cites this paper.

When Text Hijacks Vision: Benchmarking and Mitigating Text Overlay-Induced Hallucination in Vision Language Models TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:46:37.546797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T06:41:59.641410Z digest=sha256:07b0656da2956799c3fab2c059ab58ecbba1e0e6c091545e21d00bd800ca5630

Observation ebfb5115-25aa-42a9-8026-15d9f12fafa5 · inbound

VideoNet: A Large-Scale Dataset for Domain-Specific Action Recognition cites this paper.

VideoNet: A Large-Scale Dataset for Domain-Specific Action Recognition TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:20:41.264335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T18:35:48.379198Z digest=sha256:3a93ddb80a7dd0c64d23a0d9ba073dd486e0b1491aeec5d7919397c4b44616bf

Observation 83c7d2dd-2323-4155-b984-2468ff757f85 · inbound

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models cites this paper.

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:31:26.606021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T04:13:21.487431Z digest=sha256:5f0231cea9d2653dcbe52ae9fde6965b3432d8b24d5f996267bc3fc23289d8f5

Observation 7d3ca994-57b8-40e7-917c-a5a53953eb14 · inbound

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models cites this paper.

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:57:28.298673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T06:53:42.726350Z digest=sha256:26de72f9fd3d5d123f28a8db8a0f2cb736baa8aa07837495cacddfc60c55c4e2

Observation 7559d5ad-eb08-48c6-aa60-ead607b633ac · inbound

EvoGround: Self-Evolving Video Agents for Video Temporal Grounding cites this paper.

EvoGround: Self-Evolving Video Agents for Video Temporal Grounding TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:32:52.564646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T19:29:47.356665Z digest=sha256:65418e57956dc8430378033017748861a6afb9855c6c08ffe2c24a0267a337a7

Observation f1700d39-3c5f-45fb-8f20-19724b4ed41a · inbound

Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding cites this paper.

Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:03:08.044136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T16:02:53.887605Z digest=sha256:a5ddf45d9651450a58bf8b7a59f5b0996675bdc03bd419e55ed545a348677ede

Observation aaa66853-e66b-4ffb-8c41-8f8517e6def5 · inbound

FineBench: Benchmarking and Enhancing Vision-Language Models for Fine-grained Human Activity Understanding cites this paper.

FineBench: Benchmarking and Enhancing Vision-Language Models for Fine-grained Human Activity Understanding TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:13:05.462877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T06:08:12.887394Z digest=sha256:42b5e607c710134995476044ddbb60e014502ca784396af2dbe315c955cd20ee

Observation a057f476-9899-436f-91f4-48608cfcacc4 · inbound

FineBench: Benchmarking and Enhancing Vision-Language Models for Fine-grained Human Activity Understanding cites this paper.

FineBench: Benchmarking and Enhancing Vision-Language Models for Fine-grained Human Activity Understanding TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:04:02.791065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T07:59:57.488107Z digest=sha256:eb7f57521b0922819a3bdb34f3c69d3aaaed8d7d2e97b26a6558cea22c7bbb9b

Observation 165ec764-c51b-4b13-8e50-4780a9e9f8e9 · inbound

FineBench: Benchmarking and Enhancing Vision-Language Models for Fine-grained Human Activity Understanding cites this paper.

FineBench: Benchmarking and Enhancing Vision-Language Models for Fine-grained Human Activity Understanding TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:14:59.999305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T18:09:49.178090Z digest=sha256:720b8de8b18e4720fe01f3a0b5417560f4b6f85e760771c6db079ec0d2e4b601

Observation e6080247-df3f-4d8b-944b-8736e0d26d6a · inbound

Learning Spatiotemporal Sensitivity in Video LLMs via Counterfactual Reinforcement Learning cites this paper.

Learning Spatiotemporal Sensitivity in Video LLMs via Counterfactual Reinforcement Learning TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:46:14.880613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T07:45:56.473188Z digest=sha256:5db0804e41b3e986888f6ec860cf1114ddff90d938b20322be56ff58b3a58217

Observation 9aca9f84-e2ee-4eb9-a074-cb6fe0b21f5e · inbound

The TIME Machine: On The Power of Motion for Efficient Perception cites this paper.

The TIME Machine: On The Power of Motion for Efficient Perception TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:46:40.010650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T05:40:33.752341Z digest=sha256:68ee1981e0262d9d355560f41250bb1b3ec046583106b069a174ae469362cf41

Observation 5fa44cfa-c497-42cd-8aa1-84798f581a31 · inbound

The TIME Machine: On The Power of Motion for Efficient Perception cites this paper.

The TIME Machine: On The Power of Motion for Efficient Perception TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-15T11:06:23.089564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T11:06:23.089564Z digest=sha256:cd5a1d1e2ab6a9e41b4ab782f19f94b58ab14143a2913f9dfe2e7c0073aff990

Observation de64e0c2-82c2-414a-a70c-739353ac54ac · inbound

The TIME Machine: On The Power of Motion for Efficient Perception cites this paper.

The TIME Machine: On The Power of Motion for Efficient Perception TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T13:27:48.232517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:27:48.232517Z digest=sha256:aea5d1bc4bed405b92459bf310cf7d100d0997abae5a5eaf41b824206753f833

Observation b5eb221b-48ea-48d7-8e50-f892bc12077b · inbound

IPIBench: Evaluating Interactive Proactive Intelligence of MLLMs under Continuous Streams cites this paper.

IPIBench: Evaluating Interactive Proactive Intelligence of MLLMs under Continuous Streams TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:33:51.073077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T18:24:57.881644Z digest=sha256:f33c0dfa2286a1042095653e74241e93e6ac1978e3e041b9268466cf61d176fd

Observation 2a848339-a283-4131-80eb-97e8b045757f · inbound

YoCausal: How Far is Video Generation from World Model? A Causality Perspective cites this paper.

YoCausal: How Far is Video Generation from World Model? A Causality Perspective TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:33:15.588037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T08:27:03.674229Z digest=sha256:a93cf414f3a3d6f3c3cbd79bb26db78f7aa3ec478ece5b0cd5efaa110fa1490e

Observation f8049609-c217-4dd0-9b2a-88c25912653e · inbound

TeachObs: A Human-Validated Benchmark for Multimodal Teaching Observation and Model Evaluation cites this paper.

TeachObs: A Human-Validated Benchmark for Multimodal Teaching Observation and Model Evaluation TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:12:46.638618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T23:11:43.712391Z digest=sha256:c1c20c7bbc440904471af4fcbc63ad0519851e2303ed9684f99f368e1104c912

Observation 8ba65bc2-aaee-4236-9f7a-44708d0b76f4 · inbound

Benchmarking Visual State Tracking in Multimodal Video Understanding cites this paper.

Benchmarking Visual State Tracking in Multimodal Video Understanding TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:46:27.984113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T10:43:06.811228Z digest=sha256:1191864cd92b6b590ce839b638e75144b676089024e7145ef6b2306b48e2607f

Observation a21a8d1f-e3e0-4d18-957e-68a4b6df0fbb · inbound

MAOAM: Unified Object and Material Selection with Vision-Language Models cites this paper.

MAOAM: Unified Object and Material Selection with Vision-Language Models TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 142

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:16:26.337706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T11:08:59.900161Z digest=sha256:99a2ab8e0e08c0147914263085449f27136448496b411dc89733e6265f9fae4f

Observation 4544884c-2f76-421b-ae9b-7d6aabf2f597 · inbound

APT: Atomic Physical Transitions for Causal Video-Language Understanding cites this paper.

APT: Atomic Physical Transitions for Causal Video-Language Understanding TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:39:03.236843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T22:06:00.249819Z digest=sha256:a9bdab04e0a9061ce235cd3fa989242433bb371de86b6218357e670481e1597b

Observation fe62423f-d61e-4782-aac6-fa2f3577de9c · inbound

Animation2Code: Evaluating Temporal Visual Reasoning in Video-to-Code Generation cites this paper.

Animation2Code: Evaluating Temporal Visual Reasoning in Video-to-Code Generation TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-01T16:05:49.779006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T00:53:52.724902Z digest=sha256:fd523ef0f61753300fcaaa6943aa4d2db4e30a13396c00085e45b2dbe711055c