Pith. sign in

Paper Citation Record · LEDGER

TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 44 inbound Pith citation observations for arXiv:2410.10818.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.10818 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 44 of 44 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T19:06:46.147800Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T23:39:03.234328Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d5b6241f-c506-4dde-888a-5ebddaa3d455 · inbound

LLaVA-Video: Video Instruction Tuning With Synthetic Data cites this paper.

LLaVA-Video: Video Instruction Tuning With Synthetic Data TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 216

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:20:33.170271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:e74a2cbf594c633e0d7f8ee27229fa2411bb3976cdf03ee3c925dcc54b887e41

Observation 43e812f6-e863-40dd-8404-0d52591f11a0 · inbound

Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! cites this paper.

Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T19:06:46.147800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:06:46.147800Z digest=sha256:084f3ca5af49369beb6d220ff6d3d2ffb85f1058428276c5fcaa899e95b0a616

Observation 12dd80fa-586b-45d7-9b83-9d90bb319641 · inbound

MMVU: Measuring Expert-Level Multi-Discipline Video Understanding cites this paper.

MMVU: Measuring Expert-Level Multi-Discipline Video Understanding TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-10T17:15:23.810158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:15:23.810158Z digest=sha256:a914f06c8d8e746769246e253781372661c1818e7b95ae9bbcbad6a73860a8b4

Observation 8728683d-c1dd-4cae-8913-cb685f76a829 · inbound

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos cites this paper.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:32:41.302744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:0d11714a30b6ec63448bff3e1a7ec2e4d5287e399d02e0eb5ed21db3a2056ce1

Observation 333f9254-8eb1-45d4-bb86-ce42e95dab06 · inbound

HD-EPIC: A Highly-Detailed Egocentric Video Dataset cites this paper.

HD-EPIC: A Highly-Detailed Egocentric Video Dataset TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T23:26:20.268946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T23:26:20.268946Z digest=sha256:4c72912e884df26d9c37ad18fbfcb8a286d63bdb58f7f3b53df154baddc40dfa

Observation 4d7bc3d6-cca5-45e2-9860-084fb8cfc5d9 · inbound

Seed1.5-VL Technical Report cites this paper.

Seed1.5-VL Technical Report TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:05.681583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:97e445f822d81dd43cd3e672f204fb60960e671f362d12b172bddf076b59bc0c

Observation a6bc6d7e-b309-427f-aa28-fe817f779cc3 · inbound

RTime-QA: A Benchmark for Atomic Temporal Event Understanding in Large Multi-modal Models cites this paper.

RTime-QA: A Benchmark for Atomic Temporal Event Understanding in Large Multi-modal Models TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:07.209929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:24:07.209929Z digest=sha256:173b6dde43ab2f9548b90bd07297456a3bc88982aa857619f2467dd5662f594c

Observation 0068b704-aecd-46cc-a8c5-175b3944aa05 · inbound

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos cites this paper.

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:59.572640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:02:59.572640Z digest=sha256:e17678711373a003eda844dc2262b47e06a27781acf8d96a61155b7ccaed6e12

Observation 358b7850-828e-4e44-905f-3a296da34640 · inbound

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding cites this paper.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:17:18.545824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:c47ae3def191b9676a0ce3da8574b4ba0870db7df9c9346ed862a31e456ec72c

Observation 8c97c53f-09c4-49cc-a2be-eebe6f273efc · inbound

Fostering Video Reasoning via Next-Event Prediction cites this paper.

Fostering Video Reasoning via Next-Event Prediction TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:40.431117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:40.431117Z digest=sha256:bf5981c95c11e9907f7d6942d50f5cb9d7c3c3677a00cf64402ce5768d682408

Observation 305f6973-6c73-49c7-a805-d19b14918540 · inbound

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times cites this paper.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.121411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.121411Z digest=sha256:67d41b073ac6c750cfebd96883bb7cdf3a3de7b12ca38f0b8c2aaa2e080d9511

Observation 0a22fff6-cf20-4c15-89bc-b46ffd20a18e · inbound

Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs cites this paper.

Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:07:15.380828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T11:03:59.222849Z digest=sha256:26414eb548e07165a4aec619e3c01f31adb6dc16778ea9eeb8197408e21d1ee5

Observation b3e7157c-7b90-4817-989b-f5200612b389 · inbound

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning cites this paper.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:33:50.948202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:bcae1ccfe0905a775eac12d0efa3da3f71bdc2dbb81aced75c3842ecaf0623be

Observation e849daa4-bf69-475c-8e19-353675a6c28c · inbound

ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models cites this paper.

ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:13.972207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:03:13.972207Z digest=sha256:af87c9297a45fd7d8a5856f578c64a9e2e59306782405b63c9357c6434edf1c6

Observation bbe643d1-35b2-4d2e-9e1a-e0b64b67fc02 · inbound

GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? cites this paper.

GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T17:58:26.999928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:58:26.999928Z digest=sha256:326f4d1f57c5e79b02353adc12993760a965f506aca6b97cad73c03457381ce1

Observation f91d00e8-c6f6-491f-b1a5-7bdc970597a8 · inbound

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding cites this paper.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:31.035101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:31.035101Z digest=sha256:baee25117ddb1a02f1f2d67c73b25f7a6c11af98f1af619af970ed2a2d62c03e

Observation ad492195-bd08-400d-a93c-a65b26054333 · inbound

CausalStep: A Benchmark for Explicit Stepwise Causal Reasoning in Videos cites this paper.

CausalStep: A Benchmark for Explicit Stepwise Causal Reasoning in Videos TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T15:14:17.287285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:14:17.287285Z digest=sha256:d596b0b977a6f7f50bd1751cc1b70b5f4dda049145ff6b2a40ce2b69553bd834

Observation 15f9e0ce-387e-4e56-a998-6d9afa19ef88 · inbound

AdsQA: Towards Advertisement Video Understanding cites this paper.

AdsQA: Towards Advertisement Video Understanding TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T20:20:36.661579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:20:36.661579Z digest=sha256:84f624a691e3c259cf3ba0f852eaac8997fc6dbc11fef09e58b8b31806d7a6c7

Observation af06448f-fba1-4f4c-bcbd-94a2990c6610 · inbound

NeMo: Needle in a Montage for Video-Language Understanding cites this paper.

NeMo: Needle in a Montage for Video-Language Understanding TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T13:54:11.277795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:54:11.277795Z digest=sha256:c5d218c2e124c5545a9245aca39dbb41a836ee267e95c439198f21111f874d96

Observation 6a8aaab0-b1f6-4fde-bec6-dbbb05df58b9 · inbound

LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling cites this paper.

LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-22T12:31:32.156486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T12:26:35.347190Z digest=sha256:44bb3255a792692ef2197426c0aa5fa324555098021778a3dd1732dab095e20f

Observation d12b6c69-c134-451c-910b-20cabb957091 · inbound

See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models cites this paper.

See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:13:52.103038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T02:12:55.170296Z digest=sha256:d51663fac15a1180b093d0ed03db8e16de25fa276d78a9435dc910b5f30671ed

Observation c84d5d53-66bb-46aa-8f21-429c8c812f36 · inbound

From Segments to Scenes: Temporal Understanding for Agentic Autonomous Driving via Vision-Language Models cites this paper.

From Segments to Scenes: Temporal Understanding for Agentic Autonomous Driving via Vision-Language Models TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T18:28:15.922879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:28:15.922879Z digest=sha256:c95d9d1ba5fadfe4049534598206d4b165e2481fc704ff42bd3493cb89d5f0b5

Observation f6290141-24a4-455e-b8f9-c5e6fe3d58e8 · inbound

GTASA: Ground Truth Annotations for Spatiotemporal Analysis, Evaluation and Training of Video Models cites this paper.

GTASA: Ground Truth Annotations for Spatiotemporal Analysis, Evaluation and Training of Video Models TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:26:01.762020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:10.956726Z digest=sha256:38178919cb8c92a9cc7e1e81515f2510035106957876bee9ae4615dbe1aaa614

Observation 6968e55c-af11-43d4-910d-b458eb5ba282 · inbound

POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs cites this paper.

POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:41:04.015830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T15:23:08.671342Z digest=sha256:c7a52e7ec204615ff599ac28e6a0815c39f5d0c6de00837d9c99da325fae1710

Observation 4d934730-f911-4785-8ed2-1e50700138de · inbound

When Text Hijacks Vision: Benchmarking and Mitigating Text Overlay-Induced Hallucination in Vision Language Models cites this paper.

When Text Hijacks Vision: Benchmarking and Mitigating Text Overlay-Induced Hallucination in Vision Language Models TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:46:37.546797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T06:41:59.641410Z digest=sha256:8d23c02cff902bb8718b676ea4d1914bfd4995d4fa107c658e73cbf5bdeb1ebd

Observation ebfb5115-25aa-42a9-8026-15d9f12fafa5 · inbound

VideoNet: A Large-Scale Dataset for Domain-Specific Action Recognition cites this paper.

VideoNet: A Large-Scale Dataset for Domain-Specific Action Recognition TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:20:41.264335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T18:35:48.379198Z digest=sha256:b13882f6e6d31fcc931b90b5b25df7d02a3628328f618aad4be21e32e1f3eacd

Observation 83c7d2dd-2323-4155-b984-2468ff757f85 · inbound

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models cites this paper.

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:31:26.606021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:13:21.487431Z digest=sha256:46bad3df19808aec49404fc112a1e03265caab515967e6a088d6b9f20a94a22f

Observation 7d3ca994-57b8-40e7-917c-a5a53953eb14 · inbound

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models cites this paper.

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:57:28.298673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T06:53:42.726350Z digest=sha256:29f2892e0683eef1a62daa413ffe8cfb20d8f8e1cba87f3e686c9062562d90fb

Observation 7559d5ad-eb08-48c6-aa60-ead607b633ac · inbound

EvoGround: Self-Evolving Video Agents for Video Temporal Grounding cites this paper.

EvoGround: Self-Evolving Video Agents for Video Temporal Grounding TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:32:52.564646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T19:29:47.356665Z digest=sha256:e1d62e81e931107ae48f3b5743e5213032b439835c4f87ee705722fa59c66c9e

Observation f1700d39-3c5f-45fb-8f20-19724b4ed41a · inbound

Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding cites this paper.

Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:03:08.044136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T16:02:53.887605Z digest=sha256:f1a3968b4e40179d0e6315d637b1c9e3b3851b3c6eb8e6b2f23761d33fa4061d

Observation aaa66853-e66b-4ffb-8c41-8f8517e6def5 · inbound

FineBench: Benchmarking and Enhancing Vision-Language Models for Fine-grained Human Activity Understanding cites this paper.

FineBench: Benchmarking and Enhancing Vision-Language Models for Fine-grained Human Activity Understanding TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:13:05.462877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T06:08:12.887394Z digest=sha256:79614251c1956ee009d64aa9fb1084185b96c91c8cf3df7ceb75754b294a61f4

Observation a057f476-9899-436f-91f4-48608cfcacc4 · inbound

FineBench: Benchmarking and Enhancing Vision-Language Models for Fine-grained Human Activity Understanding cites this paper.

FineBench: Benchmarking and Enhancing Vision-Language Models for Fine-grained Human Activity Understanding TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:04:02.791065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T07:59:57.488107Z digest=sha256:38b7e8373ea72a8cae95d0662403adf819e153ada7f1f94d064000ae6de9c1da

Observation 165ec764-c51b-4b13-8e50-4780a9e9f8e9 · inbound

FineBench: Benchmarking and Enhancing Vision-Language Models for Fine-grained Human Activity Understanding cites this paper.

FineBench: Benchmarking and Enhancing Vision-Language Models for Fine-grained Human Activity Understanding TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:14:59.999305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T18:09:49.178090Z digest=sha256:d4e8391a7c643b6f72d3ca0c9e2aaa599f815cab8d4328412f2bff8a4754d68f

Observation e6080247-df3f-4d8b-944b-8736e0d26d6a · inbound

Learning Spatiotemporal Sensitivity in Video LLMs via Counterfactual Reinforcement Learning cites this paper.

Learning Spatiotemporal Sensitivity in Video LLMs via Counterfactual Reinforcement Learning TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:46:14.880613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T07:45:56.473188Z digest=sha256:852c07e9c8df3e9b0f25ed5c5a691edf536619daebdfb9e65d14e3adc747292d

Observation 9aca9f84-e2ee-4eb9-a074-cb6fe0b21f5e · inbound

The TIME Machine: On The Power of Motion for Efficient Perception cites this paper.

The TIME Machine: On The Power of Motion for Efficient Perception TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:46:40.010650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T05:40:33.752341Z digest=sha256:1c6b7953372871e3a7a5f6cca73f26ad5cc3f54c4cc1c9899e7ec7c75548c1e9

Observation 5fa44cfa-c497-42cd-8aa1-84798f581a31 · inbound

The TIME Machine: On The Power of Motion for Efficient Perception cites this paper.

The TIME Machine: On The Power of Motion for Efficient Perception TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-15T11:06:23.089564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T11:06:23.089564Z digest=sha256:d7f4974713b9a8c6889a7ec26c353cd6a61f5450aadabcaec1c5b3d053eb2971

Observation de64e0c2-82c2-414a-a70c-739353ac54ac · inbound

The TIME Machine: On The Power of Motion for Efficient Perception cites this paper.

The TIME Machine: On The Power of Motion for Efficient Perception TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T13:27:48.232517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:27:48.232517Z digest=sha256:ba58bdf81b6363d4c929c5a9e98b235390e7096eb1e29e4bb5dceecfaaca9504

Observation b5eb221b-48ea-48d7-8e50-f892bc12077b · inbound

IPIBench: Evaluating Interactive Proactive Intelligence of MLLMs under Continuous Streams cites this paper.

IPIBench: Evaluating Interactive Proactive Intelligence of MLLMs under Continuous Streams TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:33:51.073077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T18:24:57.881644Z digest=sha256:1d8fde8bbc3d4829c2cd8c99ff6f1b9ecd8373b97e19be7b67e9678d79df6c74

Observation 2a848339-a283-4131-80eb-97e8b045757f · inbound

YoCausal: How Far is Video Generation from World Model? A Causality Perspective cites this paper.

YoCausal: How Far is Video Generation from World Model? A Causality Perspective TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:33:15.588037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T08:27:03.674229Z digest=sha256:9f607bba53c96523952cd694d1e361a68f41da9a14fc1cf472d3f6d33fa72f62

Observation f8049609-c217-4dd0-9b2a-88c25912653e · inbound

TeachObs: A Human-Validated Benchmark for Multimodal Teaching Observation and Model Evaluation cites this paper.

TeachObs: A Human-Validated Benchmark for Multimodal Teaching Observation and Model Evaluation TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:12:46.638618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T23:11:43.712391Z digest=sha256:abf5be7ad9b09672d2cdca6a4b4e702e9b9927b8e371f6e4694fc559632ede69

Observation 8ba65bc2-aaee-4236-9f7a-44708d0b76f4 · inbound

Benchmarking Visual State Tracking in Multimodal Video Understanding cites this paper.

Benchmarking Visual State Tracking in Multimodal Video Understanding TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:46:27.984113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:43:06.811228Z digest=sha256:fd7eb1c5e7084a5d44c8c0db8b6619cd1daa5c33943bab00349b321321b2a2b9

Observation a21a8d1f-e3e0-4d18-957e-68a4b6df0fbb · inbound

MAOAM: Unified Object and Material Selection with Vision-Language Models cites this paper.

MAOAM: Unified Object and Material Selection with Vision-Language Models TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 142

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:16:26.337706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T11:08:59.900161Z digest=sha256:7425b4ab9fac2016d605cf01e32abe9d95dd7a7aab690ba20e81493fdd6295d2

Observation 4544884c-2f76-421b-ae9b-7d6aabf2f597 · inbound

APT: Atomic Physical Transitions for Causal Video-Language Understanding cites this paper.

APT: Atomic Physical Transitions for Causal Video-Language Understanding TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:39:03.236843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T22:06:00.249819Z digest=sha256:16bade3f911e512258cabdfc2effca92b0cabe651910e4a982a1444c83870d57

Observation fe62423f-d61e-4782-aac6-fa2f3577de9c · inbound

Animation2Code: Evaluating Temporal Visual Reasoning in Video-to-Code Generation cites this paper.

Animation2Code: Evaluating Temporal Visual Reasoning in Video-to-Code Generation TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-01T16:05:49.779006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T00:53:52.724902Z digest=sha256:f6529ecf3ed4a929e92f3f2ac8c451b0b8d7e773786e2d95405f3dee92473c72