Pith. sign in

Paper Citation Record · LEDGER

TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 35 inbound Pith citation observations for arXiv:2312.02051.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.02051 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 35 of 35 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:26:44.493162Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:39:37.643725Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2894a04d-1d74-4417-a654-57b721d1551d · inbound

TempCompass: Do Video LLMs Really Understand Videos? cites this paper.

TempCompass: Do Video LLMs Really Understand Videos? TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 116

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:46:16.807356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-17T02:46:16.632743Z digest=sha256:9dc5f3f683d77e3d0815d7ab1665f565aa2a6db5d944fd4597df7ec64051e5bb

Observation cd0c9261-42f4-4beb-8ad3-e83f25ecae50 · inbound

MLVU: Benchmarking Multi-task Long Video Understanding cites this paper.

MLVU: Benchmarking Multi-task Long Video Understanding TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:55:26.427758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-14T19:55:26.333923Z digest=sha256:ea07a272953549bd0f3333459a85a6ecb47291c9044f308fce8b9bc7d522f56c

Observation e1fd15f6-ca78-464e-9243-f4693104d822 · inbound

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs cites this paper.

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-11T02:44:53.620986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-11T02:44:53.284345Z digest=sha256:b23f06fd1cd33214ba513aae0fc32d1f9875dcc954e3274fd22a062b66387995

Observation ed7db905-7503-481c-b1d0-c2d6a37bc33b · inbound

LVBench: An Extreme Long Video Understanding Benchmark cites this paper.

LVBench: An Extreme Long Video Understanding Benchmark TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:55:30.099872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T11:55:30.048525Z digest=sha256:3ddae6df0125fc9fffe6b5bf9a207e0d2d179cb3a7c0cc895990e5300300397d

Observation ca324a57-d643-4590-a8bb-7f477c6345fe · inbound

CogVLM2: Visual Language Models for Image and Video Understanding cites this paper.

CogVLM2: Visual Language Models for Image and Video Understanding TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:10:27.687282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:9074f3d2a6219fe8c1a38f6189ca707c98eb7f1a5b32c67489178d90f23389d1

Observation b67c3742-40f5-42b3-b43b-dd0707f0f87f · inbound

Is this Generated Person Existed in Real-world? Fine-grained Detecting and Calibrating Abnormal Human-body cites this paper.

Is this Generated Person Existed in Real-world? Fine-grained Detecting and Calibrating Abnormal Human-body TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T15:28:30.295762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:28:30.295762Z digest=sha256:af533fa7ef197defd972b6a5ce524c62a599190329c81448579443c0270bb574

Observation a4734340-4adb-44c2-9d6d-602906a85036 · inbound

TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability cites this paper.

TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:56.200232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:56.200232Z digest=sha256:d539b734771471e34dbb77e3ab47212ce3a7c77b68659beb04f74291b503f16c

Observation 00f39dc4-3c86-4e6e-8c63-c06ad75c3438 · inbound

VCA: Video Curious Agent for Long Video Understanding cites this paper.

VCA: Video Curious Agent for Long Video Understanding TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.245087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.245087Z digest=sha256:1de5c6dbf2ecdfe4e6e456b241b830b2a7caa71c8ae725f80cfbb2e22701035c

Observation a3e98c36-3502-47d6-b542-48226d44bbbe · inbound

Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment cites this paper.

Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T00:47:12.355668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:47:12.355668Z digest=sha256:d774b8828f2813ab0f7dd37e64e7f81b59fdc364663e8ec4c6c4a4c5ed983db5

Observation eebf7881-e136-4104-8bed-c34d43582f63 · inbound

GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models cites this paper.

GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:55.519932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:55.519932Z digest=sha256:ae539323bb46aa7558ee7a7005ed088004a4c574e1cc9bea52c1583dcef0fe80

Observation 980131d7-c4f1-4d64-a340-6fb74a1ff9e0 · inbound

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs cites this paper.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.497615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.497615Z digest=sha256:0f1782d1c2b3f55e6d51736e8786256f653df146c7610ea7a388bb7ce80c6f18

Observation 1124415d-0c4b-421a-a124-df65e9d5acbc · inbound

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models cites this paper.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T20:34:12.264313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:34:12.264313Z digest=sha256:8f6709109d6297d18fc38b9be2087dd6094aedf6a6bcf27844fd657891c1020f

Observation 74733113-350f-4079-852c-d6a93c738232 · inbound

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling cites this paper.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.762202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:eb77011483f54432473d94434cfe39f2903970560465bab44b4eba33d43e8df4

Observation b7bb7727-ba7d-4ecc-ba1e-6079cc4e9f21 · inbound

MTPChat: A Multimodal Time-Aware Persona Dataset for Conversational Agents cites this paper.

MTPChat: A Multimodal Time-Aware Persona Dataset for Conversational Agents TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T17:37:38.120634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:37:38.120634Z digest=sha256:4e11f6989c3753b32a54d1feac3c9c125f929d793f78898a1bcd68d5f32f9505

Observation 5050fcb7-e5de-406d-afe0-4055c9267418 · inbound

ClassComet: Exploring and Designing AI-generated Danmaku in Educational Videos to Enhance Online Learning cites this paper.

ClassComet: Exploring and Designing AI-generated Danmaku in Educational Videos to Enhance Online Learning TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-16T10:26:44.493162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:26:44.493162Z digest=sha256:4202ea16f6a398570cfe78c69cbf9e08d41b323d5c323b7b5d3790d4b37fde6f

Observation 0bbf7527-2e0a-43dd-8862-846ec2812208 · inbound

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction cites this paper.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:06.501876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:06.501876Z digest=sha256:d418c73d6b3e9a49cf700f2c706bbd68e51f7b0b59f3b01debae898acd6a73ef

Observation db37397a-99a4-4b89-a1ef-8f05e1286840 · inbound

EASG-Bench: Video Q&A Benchmark with Egocentric Action Scene Graphs cites this paper.

EASG-Bench: Video Q&A Benchmark with Egocentric Action Scene Graphs TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:17.607966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:17.607966Z digest=sha256:044a97695db9a987d89e5a1536bed6f78fe1c2aa0d3c6f443628ce8c7728dfd7

Observation 039f19b2-59e5-41e2-b0c1-42a3e1a0c856 · inbound

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs cites this paper.

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T22:24:31.342995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:24:31.342995Z digest=sha256:ffb59b929fc00ed21284c0fc27bf3f5d512332e65c53262fd2f0139df9166064

Observation c76d8af5-6ecf-405e-8499-77ad6f20ccc0 · inbound

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025 cites this paper.

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025 TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:20:12.971888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:20:12.971888Z digest=sha256:2325dbcaff58350d94c69e6f3903fa5553cb26d770736cc767a28ef4ca2b5da0

Observation ce5d04b7-30d0-4adc-8fce-0d5913fe9b60 · inbound

MANTA: Cross-Modal Semantic Alignment and Information-Theoretic Optimization for Long-form Multimodal Understanding cites this paper.

MANTA: Cross-Modal Semantic Alignment and Information-Theoretic Optimization for Long-form Multimodal Understanding TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:00:06.968000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:00:06.968000Z digest=sha256:5e7e99cc08290f9183546d9f794e074ce4becc7e2fd31c47b583e40eedb2d3ca

Observation b85caa48-9723-4b62-a11a-6e2c148e0057 · inbound

A Paradigm Shift: Fully End-to-End Training for Temporal Sentence Grounding in Videos cites this paper.

A Paradigm Shift: Fully End-to-End Training for Temporal Sentence Grounding in Videos TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-13T19:48:11.525080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T19:43:57.298338Z digest=sha256:c7c71cb7a8c9b658d1748b6cc21d14b7778e8f606d2e0036a14105ad3d3c9a3a

Observation bbc5f630-013a-40d7-a1a7-2ec06f18f519 · inbound

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding cites this paper.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T00:46:13.738281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-07T14:10:27.416341Z digest=sha256:e20f0c38c48993c14fc4631a245610dd025e3a1350d6b32889c7f15ef2d43a34

Observation 7ba36f90-8d5b-4224-860b-09a8bac66510 · inbound

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding cites this paper.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T08:45:34.755963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:61fb3071a30b6d09eadb44c778a668fe8c6ed40a3e740d9ede7cb19d2f6c5245

Observation 62853dd5-b767-49ba-a9a9-4a13f37bcfab · inbound

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence cites this paper.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.539308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:04aed0b4067c58ba34fb8f42eea7306edd5cead2e65a399769b9e37e07bd97b1

Observation c5036575-fc9c-419e-84c4-37101201c4e5 · inbound

VTI-CoT: Visual-Textual Interleaved Chain of Thought for Video Reasoning cites this paper.

VTI-CoT: Visual-Textual Interleaved Chain of Thought for Video Reasoning TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:46:56.832350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T01:52:44.785582Z digest=sha256:b9861ff72b645b98b934f1f2a24ac59c0649f6991ab8ab92e0ef398281204428

Observation 99f19cce-399a-45ea-af49-7a991abf30fc · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.504136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:fae152a10feee5e2d66093aa483dd643aa2dd6bbba443c19e50b5f96bd043b06

Observation 720ab0c3-e66d-454c-a5d5-e85fc57d5822 · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 193

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:39:37.645127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:e9a5723a77a468fa9c37e607964723d8080442e53ede99b3b22e841c21a2bb81

Observation 3e1094fb-58b5-43cf-8056-6f2b0e2139b4 · inbound

EFlow: Learning Evidence Flow for Long-Video Reasoning with Adaptive Reflection cites this paper.

EFlow: Learning Evidence Flow for Long-Video Reasoning with Adaptive Reflection TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T14:27:03.324975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-02T14:19:04.890074Z digest=sha256:29c03f52037c09a8823b51a9378250ed3f029980a1af4c2e8481ff29701b5e39

Observation af988246-9c7f-4120-a332-cdde69102010 · inbound

EFlow: Learning Evidence Flow for Long-Video Reasoning with Adaptive Reflection cites this paper.

EFlow: Learning Evidence Flow for Long-Video Reasoning with Adaptive Reflection TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-12T09:13:38.446621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:13:38.446621Z digest=sha256:6907d5d8e30166ed3d97068199645dad21f6f7edc847245938fca048aca2f944

Observation 66d9599d-a935-4146-8499-6ec150eea4e5 · inbound

EFlow: Learning Evidence Flow for Long-Video Reasoning with Adaptive Reflection cites this paper.

EFlow: Learning Evidence Flow for Long-Video Reasoning with Adaptive Reflection TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T09:17:20.566842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:17:20.566842Z digest=sha256:e3c040ad71de0d386546b4c0d0f5f05f8f0071c26cf61ce8efefdf7bf709b4ec

Observation 0e22faf1-5308-4869-a1b5-4ca00032fe8c · inbound

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences cites this paper.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:18b582ef1ffa78a5770c5fc11509eb675fb68c009fd0238fe58f829a1f71c197

Observation ddf24105-f1b7-4d79-bbad-836acba77bba · inbound

Mixture of Probes: Learning from Privileged Modalities in Multimodal LLMs Through Probing cites this paper.

Mixture of Probes: Learning from Privileged Modalities in Multimodal LLMs Through Probing TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-13T06:21:17.706173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T06:21:17.706173Z digest=sha256:53bc3def379d090c6a950e4811cf73c9fec59e67eb397edb021cfa7fe82e8dc1

Observation b0c52bc0-b812-4f45-9d1b-c2b9ab76835b · inbound

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs cites this paper.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:15.439960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:15.439960Z digest=sha256:a7f28c3601ae2c883bd976292ee354e8377c4450b15bab056ae9b69f34028a72

Observation 75c7b714-5509-4238-bc47-182e8a457c22 · inbound

TimePLE: Rethinking Temporal Representation for Video Temporal Grounding cites this paper.

TimePLE: Rethinking Temporal Representation for Video Temporal Grounding TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-31T23:32:53.515428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:32:53.515428Z digest=sha256:254269e48da6129f9a469c72d418c62ec1b78914b185736d15825c23bd4f8c49

Observation 7373f027-c75e-455f-a08b-410c4054f616 · inbound

AdaThinkV: Adaptive Thinking for Token-Efficient Video Reasoning cites this paper.

AdaThinkV: Adaptive Thinking for Token-Efficient Video Reasoning TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-04T17:24:16.109594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:24:16.109594Z digest=sha256:e5852c6053e526707d4e6150f735615ef46446bf26726c6a58e319493675865f