Pith. sign in

Paper Citation Record · LEDGER

TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2312.02051.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.02051 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T17:37:38.120634Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:39:37.643725Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2894a04d-1d74-4417-a654-57b721d1551d · inbound

TempCompass: Do Video LLMs Really Understand Videos? cites this paper.

TempCompass: Do Video LLMs Really Understand Videos? TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 116

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:46:16.807356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-17T02:46:16.632743Z digest=sha256:6de9bd923dca86da51d56c4e773a6f8c3d1cbb01b932f387142ac87b54ba184c

Observation cd0c9261-42f4-4beb-8ad3-e83f25ecae50 · inbound

MLVU: Benchmarking Multi-task Long Video Understanding cites this paper.

MLVU: Benchmarking Multi-task Long Video Understanding TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:55:26.427758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T19:55:26.333923Z digest=sha256:cb78d9a10ef531b651e9a7e334911f3fa3029927e6431e5a896e02ea7358f49d

Observation e1fd15f6-ca78-464e-9243-f4693104d822 · inbound

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs cites this paper.

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-11T02:44:53.620986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T02:44:53.284345Z digest=sha256:c8cb36f74e8e0a4dde870a9acf08fb9c5f15f86aa56fee6a25d5520abe037852

Observation ed7db905-7503-481c-b1d0-c2d6a37bc33b · inbound

LVBench: An Extreme Long Video Understanding Benchmark cites this paper.

LVBench: An Extreme Long Video Understanding Benchmark TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:55:30.099872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T11:55:30.048525Z digest=sha256:5bfa4ef6812e66e8ee8a1f675b494b6399e54ef94d5336d1ed47528d77041ba9

Observation ca324a57-d643-4590-a8bb-7f477c6345fe · inbound

CogVLM2: Visual Language Models for Image and Video Understanding cites this paper.

CogVLM2: Visual Language Models for Image and Video Understanding TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:10:27.687282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:bbcd27829d1b49e22290a2ff77db68a3332daa6db970feb434ab85de46b1e90b

Observation 74733113-350f-4079-852c-d6a93c738232 · inbound

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling cites this paper.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.762202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:c3c6f39a3add51d83b271721e5d360edbf8965e45a2c6f8d7d2167fcb926114b

Observation b7bb7727-ba7d-4ecc-ba1e-6079cc4e9f21 · inbound

MTPChat: A Multimodal Time-Aware Persona Dataset for Conversational Agents cites this paper.

MTPChat: A Multimodal Time-Aware Persona Dataset for Conversational Agents TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T17:37:38.120634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:37:38.120634Z digest=sha256:13a7ced027a12afc2c8882d16d746552239c1750781db662c94d6e5caad26d85

Observation 0bbf7527-2e0a-43dd-8862-846ec2812208 · inbound

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction cites this paper.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:06.501876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:06.501876Z digest=sha256:f53e437ac9912274c124e6689eff460a4d6db3fbaa42615f370633ed61067973

Observation db37397a-99a4-4b89-a1ef-8f05e1286840 · inbound

EASG-Bench: Video Q&A Benchmark with Egocentric Action Scene Graphs cites this paper.

EASG-Bench: Video Q&A Benchmark with Egocentric Action Scene Graphs TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:17.607966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:17.607966Z digest=sha256:8215695c7ff573ca61fe33e4de98ef16edda74eb22407db1cdf73db788cc2481

Observation 039f19b2-59e5-41e2-b0c1-42a3e1a0c856 · inbound

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs cites this paper.

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T22:24:31.342995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:24:31.342995Z digest=sha256:0119a94e2fef7d79e1b997f3649fb5d6715e29a1a0218ac7b7335700dda87f0c

Observation c76d8af5-6ecf-405e-8499-77ad6f20ccc0 · inbound

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025 cites this paper.

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025 TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:20:12.971888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:20:12.971888Z digest=sha256:ce76b91b2ee86066cf08ac9b784a575d801a323040fdd07385708bd75e8cd166

Observation ce5d04b7-30d0-4adc-8fce-0d5913fe9b60 · inbound

MANTA: Cross-Modal Semantic Alignment and Information-Theoretic Optimization for Long-form Multimodal Understanding cites this paper.

MANTA: Cross-Modal Semantic Alignment and Information-Theoretic Optimization for Long-form Multimodal Understanding TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:00:06.968000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:00:06.968000Z digest=sha256:77aa9f3f6a22ce76905d07076938d53046471811ed47eea40d2cb5ea74b7f4fc

Observation b85caa48-9723-4b62-a11a-6e2c148e0057 · inbound

A Paradigm Shift: Fully End-to-End Training for Temporal Sentence Grounding in Videos cites this paper.

A Paradigm Shift: Fully End-to-End Training for Temporal Sentence Grounding in Videos TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-13T19:48:11.525080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T19:43:57.298338Z digest=sha256:d575c47f1e81861b324788ad5986722f08e319e08c078b567624f468e3617d6f

Observation bbc5f630-013a-40d7-a1a7-2ec06f18f519 · inbound

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding cites this paper.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T00:46:13.738281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-07T14:10:27.416341Z digest=sha256:3f450bbd2abed5dd31e21625cfc1351bd34f811be804fd312d0340941d01f2c0

Observation 7ba36f90-8d5b-4224-860b-09a8bac66510 · inbound

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding cites this paper.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T08:45:34.755963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:f4d594adfc01f66056d1cd10494da755092fa247268af58ad850563e158a339b

Observation 62853dd5-b767-49ba-a9a9-4a13f37bcfab · inbound

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence cites this paper.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.539308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:cb12bebb08784dd38d07412a90218db7ad9f84d8dcc905f2ee042dee114e558e

Observation c5036575-fc9c-419e-84c4-37101201c4e5 · inbound

VTI-CoT: Visual-Textual Interleaved Chain of Thought for Video Reasoning cites this paper.

VTI-CoT: Visual-Textual Interleaved Chain of Thought for Video Reasoning TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:46:56.832350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T01:52:44.785582Z digest=sha256:fbdbe2be42e796f0e1b008e4c80bd4b6259c9df6d3203805625c5a742213d8ac

Observation 99f19cce-399a-45ea-af49-7a991abf30fc · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.504136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:23105a4967e3bd8163c8b67de8cc3e261a497156c25e4766a2790f67aaf72bfe

Observation 720ab0c3-e66d-454c-a5d5-e85fc57d5822 · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 193

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:39:37.645127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:254611c4d03af7f52a02e1248083a9b7c613befb16e3d6e6a91db9484facd2f8

Observation 3e1094fb-58b5-43cf-8056-6f2b0e2139b4 · inbound

EFlow: Learning Evidence Flow for Long-Video Reasoning with Adaptive Reflection cites this paper.

EFlow: Learning Evidence Flow for Long-Video Reasoning with Adaptive Reflection TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T14:27:03.324975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T14:19:04.890074Z digest=sha256:b30666e7d38bec139c2be1dd389076dcccb7a42514c8eae1a4456054c8b937c7

Observation af988246-9c7f-4120-a332-cdde69102010 · inbound

EFlow: Learning Evidence Flow for Long-Video Reasoning with Adaptive Reflection cites this paper.

EFlow: Learning Evidence Flow for Long-Video Reasoning with Adaptive Reflection TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-12T09:13:38.446621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:13:38.446621Z digest=sha256:133f51a099f031ddfa20d79535a302d246f657adf56e6a8ac26452b507d23ac5

Observation 66d9599d-a935-4146-8499-6ec150eea4e5 · inbound

EFlow: Learning Evidence Flow for Long-Video Reasoning with Adaptive Reflection cites this paper.

EFlow: Learning Evidence Flow for Long-Video Reasoning with Adaptive Reflection TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T09:17:20.566842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:17:20.566842Z digest=sha256:7e22d19e099cf14b7ee0cb5cd5283047674b41c318cf5097b95c58cbdb9d4236

Observation 0e22faf1-5308-4869-a1b5-4ca00032fe8c · inbound

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences cites this paper.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:39930770a98982e357e887112f914e18e3f46f238037e4186d04b901cc56a8ed

Observation ddf24105-f1b7-4d79-bbad-836acba77bba · inbound

Mixture of Probes: Learning from Privileged Modalities in Multimodal LLMs Through Probing cites this paper.

Mixture of Probes: Learning from Privileged Modalities in Multimodal LLMs Through Probing TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-13T06:21:17.706173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T06:21:17.706173Z digest=sha256:12509027b3252319a96fa41ccd33eafaa3bfb3f238e3de160b4d2191208ba709

Observation b0c52bc0-b812-4f45-9d1b-c2b9ab76835b · inbound

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs cites this paper.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:15.439960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:15.439960Z digest=sha256:75e928cf56d433c8269148004eb1e82392f4afa7dc73390a714b399f8356cbc4

Observation 75c7b714-5509-4238-bc47-182e8a457c22 · inbound

TimePLE: Rethinking Temporal Representation for Video Temporal Grounding cites this paper.

TimePLE: Rethinking Temporal Representation for Video Temporal Grounding TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-31T23:32:53.515428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:32:53.515428Z digest=sha256:15a9e1c3559f9033cdb90fcb91321fdf52056f3946523ea83761394e3330a4ef

Observation 7373f027-c75e-455f-a08b-410c4054f616 · inbound

AdaThinkV: Adaptive Thinking for Token-Efficient Video Reasoning cites this paper.

AdaThinkV: Adaptive Thinking for Token-Efficient Video Reasoning TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-04T17:24:16.109594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:24:16.109594Z digest=sha256:0b7f6f6465e36a4f923724c5f9ccbbd84572d4a677a58c70d23c3667fdb10982