Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T17:07:00.966961Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 0 inbound Pith citation observations for arXiv:2411.12951.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T17:07:00.966961Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
59 of 59 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 08c990da-1e8e-4903-88b8-653ca79ca3cd · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b955f145-fd79-42bf-b5b4-2da3da0f60c4 · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension The surprising effectiveness of multimodal large language models for video moment retrieval
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3c56589-7156-44e0-8e30-add47b684379 · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension End- to-end object detection with transformers
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7f1de3a9-2cdc-4bca-ae3d-de878c1e28fa · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad1d2f73-aa96-4fe9-a1c5-a424690178f3 · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension Instructblip: Towards general-purpose vision- language models with instruction tuning, 2023
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6a659422-42fe-42db-b2f5-a86f256dd7a9 · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension Measuring and improving consistency in pretrained language models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 628ecb47-2ece-4484-bbbe-89b43659ec8e · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension Tall: Temporal activity localization via language query
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1c10303-a580-4b2c-85bb-671a12b0ab79 · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension VTG-LLM: Integrating Timestamp Knowledge into Video LLMs for Enhanced Video Temporal Grounding
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c50191b1-4e91-419b-94f6-ceb545b407d0 · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension Vtimellm: Empower llm to grasp video moments
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 83e26ae3-5819-42ba-a3f9-862648089c34 · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension LITA: Language Instructed Temporal-Localization Assistant
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0cca322-944d-419b-a950-650c1a0b051e · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension Modal-specific pseudo query genera- tion for video corpus moment retrieval
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 68388e89-0a61-4557-ba7c-0b2692f18aef · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension Background-aware moment detection for video moment retrieval
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 129af2f7-9a6b-4b7e-bc82-8b43005c2458 · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension Language Repository for Long Video Understanding
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9a11936-c5ec-4787-9ec8-97e76b1767f8 · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension Dense-captioning events in videos
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6ff56cb3-4f18-41ec-ae05-55a96c0bb826 · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension Detecting mo- ments and highlights in videos via natural language queries
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3b98d44e-4679-4ccc-9437-5f7e00f4ee6b · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension VideoChat: Chat-Centric Video Understanding
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation daf696bf-3bbd-4c77-a87b-b3a180fd56cd · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension Mvbench: A comprehensive multi-modal video understand- ing benchmark
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6eb5061f-cf7e-4ea8-b19a-63b639977b23 · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension VITATECS: A Diagnostic Dataset for Temporal Concept Understanding of Video-Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb8b437f-aed0-497b-b36b-f42d083d8a54 · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension Benchmarking and Improving Generator-Validator Consistency of Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5ef0812-feb0-4a45-9f1d-dbc5e192265f · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8fe26f4-9732-427d-ab48-ac4823d85ce5 · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension Visual instruction tuning, 2023
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcc114b1-bc9b-4935-b563-bf168f02d0ec · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension TempCompass: Do Video LLMs Really Understand Videos?
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6af2b771-a107-4a7a-a8d2-f466382770b7 · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd341e08-3ec7-49fb-a242-d8f604083cb9 · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension Video-chatgpt: Towards detailed video understanding via large vision and language models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation cf2e423e-53b8-461c-a37a-a13460241155 · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension Query-dependent video representation for moment retrieval and highlight detection
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f10669a0-866d-4530-b045-cb19b8ef92db · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension Local-global video-text interactions for temporal grounding
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 081f7dd8-e3a6-4ac6-a1ad-3459c8ad70bb · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f32a4cb9-4c3a-4245-bf2f-05710e46e5d8 · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension Uncovering Hidden Challenges in Query-Based Video Moment Retrieval
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52ba0e91-0b05-483b-a6b9-dbba00470632 · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7fd9c77-a4ce-4769-99bc-27b13b521333 · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7525b355-3d6c-414c-b3e2-48a18b4b3b0f · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension Timechat: A time-sensitive multimodal large lan- guage model for long video understanding
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation edd3e7d5-3664-4011-9117-43be9ee7776d · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension Self-Consistency Improves Chain of Thought Reasoning in Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ad9f57c-8738-44b4-9e96-6089277393f9 · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension HawkEye: Training Video-Text LLMs for Grounding Text in Videos
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c426dd8-6ad6-4685-a14d-c9b05c6a5b37 · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension Negative sample matters: A renaissance of metric learning for temporal grounding
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c0706820-e09c-44c9-b3e5-afab73f553b6 · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension Chain-of- thought prompting elicits reasoning in large language models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 97b95638-3e78-4d88-8d31-996da5028e42 · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension VideoQA in the Era of LLMs: An Empirical Study
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 574da5c1-99cd-4299-8ab0-9072c58b7e7d · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension Can i trust your answer? visually grounded video question answering
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f9ec1c82-fc35-44e0-b10d-bd132f8e91b3 · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension A closer look at temporal sentence ground- ing in videos: Dataset and metric
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7101815f-3ff4-4f8e-80eb-738370f61abe · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension Sc-tune: Unleashing self-consistent referential compre- hension in large vision language models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5bdc71a0-763f-44ab-8f84-70530c3739b9 · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension Span-based Localizing Network for Natural Language Video Localization
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27c0a6f0-4963-4354-98a4-d5c6d0850ad0 · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6964b318-5963-4117-aa1a-4d7bd5816eca · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension Learning 2d temporal adjacent networks for moment local- ization with natural language
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f458d21-c5b7-4dbf-a0ca-533ea3693a1b · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension Unveiling the Tapestry of Consistency in Large Vision-Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab874dcf-3866-4f0b-80a6-0a7d39a40a66 · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension Prompt Consistency for Zero-Shot Task Generalization
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6af2e0c8-f895-43dd-854b-1d6da6277b2e · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension Towards auto- matic learning of procedures from web instructional videos
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07b15573-2b1f-4fb0-b96f-1eb1ac93e7a7 · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fff885c0-7168-462e-bd0c-cd1a399c9171 · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension It shows a remarkable zero-shot audio understanding capability and also generates responses to the visual and audio information presented in the videos
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b40b6262-0a19-4173-b722-e609557a1753 · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension To do this, Video- LLaV A collects both image and video-text datasets and incorporates them in its instruction tuning
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a35b7f22-f46b-48f1-ae24-83cb385aaf2d · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension It introduces a new dataset for video instruction tuning, containing 100,000 high-quality video-instruction pairs
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 17f59f9c-44dc-49ae-827c-ec549a3eddba · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension Unresolved cited work
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3996aac8-ead6-49ad-879d-829a2fe2e155 · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension The format should be: ’start time - end seconds’
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a304e0bb-3c63-451d-a8fa-d2311d2200bc · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension The output format should be: ’start - end seconds’
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5cc8dc9c-75cd-4577-bfff-c5c5fa42642f · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension Specifically, they aim to align vision and text in the first stage and then generate captions from various image-text pairs
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 57a7ffd1-1db5-4dbe-bc4c-4e3989e145c7 · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension They seamlessly integrate both visual and audio modalities in videos and propose STC connector to understand spatiotemporal video informa- tion
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3bb54507-8220-42f6-8272-5a8da55627ff · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension Unresolved cited work
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 46f06485-429f-4d83-8e75-e4127754a740 · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension Unresolved cited work
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 041bec98-5982-4e8c-8da6-4fb868e61657 · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension I’m unable to find timestamps in the video
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 83407866-e070-4e37-98b1-af452de9d1ae · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension Experiments on Charades-CON with TimeChat
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 64614f41-df1e-4010-a7fd-abc914e6d17b · outbound
On the Consistency of Video Large Language Models in Temporal Comprehension 16 Figure 11
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
No inbound Pith citation observations are available.