Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:32:55.846636Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 5 inbound Pith citation observations for arXiv:2505.24329.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:32:55.846636Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-11T18:39:43.071937Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-16T08:40:46.400243Z
59 of 59 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e983d84c-ad2e-4741-9dd1-b6b1a56c401b · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 046fa311-fe12-4a19-b887-0a6fd01b6ce5 · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models Frozen in time: A joint video and image encoder for end-to-end retrieval
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e150a606-02fb-4a4e-8b6f-c23c42111155 · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 866881c4-15df-4a8a-b47b-10bf6d8e8030 · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 538b5115-2b18-4501-8090-17fa6c5d4c27 · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f47d592-0777-426d-a89d-ac41ca6148e3 · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af726b0c-8c59-43c4-b747-5ff92d29d4e0 · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7fee8cdc-d608-4d80-879c-706639178c51 · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 872d5253-b461-4aaa-a439-6770b4e4f72a · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models Tall: Temporal activity localization via language query
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e2bdee07-8166-41f5-a9a7-78f04a2aaf8d · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models LinVT: Empower Your Image-level Large Language Model to Understand Videos
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d62473d-8e26-4913-b3c3-cf1625e38c8f · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models Saliency-guided detr for mo- ment retrieval and highlight detection
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a097448b-0aec-4dac-9ec7-7ab4de3fd9dd · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models Ego4d: Around the world in 3,000 hours of egocentric video
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 4a7dd681-5851-4269-98a3-7d8710374cc3 · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models VTG-LLM: Integrating Timestamp Knowledge into Video LLMs for Enhanced Video Temporal Grounding
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46d40a68-0663-4594-a67f-c871ece1aa77 · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models Lora: Low-rank adaptation of large language models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ad03b8fa-587b-4f00-b849-d94343e2cd88 · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models Vtimellm: Empower llm to grasp video moments
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 99209f1d-808b-46a7-b8ad-cc1eda223349 · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models Dense-captioning events in videos
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0516f95-5d10-48ae-b0aa-9690bfc48e09 · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models Detecting mo- ments and highlights in videos via natural language queries
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ca54fab8-e018-4a59-a862-6e7ff182cc8e · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models LLaVA-OneVision: Easy Visual Task Transfer
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6a31558-7800-4065-9018-751a255c5cb3 · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 21d7e694-12db-4d82-a923-8fecfdde71d9 · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models VideoChat: Chat-Centric Video Understanding
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6db8b069-5dc4-4c7b-b4e0-4830daab98cf · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models Mvbench: A comprehensive multi-modal video understand- ing benchmark
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1c5b0dde-dbb6-429c-8e15-e36514292473 · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models Generalized focal loss: Learning qualified and distributed bounding boxes for dense 9 object detection
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 236837cc-17ff-4c9b-845d-37a95944f148 · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models Llama-vid: An image is worth 2 tokens in large language models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4644ef79-c8e4-4276-aac2-14c94566a4e0 · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models GroundingGPT:Language Enhanced Multi-modal Grounding Model
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bffae752-2bbc-48d1-afa2-3ff47b1f1019 · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models Detal: Open-vocabulary temporal action localization with decoupled networks
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation de9c4524-36a8-459d-9684-22869a55e3dc · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models Univtg: Towards unified video- language temporal grounding
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 09668194-d388-476e-b1b1-3a5cdb596534 · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models Visual instruction tuning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 039de858-bae8-427b-a797-8f0160f23146 · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models E.T. Bench: Towards Open-Ended Event-Level Video-Language Understanding
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d2a4862-2b69-4d78-8e00-51a1e96f90cb · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models Decoupled Weight Decay Regularization
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8c5eb12-65ec-487d-8628-b0f2bed9b148 · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1b43079-4c42-428b-a47a-aff1f18e0f89 · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4941de7b-9f8d-4784-aaf8-718110310073 · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models The surprising effectiveness of multimodal large language models for video moment retrieval
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfe258fe-aad4-4e4b-8b4d-450c4bd1da09 · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models Correlation-Guided Query-Dependency Calibration for Video Temporal Grounding
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d5c674d-ccd7-4949-80ac-ace48b687509 · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models Perceptiongpt: Effectively fusing visual perception into llm
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f4567e2-5e6e-4036-9683-4241c6b3734d · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37c0725f-863e-457e-bda2-48f82bcbb0eb · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models Chatvtg: Video temporal grounding via chat with video dialogue large language models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation be224410-daab-4624-ada5-1d69663abbef · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models Timechat: A time-sensitive multimodal large lan- guage model for long video understanding
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation dae1cfa4-b820-47e8-a865-124eaad7325d · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3caff01c-d124-4cdf-9fe6-4f42d9280603 · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models React: Temporal action detection with relational queries
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 551e7c59-7076-4afa-91cd-d51c2096a59b · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models Temporal Action Localization with Enhanced Instant Discriminability
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5128f6d0-93b9-45d2-abb8-e424557f2f22 · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models Tridet: Temporal action detection with relative boundary modeling
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 213c9c4c-bd20-43cf-8c8d-d90308fb481e · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1569b71e-42fe-4991-9cb7-2e6b70e55ea5 · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models Internvideo2: Scaling foundation models for mul- timodal video understanding
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ee4104bc-0d5b-416d-9007-440cd0f92730 · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee2793be-62f1-48b5-964c-1c8c4c7b02b7 · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models Longvideobench: A benchmark for long-context interleaved video-language understanding
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 0bd14b92-e188-48bd-af3a-0d310dfc8100 · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models Can i trust your answer? visually grounded video question answering
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e29da4d6-2586-4b7a-b153-79b9c7f0bd1f · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4afb9bc-f957-4eac-9ddd-24f633e3724d · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models Zero-shot video question answering via 10 frozen bidirectional language models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a1c6c84f-88c8-4d7f-876c-e230f0f6b518 · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bce61b1b-d678-4b3d-b4b9-22fe8ddce6b3 · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models Self-chained image-language model for video localization and question answering
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 044d7d47-c453-4bd2-b160-9d8f5f843d4e · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models Activitynet-qa: A dataset for understanding complex web videos via question answering
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8466171-b24c-4362-a5f6-34e316667057 · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models Unimd: Towards unifying moment retrieval and temporal ac- tion detection
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 53e72cc7-89e6-494b-a7da-65fab8cf662d · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01593c2e-f8b1-48d3-99d3-1f3a35340e4f · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models LD-DETR: Loop Decoder DEtection TRansformer for Video Moment Retrieval and Highlight Detection
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4113345c-9176-4ccf-b553-d34527587dff · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models Training-free video temporal grounding using large-scale pre-trained models
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5b333f8d-3dec-4341-a7f5-6a4bb8e52b42 · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models Towards automatic learning of procedures from web instructional videos
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6e564994-cdc0-40c3-8e92-e5c81178fd94 · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be464263-fe35-4f22-9094-5c98ea8d6d8d · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models Dedicated
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8b2765c7-1c23-42b5-ba1f-7890ff6a0025 · outbound
DisTime: Distribution-based Time Representation for Video Large Language Models Give you the textual query: ‘thereis an orange barrier out of which people can stand’
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ec3e9b74-62dc-400b-add3-28e3b82cbe76 · inbound
RoboTron-Drive: All-in-One Large Multimodal Model for Autonomous Driving DisTime: Distribution-based Time Representation for Video Large Language Models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48cf3857-f29c-42f7-8e63-d7dbd67fd0c2 · inbound
Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation DisTime: Distribution-based Time Representation for Video Large Language Models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 4944cbca-33ab-44b1-a708-1985c1e860aa · inbound
Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation DisTime: Distribution-based Time Representation for Video Large Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f28223d-7405-4cd7-83c5-6a4824473fa1 · inbound
OmniVTG: A Large-Scale Dataset and Training Paradigm for Open-World Video Temporal Grounding DisTime: Distribution-based Time Representation for Video Large Language Models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation fb0229f2-85fa-4610-ba82-d25d5f94f700 · inbound
TimePLE: Rethinking Temporal Representation for Video Temporal Grounding DisTime: Distribution-based Time Representation for Video Large Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.