Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T10:48:17.950215Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 1 inbound Pith citation observation for arXiv:2501.16786.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T10:48:17.950215Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-19T13:13:40.485342Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-19T13:17:18.568779Z
29 of 29 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 75b632ad-046b-49b4-8fcd-17c3a05b0fbf · outbound
Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73a91c07-9dad-4382-a05a-62699548d179 · outbound
Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 500db361-7377-474b-9f70-21366cad3cce · outbound
Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2bd6e7a-17bb-4522-8487-680d3125d74d · outbound
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 290033b8-cb34-45a7-b3af-d47a703105ba · outbound
Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f598d0c2-0a34-4639-9c67-ec2de9c01e16 · outbound
Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding PaLM-E: An Embodied Multimodal Language Model
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fae8d7b2-9d0d-43e9-b3e4-73f8e60ef4e0 · outbound
Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47d3c21a-0711-48a9-ba7d-54d3919f7cd2 · outbound
Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding Decoder-Only or Encoder-Decoder? Interpreting Language Model as a Regularized Encoder-Decoder
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d90a5ac2-0c5c-4817-baa2-6507bebe835e · outbound
Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding Mistral 7B
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7538c262-4576-4bdc-8932-834bf4b6d100 · outbound
Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding A diagram is worth a dozen images
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff23c886-4446-4908-a7b9-a8bf1dff7b4f · outbound
Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a01c820-c8a5-43b0-a3be-fa9537e63bad · outbound
Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e29ab880-dcb7-4151-805d-5fe5b1837272 · outbound
Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a168a47-981a-4c8c-b045-3b6e3d651477 · outbound
Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea803d1b-9603-41f8-960f-ee58bbba00e3 · outbound
Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 145a6331-b68b-437c-8687-a1e0c029a285 · outbound
Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce7ce885-3b72-4e58-853e-092787676375 · outbound
Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding Qwen2 Technical Report
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6916f56e-e1ed-48d5-a2c8-4beb7bebcedc · outbound
Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding A Survey on Multimodal Large Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43dccc6f-11bd-4bd9-8b36-83a012f1d99e · outbound
Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33be5698-4a49-4642-8d2b-1211d59d4c0c · outbound
Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec5a6765-7fa8-4da4-bca8-93e9a593e3ea · outbound
Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d3bcfc1-2645-4076-9670-406a32aaac5c · outbound
Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding LLaV A-OV-STE and LLaV A-Video-STE refer to LLaV A-OV-STE-3-(2:2) and LLaV A- Video-STE-3-(2:2), respectively.) A
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4d1c987c-b698-48d1-b4b9-03cd5fbcb579 · outbound
Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding LLaV A-Video-STE-1/2/3-(2:2) represents LLaV A-Video-STE using 1/2/3 layers of (2:2)
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation aab48173-787c-4e01-acf6-46778c48d019 · outbound
Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding LLaVA-OneVision: Easy Visual Task Transfer
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e007916-5c72-4fef-9a00-8bab48f022fa · outbound
Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f975545c-17ed-491f-9460-32ab7b9b94f4 · outbound
Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding Encoder vs Decoder: Comparative Analysis of Encoder and Decoder Language Models on Multilingual NLU Tasks
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4dc51d34-64a4-4a73-a5f3-bd76d8e8c094 · outbound
Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding Are We on the Right Way for Evaluating Large Vision-Language Models?
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8438e5c6-47a7-45db-a941-39bef0466c81 · outbound
Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ef87b41-155a-4009-a522-fa0e77db1a6c · outbound
Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f9d51a0-8931-4a44-990f-d86b12465850 · inbound
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.