Pith. sign in

Paper Citation Record · LEDGER

Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding

As of 10 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 1 inbound Pith citation observation for arXiv:2501.16786.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.16786 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T10:48:17.950215Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-19T13:13:40.485342Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T13:17:18.568779Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 75b632ad-046b-49b4-8fcd-17c3a05b0fbf · outbound

This paper cites GPT-4 Technical Report.

Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T10:48:17.758736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:48:17.758736Z digest=sha256:1150e23382a21cb78084078d7d92d10f212d62c5da8393e50a2b68efe49a4310

Observation 73a91c07-9dad-4382-a05a-62699548d179 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T10:48:17.776805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:48:17.776805Z digest=sha256:086e38e26f0db38aa2f71cadb9583ec726c3255b4de64d6bec2d66027bd13912

Observation 500db361-7377-474b-9f70-21366cad3cce · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T10:48:17.792593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:48:17.792593Z digest=sha256:d3333184d3019f8c63862ba698a8cbb643b3a88839a41807b28b7d37a5a05174

Observation f2bd6e7a-17bb-4522-8487-680d3125d74d · outbound

This paper cites E., et al.

Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding E., et al

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T10:48:17.798591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:48:17.798591Z digest=sha256:1eeb1d2d8abccc774b7d98b09765f5e1a2d19607d966bddb682356bd92467d00

Observation 290033b8-cb34-45a7-b3af-d47a703105ba · outbound

This paper cites InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model.

Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T10:48:17.803897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:48:17.803897Z digest=sha256:4b7114fe9eb0aa370e9c718bee8905f564b558ad555777b36161d19b0fc29ff5

Observation f598d0c2-0a34-4639-9c67-ec2de9c01e16 · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding PaLM-E: An Embodied Multimodal Language Model

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T10:48:17.809722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:48:17.809722Z digest=sha256:3837ed87bdffcc4c77cd1a4566841399379ac0885853493818d53c3ebcf0fe95

Observation fae8d7b2-9d0d-43e9-b3e4-73f8e60ef4e0 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T10:48:17.816075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:48:17.816075Z digest=sha256:fdcfde7c6bc717241a247deab4d83e356b7de580def25ef6bde3058a503cf8a9

Observation 47d3c21a-0711-48a9-ba7d-54d3919f7cd2 · outbound

This paper cites Decoder-Only or Encoder-Decoder? Interpreting Language Model as a Regularized Encoder-Decoder.

Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding Decoder-Only or Encoder-Decoder? Interpreting Language Model as a Regularized Encoder-Decoder

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T10:48:17.821361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:48:17.821361Z digest=sha256:be954d1de873e8615a7968949e2f130470e564ea5d16cc15eeccbb373b3760da

Observation d90a5ac2-0c5c-4817-baa2-6507bebe835e · outbound

This paper cites Mistral 7B.

Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding Mistral 7B

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T10:48:17.826834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:48:17.826834Z digest=sha256:c8ae38661985c9c42a0728ad28883187024e363704691015d85d62e4bc2941b7

Observation 7538c262-4576-4bdc-8932-834bf4b6d100 · outbound

This paper cites A diagram is worth a dozen images.

Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding A diagram is worth a dozen images

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T10:48:17.832020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:48:17.832020Z digest=sha256:0efae61c401a83e8878821285d45f9f040c7ceb0dc37616991a17d63e9ec8c71

Observation ff23c886-4446-4908-a7b9-a8bf1dff7b4f · outbound

This paper cites Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input.

Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T10:48:17.852698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:48:17.852698Z digest=sha256:9f142cab3b8c26b4ec8ada05e2f20344ad7178a67cecf5dd8a0caddcaa0ce45c

Observation 3a01c820-c8a5-43b0-a3be-fa9537e63bad · outbound

This paper cites Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al.

Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:48:18.628744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T10:48:17.868669Z digest=sha256:59f401784aff3172fd0ebf89a32f8d5cd50686fad0e3d1855e9c90e4f5eb9f11

Observation e29ab880-dcb7-4151-805d-5fe5b1837272 · outbound

This paper cites LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T10:48:17.877288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:48:17.877288Z digest=sha256:68758adbce502969a21cee9d9ea1b4023617b9148a2eae920383654b1294ae35

Observation 0a168a47-981a-4c8c-b045-3b6e3d651477 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T10:48:17.882899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:48:17.882899Z digest=sha256:accef8ab699ff7f98eb5749e3a91696550737e7720ed2f8afabca03a8f3153fd

Observation ea803d1b-9603-41f8-960f-ee58bbba00e3 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T10:48:17.888246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:48:17.888246Z digest=sha256:121bdefbcdde4fe01c534a888803021ffc4617c54d18f97a9fcd12968860019a

Observation 145a6331-b68b-437c-8687-a1e0c029a285 · outbound

This paper cites MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding.

Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T10:48:17.895385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:48:17.895385Z digest=sha256:234165f99347224e538bd8a83041050e321301247573c1c97085f3c8b30e38c6

Observation ce7ce885-3b72-4e58-853e-092787676375 · outbound

This paper cites Qwen2 Technical Report.

Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding Qwen2 Technical Report

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T10:48:17.912512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:48:17.912512Z digest=sha256:fb78d0e3153528027de129e54709557d31be0c721504f5b72e14f4a534da0895

Observation 6916f56e-e1ed-48d5-a2c8-4beb7bebcedc · outbound

This paper cites A Survey on Multimodal Large Language Models.

Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding A Survey on Multimodal Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T10:48:17.918447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:48:17.918447Z digest=sha256:5817e14dd446b6985fda579427f3e08f7f8c87c01a9808019d8f7f03a6537970

Observation 43dccc6f-11bd-4bd9-8b36-83a012f1d99e · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T10:48:17.923525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:48:17.923525Z digest=sha256:de56de571ce73ba5010e14ee43ad9712466e9a28f3103743d255db2ea7a62bb6

Observation 33be5698-4a49-4642-8d2b-1211d59d4c0c · outbound

This paper cites Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward.

Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T10:48:17.929538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:48:17.929538Z digest=sha256:8e4833913af6994d3344aaa071e189a222bb8e39f60360774f8b4378c61f3b9f

Observation ec5a6765-7fa8-4da4-bca8-93e9a593e3ea · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T10:48:17.936475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:48:17.936475Z digest=sha256:d731dadea8558a08c330ce89a192caec412655b627922fd5846105ec0b16cacc

Observation 1d3bcfc1-2645-4076-9670-406a32aaac5c · outbound

This paper cites LLaV A-OV-STE and LLaV A-Video-STE refer to LLaV A-OV-STE-3-(2:2) and LLaV A- Video-STE-3-(2:2), respectively.) A.

Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding LLaV A-OV-STE and LLaV A-Video-STE refer to LLaV A-OV-STE-3-(2:2) and LLaV A- Video-STE-3-(2:2), respectively.) A

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:48:18.608821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T10:48:17.942857Z digest=sha256:dddb88f185c19bd478891e5f8ce467ed4ae49de4e5a961e507f65f68f281080f

Observation 4d1c987c-b698-48d1-b4b9-03cd5fbcb579 · outbound

This paper cites LLaV A-Video-STE-1/2/3-(2:2) represents LLaV A-Video-STE using 1/2/3 layers of (2:2).

Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding LLaV A-Video-STE-1/2/3-(2:2) represents LLaV A-Video-STE using 1/2/3 layers of (2:2)

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:48:18.582282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T10:48:17.950215Z digest=sha256:ba9fdf09cba1d21a7f533091659bb96c163598ca8eb9a28de849be2d1e209e0b

Observation aab48173-787c-4e01-acf6-46778c48d019 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-10T10:48:17.837710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:48:17.837710Z digest=sha256:0c4ff9389129d950f04ba2051a95ad74b104a5c071cae945571ca1d8bb6cc61c

Observation 1e007916-5c72-4fef-9a00-8bab48f022fa · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-10T10:48:17.902301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:48:17.902301Z digest=sha256:2f19a9bcfc912d88f814a0d86c4f7753761d45406d56a8e3574c85210736fcce

Observation f975545c-17ed-491f-9460-32ab7b9b94f4 · outbound

This paper cites Encoder vs Decoder: Comparative Analysis of Encoder and Decoder Language Models on Multilingual NLU Tasks.

Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding Encoder vs Decoder: Comparative Analysis of Encoder and Decoder Language Models on Multilingual NLU Tasks

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-10T10:48:17.862854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:48:17.862854Z digest=sha256:f85aa5bac88a800757a3597d2b7db3e2bdcb153fb22e94e28824278933fb4d14

Observation 4dc51d34-64a4-4a73-a5f3-bd76d8e8c094 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T10:48:17.784123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:48:17.784123Z digest=sha256:9a2d2180c316e5bb6e561e4f2628a235d65c0795b4e521cde68b313aa39c5ffe

Observation 8438e5c6-47a7-45db-a941-39bef0466c81 · outbound

This paper cites MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens.

Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T10:48:17.769194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:48:17.769194Z digest=sha256:16de64dcf9488cb1183fea399d454869fd80b5f87035ed4525ac71dc047f80d8

Observation 9ef87b41-155a-4009-a522-fa0e77db1a6c · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-10T10:48:17.846752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:48:17.846752Z digest=sha256:8577600c2197bcc8341af89a1783b89c72fb03d3e57deb2ff729d6a2131b4c04

Pith citing papers

Observation 1f9d51a0-8931-4a44-990f-d86b12465850 · inbound

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding cites this paper.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:17:18.571134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:e7ff1a68d50fd94791d6d142693888f6cf075564f2f0e48cd57de52db5abb254