Pith. sign in

Paper Citation Record · LEDGER

DisTime: Distribution-based Time Representation for Video Large Language Models

As of 15 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 5 inbound Pith citation observations for arXiv:2505.24329.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24329 v2

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:32:55.846636Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T18:39:43.071937Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T08:40:46.400243Z

Reference resolution

59 of 59 outbound references displayed

  • verified exact0
  • verified fuzzy24
  • unresolved34
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e983d84c-ad2e-4741-9dd1-b6b1a56c401b · outbound

This paper cites GPT-4 Technical Report.

DisTime: Distribution-based Time Representation for Video Large Language Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:49.930599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:49.930599Z digest=sha256:5f59f130e1001287be3a13d93c65b223c5dc392fa78cadfa6137fdb86e45facb

Observation 046fa311-fe12-4a19-b887-0a6fd01b6ce5 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

DisTime: Distribution-based Time Representation for Video Large Language Models Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:50.034969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:50.034969Z digest=sha256:67d06b2cda842476d269fa56489d528d6fe22017dedb704f62ba51c2f895d8dd

Observation e150a606-02fb-4a4e-8b6f-c23c42111155 · outbound

This paper cites CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding.

DisTime: Distribution-based Time Representation for Video Large Language Models CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:50.149539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:50.149539Z digest=sha256:3af0222aceb35a424952059a0a41daf61553d6788765fccc093abc79c89c751a

Observation 866881c4-15df-4a8a-b47b-10bf6d8e8030 · outbound

This paper cites TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability.

DisTime: Distribution-based Time Representation for Video Large Language Models TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:50.318296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:50.318296Z digest=sha256:6cbad6498125c32c71eb46a63f479107c36d45ed295c1c24e3f0f8c098c323bd

Observation 538b5115-2b18-4501-8090-17fa6c5d4c27 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

DisTime: Distribution-based Time Representation for Video Large Language Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:50.466205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:50.466205Z digest=sha256:206bd85ec048f4d8211db780638c5ca5c14050983e9374f2746ae7ab3eb2e7fd

Observation 5f47d592-0777-426d-a89d-ac41ca6148e3 · outbound

This paper cites How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.

DisTime: Distribution-based Time Representation for Video Large Language Models How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:50.609800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:50.609800Z digest=sha256:f1f611d153b49d43bbef1e066d7d5dc00b19a7ac82c79584a042a6a231b2577e

Observation af726b0c-8c59-43c4-b747-5ff92d29d4e0 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.

DisTime: Distribution-based Time Representation for Video Large Language Models Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:33:00.717368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:32:50.706528Z digest=sha256:43fc259a37727c1eb41893d4aed2bcbbdfe45f3a384151ce5117e0d071d78db2

Observation 7fee8cdc-d608-4d80-879c-706639178c51 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

DisTime: Distribution-based Time Representation for Video Large Language Models Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:50.808047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:50.808047Z digest=sha256:459ecd1a5fea80c4d4ab56216537baedfa46ac38c940f4a8e482cab55e1b16b0

Observation 872d5253-b461-4aaa-a439-6770b4e4f72a · outbound

This paper cites Tall: Temporal activity localization via language query.

DisTime: Distribution-based Time Representation for Video Large Language Models Tall: Temporal activity localization via language query

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:33:00.601125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:32:50.911068Z digest=sha256:1d2837f486437aa0bc14af062e9474a088bad2a8bdcb94e3c71d08f383ed0a62

Observation e2bdee07-8166-41f5-a9a7-78f04a2aaf8d · outbound

This paper cites LinVT: Empower Your Image-level Large Language Model to Understand Videos.

DisTime: Distribution-based Time Representation for Video Large Language Models LinVT: Empower Your Image-level Large Language Model to Understand Videos

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:51.011685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:51.011685Z digest=sha256:5dc26be1b349f46a05c4a2c6c8aba15ca122c86dee81c690f44cb28db4207395

Observation 8d62473d-8e26-4913-b3c3-cf1625e38c8f · outbound

This paper cites Saliency-guided detr for mo- ment retrieval and highlight detection.

DisTime: Distribution-based Time Representation for Video Large Language Models Saliency-guided detr for mo- ment retrieval and highlight detection

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:51.110427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:51.110427Z digest=sha256:100599ac95e5ac512dea69985293890290b93eb55f3eaf639d2c0f68bfb07a61

Observation a097448b-0aec-4dac-9ec7-7ab4de3fd9dd · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

DisTime: Distribution-based Time Representation for Video Large Language Models Ego4d: Around the world in 3,000 hours of egocentric video

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:33:00.485258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:32:51.175073Z digest=sha256:8d618b6808f630715b6a0ecdb55b166d34172e56d9f158fa6b55cb4c1d3bcdd6

Observation 4a7dd681-5851-4269-98a3-7d8710374cc3 · outbound

This paper cites VTG-LLM: Integrating Timestamp Knowledge into Video LLMs for Enhanced Video Temporal Grounding.

DisTime: Distribution-based Time Representation for Video Large Language Models VTG-LLM: Integrating Timestamp Knowledge into Video LLMs for Enhanced Video Temporal Grounding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:51.273965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:51.273965Z digest=sha256:8375bee1c2da5ba334909d01bac73e851578c62a75278b9c5ae56b14804aab08

Observation 46d40a68-0663-4594-a67f-c871ece1aa77 · outbound

This paper cites Lora: Low-rank adaptation of large language models.

DisTime: Distribution-based Time Representation for Video Large Language Models Lora: Low-rank adaptation of large language models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:33:00.364833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:32:51.368935Z digest=sha256:97495e79da052a22709e8569b2acd443f9699be15313485bfe40b537b04ab283

Observation ad03b8fa-587b-4f00-b849-d94343e2cd88 · outbound

This paper cites Vtimellm: Empower llm to grasp video moments.

DisTime: Distribution-based Time Representation for Video Large Language Models Vtimellm: Empower llm to grasp video moments

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:33:00.224325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:32:51.473579Z digest=sha256:66ec423bd307de1f07c84ea3abd097f1f8cbe55535a34f4a799f0702d9ba23e3

Observation 99209f1d-808b-46a7-b8ad-cc1eda223349 · outbound

This paper cites Dense-captioning events in videos.

DisTime: Distribution-based Time Representation for Video Large Language Models Dense-captioning events in videos

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:51.534075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:51.534075Z digest=sha256:5f2ae6615332f1f27b297f73b8769de648bd2eb5b7bb62be484444af128f5c56

Observation e0516f95-5d10-48ae-b0aa-9690bfc48e09 · outbound

This paper cites Detecting mo- ments and highlights in videos via natural language queries.

DisTime: Distribution-based Time Representation for Video Large Language Models Detecting mo- ments and highlights in videos via natural language queries

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:33:00.059185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:32:51.628353Z digest=sha256:41339f29394f3947977197b9e167dfe99b8e8528f84b15399379cc9cea7452d4

Observation ca54fab8-e018-4a59-a862-6e7ff182cc8e · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

DisTime: Distribution-based Time Representation for Video Large Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:51.722537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:51.722537Z digest=sha256:db179ec1d53756dc1ffd43ec74cdb9b31a2737f7d4b2b4261d0dc02d1b04482e

Observation c6a31558-7800-4065-9018-751a255c5cb3 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

DisTime: Distribution-based Time Representation for Video Large Language Models Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:59.896873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:32:51.820610Z digest=sha256:a2d6e3238214c103c08bcd643da44b0bbd3c1c4369db95df64fa383e4275d1a9

Observation 21d7e694-12db-4d82-a923-8fecfdde71d9 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

DisTime: Distribution-based Time Representation for Video Large Language Models VideoChat: Chat-Centric Video Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:51.944782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:51.944782Z digest=sha256:4283a82c044f919e832700ab30ba02f0811b9575eb065e05a0df6f05715ac6f8

Observation 6db8b069-5dc4-4c7b-b4e0-4830daab98cf · outbound

This paper cites Mvbench: A comprehensive multi-modal video understand- ing benchmark.

DisTime: Distribution-based Time Representation for Video Large Language Models Mvbench: A comprehensive multi-modal video understand- ing benchmark

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:59.732986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:32:52.015588Z digest=sha256:2aea14344b96a8adb76672b39593e99402369d2d558044dfcb207b3493a6497a

Observation 1c5b0dde-dbb6-429c-8e15-e36514292473 · outbound

This paper cites Generalized focal loss: Learning qualified and distributed bounding boxes for dense 9 object detection.

DisTime: Distribution-based Time Representation for Video Large Language Models Generalized focal loss: Learning qualified and distributed bounding boxes for dense 9 object detection

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:59.550281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:32:52.100242Z digest=sha256:3b767c0a50d8087acb5bafed056de8c989939e36c0470eb46477d0f2329d7ac5

Observation 236837cc-17ff-4c9b-845d-37a95944f148 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

DisTime: Distribution-based Time Representation for Video Large Language Models Llama-vid: An image is worth 2 tokens in large language models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:52.213019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:52.213019Z digest=sha256:00b7591c61400344dbae283bc5d30ee6d5781e07cb121bff763fd8135a814d2b

Observation 4644ef79-c8e4-4276-aac2-14c94566a4e0 · outbound

This paper cites GroundingGPT:Language Enhanced Multi-modal Grounding Model.

DisTime: Distribution-based Time Representation for Video Large Language Models GroundingGPT:Language Enhanced Multi-modal Grounding Model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:52.287195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:52.287195Z digest=sha256:27f6e9602f515c9fab7ca7fab6365fc6c9510ba1054ed9e4ae0fb4377f30f167

Observation bffae752-2bbc-48d1-afa2-3ff47b1f1019 · outbound

This paper cites Detal: Open-vocabulary temporal action localization with decoupled networks.

DisTime: Distribution-based Time Representation for Video Large Language Models Detal: Open-vocabulary temporal action localization with decoupled networks

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:59.375299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:32:52.373305Z digest=sha256:7dce0b60f2953705ab8f3a2df3021662c6fe438bedafca9a9854242cf4ac2bee

Observation de9c4524-36a8-459d-9684-22869a55e3dc · outbound

This paper cites Univtg: Towards unified video- language temporal grounding.

DisTime: Distribution-based Time Representation for Video Large Language Models Univtg: Towards unified video- language temporal grounding

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:59.190583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:32:52.493479Z digest=sha256:7b20e2727516a12e977268134c570097297c056ec054921a1d96d215ca00cd9a

Observation 09668194-d388-476e-b1b1-3a5cdb596534 · outbound

This paper cites Visual instruction tuning.

DisTime: Distribution-based Time Representation for Video Large Language Models Visual instruction tuning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:52.552107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:52.552107Z digest=sha256:0b92b5338944b456c0901706e3527dce5d0e113271f9dde9e26c8f34f4c10761

Observation 039de858-bae8-427b-a797-8f0160f23146 · outbound

This paper cites E.T. Bench: Towards Open-Ended Event-Level Video-Language Understanding.

DisTime: Distribution-based Time Representation for Video Large Language Models E.T. Bench: Towards Open-Ended Event-Level Video-Language Understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:52.655077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:52.655077Z digest=sha256:24b1cbf6ec30df8762a09aaef10d84403d9e07ddc782729f54a5f98925fde001

Observation 3d2a4862-2b69-4d78-8e00-51a1e96f90cb · outbound

This paper cites Decoupled Weight Decay Regularization.

DisTime: Distribution-based Time Representation for Video Large Language Models Decoupled Weight Decay Regularization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:52.766856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:52.766856Z digest=sha256:1bd9489518d5031d88154b21f6e2ca22979a19031dbf691b603f56fa2952a9ea

Observation a8c5eb12-65ec-487d-8628-b0f2bed9b148 · outbound

This paper cites LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval.

DisTime: Distribution-based Time Representation for Video Large Language Models LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:52.859271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:52.859271Z digest=sha256:16d7a6b7ae2cf409aff0921f6862e60778793e655f6c27533f71a7d8e5583828

Observation c1b43079-4c42-428b-a47a-aff1f18e0f89 · outbound

This paper cites Valley: Video Assistant with Large Language model Enhanced abilitY.

DisTime: Distribution-based Time Representation for Video Large Language Models Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:52.929275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:52.929275Z digest=sha256:93143639f535a6afb6858a8d5b6f542b15c45f72b11dcfd5d0e52cd8b79940c9

Observation 4941de7b-9f8d-4784-aaf8-718110310073 · outbound

This paper cites The surprising effectiveness of multimodal large language models for video moment retrieval.

DisTime: Distribution-based Time Representation for Video Large Language Models The surprising effectiveness of multimodal large language models for video moment retrieval

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:53.040750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:53.040750Z digest=sha256:3bd4f2f154e6f811da015ac902c155e60c4fcd6533a114e69fb3259f8ef6c9e7

Observation dfe258fe-aad4-4e4b-8b4d-450c4bd1da09 · outbound

This paper cites Correlation-Guided Query-Dependency Calibration for Video Temporal Grounding.

DisTime: Distribution-based Time Representation for Video Large Language Models Correlation-Guided Query-Dependency Calibration for Video Temporal Grounding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:53.147022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:53.147022Z digest=sha256:2fb9915cf962f1eb11213821a8f4a48105f01262df11a9fa630e86a8c2a8fabd

Observation 1d5c674d-ccd7-4949-80ac-ace48b687509 · outbound

This paper cites Perceptiongpt: Effectively fusing visual perception into llm.

DisTime: Distribution-based Time Representation for Video Large Language Models Perceptiongpt: Effectively fusing visual perception into llm

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:53.284040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:53.284040Z digest=sha256:d8a18276987e8ae73596d5a939c0c82fe9705754219f114a3341a19c009cfb52

Observation 0f4567e2-5e6e-4036-9683-4241c6b3734d · outbound

This paper cites Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning.

DisTime: Distribution-based Time Representation for Video Large Language Models Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:53.366366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:53.366366Z digest=sha256:c0975749b70865837f4e96ddd0b3be8053a7d5336554eeb0dcaac2a2c0425c35

Observation 37c0725f-863e-457e-bda2-48f82bcbb0eb · outbound

This paper cites Chatvtg: Video temporal grounding via chat with video dialogue large language models.

DisTime: Distribution-based Time Representation for Video Large Language Models Chatvtg: Video temporal grounding via chat with video dialogue large language models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:58.946972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:32:53.420330Z digest=sha256:e00e1db98a99638d7863ec0e7ae005df97519867fa44bb5e6cdda6bd582d0137

Observation be224410-daab-4624-ada5-1d69663abbef · outbound

This paper cites Timechat: A time-sensitive multimodal large lan- guage model for long video understanding.

DisTime: Distribution-based Time Representation for Video Large Language Models Timechat: A time-sensitive multimodal large lan- guage model for long video understanding

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:58.734554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:32:53.525437Z digest=sha256:3d1da72adb982a34f86c2fb9acd62c84f976670133ea0cd22681961c34b75b84

Observation dae1cfa4-b820-47e8-a865-124eaad7325d · outbound

This paper cites xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs.

DisTime: Distribution-based Time Representation for Video Large Language Models xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:53.658205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:53.658205Z digest=sha256:032fbb5e938d9f27ccc5d48203c26cfd6c5b8e264a538f3f3855473f40248640

Observation 3caff01c-d124-4cdf-9fe6-4f42d9280603 · outbound

This paper cites React: Temporal action detection with relational queries.

DisTime: Distribution-based Time Representation for Video Large Language Models React: Temporal action detection with relational queries

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:58.506942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:32:53.768815Z digest=sha256:eba1e0adfa3c4f8f1aacc4cb6f0e3f19369b43e24dfc3ea6281501bf7d900379

Observation 551e7c59-7076-4afa-91cd-d51c2096a59b · outbound

This paper cites Temporal Action Localization with Enhanced Instant Discriminability.

DisTime: Distribution-based Time Representation for Video Large Language Models Temporal Action Localization with Enhanced Instant Discriminability

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:53.865330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:53.865330Z digest=sha256:524eb697b30cdc8af76530af96122e7b293a4aa62175f500107ff8f155a41418

Observation 5128f6d0-93b9-45d2-abb8-e424557f2f22 · outbound

This paper cites Tridet: Temporal action detection with relative boundary modeling.

DisTime: Distribution-based Time Representation for Video Large Language Models Tridet: Temporal action detection with relative boundary modeling

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:58.335414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:32:53.939553Z digest=sha256:c2ef4aaec7baf0484343f66adc4c6db880f762690b5f42a8451cf63c0bec466f

Observation 213c9c4c-bd20-43cf-8c8d-d90308fb481e · outbound

This paper cites InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation.

DisTime: Distribution-based Time Representation for Video Large Language Models InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:54.042453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:54.042453Z digest=sha256:7ccafe254b6b24454c785bac1ee0200a81a8b562ede428fc4e2c700bd46aabfa

Observation 1569b71e-42fe-4991-9cb7-2e6b70e55ea5 · outbound

This paper cites Internvideo2: Scaling foundation models for mul- timodal video understanding.

DisTime: Distribution-based Time Representation for Video Large Language Models Internvideo2: Scaling foundation models for mul- timodal video understanding

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:58.125529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:32:54.145196Z digest=sha256:637393be40f82fbc5d6b475be668f585e893b002c0542e9d417d6068ae4bed10

Observation ee4104bc-0d5b-416d-9007-440cd0f92730 · outbound

This paper cites InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling.

DisTime: Distribution-based Time Representation for Video Large Language Models InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:54.252300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:54.252300Z digest=sha256:2264b2b15ea6932fae89971d724ba858f84d7d0fdac118f02c033518c9e185dc

Observation ee2793be-62f1-48b5-964c-1c8c4c7b02b7 · outbound

This paper cites Longvideobench: A benchmark for long-context interleaved video-language understanding.

DisTime: Distribution-based Time Representation for Video Large Language Models Longvideobench: A benchmark for long-context interleaved video-language understanding

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:57.949737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:32:54.341102Z digest=sha256:255af434be8a5afb8ac075248d0321ee0bb10383bf12878ab5667a58a9998ed2

Observation 0bd14b92-e188-48bd-af3a-0d310dfc8100 · outbound

This paper cites Can i trust your answer? visually grounded video question answering.

DisTime: Distribution-based Time Representation for Video Large Language Models Can i trust your answer? visually grounded video question answering

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:57.761265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:32:54.441110Z digest=sha256:ecd15f9417c3d87ecd038f818d2cb074904b44f10bd867d1f0afd90e87ad6caa

Observation e29da4d6-2586-4b7a-b153-79b9c7f0bd1f · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

DisTime: Distribution-based Time Representation for Video Large Language Models PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:54.522637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:54.522637Z digest=sha256:cabfe9b44b242ccfa77efb4ebb613bae28b39e48d726c8014a4da062fba290aa

Observation c4afb9bc-f957-4eac-9ddd-24f633e3724d · outbound

This paper cites Zero-shot video question answering via 10 frozen bidirectional language models.

DisTime: Distribution-based Time Representation for Video Large Language Models Zero-shot video question answering via 10 frozen bidirectional language models

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:57.564606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:32:54.633174Z digest=sha256:65ace96bb3a8e90608dfa1ce86092fe9ffbab9cb06202698092121f3c9646ba5

Observation a1c6c84f-88c8-4d7f-876c-e230f0f6b518 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

DisTime: Distribution-based Time Representation for Video Large Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:54.728045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:54.728045Z digest=sha256:5c6bc5f703cf51f85e684cdd44dbd7b2a51566a7d4d1e2793b560df12ea27f79

Observation bce61b1b-d678-4b3d-b4b9-22fe8ddce6b3 · outbound

This paper cites Self-chained image-language model for video localization and question answering.

DisTime: Distribution-based Time Representation for Video Large Language Models Self-chained image-language model for video localization and question answering

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:57.378825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:32:54.841533Z digest=sha256:f2283d6bf9d3ca30c5cf03993591c56a30d9fb807d896a244881c3760fd18298

Observation 044d7d47-c453-4bd2-b160-9d8f5f843d4e · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

DisTime: Distribution-based Time Representation for Video Large Language Models Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:54.974321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:54.974321Z digest=sha256:97158c48096d50f3ddbc173aef8afd5be0f4259bb9109ddd2c66b6818e426999

Observation c8466171-b24c-4362-a5f6-34e316667057 · outbound

This paper cites Unimd: Towards unifying moment retrieval and temporal ac- tion detection.

DisTime: Distribution-based Time Representation for Video Large Language Models Unimd: Towards unifying moment retrieval and temporal ac- tion detection

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:57.232076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:32:55.048733Z digest=sha256:5884d4a4b9113493ab2a305494c66357a6641324bdf004e042e442567ea342b5

Observation 53e72cc7-89e6-494b-a7da-65fab8cf662d · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

DisTime: Distribution-based Time Representation for Video Large Language Models Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:55.159633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:55.159633Z digest=sha256:158f7984cd2179a53412c0288409e733730175ab16ca25bbcb9ee0fca7ded220

Observation 01593c2e-f8b1-48d3-99d3-1f3a35340e4f · outbound

This paper cites LD-DETR: Loop Decoder DEtection TRansformer for Video Moment Retrieval and Highlight Detection.

DisTime: Distribution-based Time Representation for Video Large Language Models LD-DETR: Loop Decoder DEtection TRansformer for Video Moment Retrieval and Highlight Detection

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:55.288998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:55.288998Z digest=sha256:1dac8384de2355edb56d3ae5c4ec455cbf44d57b8ea58ea2c63b6dee65227150

Observation 4113345c-9176-4ccf-b553-d34527587dff · outbound

This paper cites Training-free video temporal grounding using large-scale pre-trained models.

DisTime: Distribution-based Time Representation for Video Large Language Models Training-free video temporal grounding using large-scale pre-trained models

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:57.043893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:32:55.399042Z digest=sha256:c291d27602593b5e486fa5902b6c86354aa6a1a35e7b648a80cfa55fbba7c2ff

Observation 5b333f8d-3dec-4341-a7f5-6a4bb8e52b42 · outbound

This paper cites Towards automatic learning of procedures from web instructional videos.

DisTime: Distribution-based Time Representation for Video Large Language Models Towards automatic learning of procedures from web instructional videos

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:56.872197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:32:55.508617Z digest=sha256:549a30e94529939f680ee84b8b6fc50b71866c61f53e32a8d78ddb17b01c1940

Observation 6e564994-cdc0-40c3-8e92-e5c81178fd94 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

DisTime: Distribution-based Time Representation for Video Large Language Models MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:55.604775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:55.604775Z digest=sha256:2bcec2f8eb74ee97c70d72344827b3e921de632bb4f495f4f9a313924214147c

Observation be464263-fe35-4f22-9094-5c98ea8d6d8d · outbound

This paper cites Dedicated.

DisTime: Distribution-based Time Representation for Video Large Language Models Dedicated

Reference 58

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T12:32:56.674125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:32:55.709982Z digest=sha256:343d678444e38107212edd5ec10ce6b0762a23f4d675ef3b85582d85969fb0a5

Observation 8b2765c7-1c23-42b5-ba1f-7890ff6a0025 · outbound

This paper cites Give you the textual query: ‘thereis an orange barrier out of which people can stand’.

DisTime: Distribution-based Time Representation for Video Large Language Models Give you the textual query: ‘thereis an orange barrier out of which people can stand’

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:56.490871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:32:55.846636Z digest=sha256:23aff1ca9d4cf35d3c0fa0f055b7271f31d195593e4dbc7772cc12543ea9d421

Pith citing papers

Observation ec3e9b74-62dc-400b-add3-28e3b82cbe76 · inbound

RoboTron-Drive: All-in-One Large Multimodal Model for Autonomous Driving cites this paper.

RoboTron-Drive: All-in-One Large Multimodal Model for Autonomous Driving DisTime: Distribution-based Time Representation for Video Large Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T18:39:43.071937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:39:43.071937Z digest=sha256:67b45d914790ae905c428d6b9d2211db4d47c11a8be894d7548c96dd1cef4849

Observation 48cf3857-f29c-42f7-8e63-d7dbd67fd0c2 · inbound

Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation cites this paper.

Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation DisTime: Distribution-based Time Representation for Video Large Language Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:40:46.401909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T08:38:49.075457Z digest=sha256:72969b3b10c0fde4febee91a9cb04f9efaa9900169d502e07fadbc2f38512bdb

Observation 4944cbca-33ab-44b1-a708-1985c1e860aa · inbound

Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation cites this paper.

Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation DisTime: Distribution-based Time Representation for Video Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T05:14:24.814323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:14:24.814323Z digest=sha256:ae5f11bb0a2e3e2bf47637b90289913b5a05edc6778e1ffcd5579dcdf3f7b0c7

Observation 1f28223d-7405-4cd7-83c5-6a4824473fa1 · inbound

OmniVTG: A Large-Scale Dataset and Training Paradigm for Open-World Video Temporal Grounding cites this paper.

OmniVTG: A Large-Scale Dataset and Training Paradigm for Open-World Video Temporal Grounding DisTime: Distribution-based Time Representation for Video Large Language Models

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:31:15.758838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-07T16:51:43.783133Z digest=sha256:23e9b01d3097b99252c80233be4f91907aaa683eaab55dc8023171fc1a3dd491

Observation fb0229f2-85fa-4610-ba82-d25d5f94f700 · inbound

TimePLE: Rethinking Temporal Representation for Video Temporal Grounding cites this paper.

TimePLE: Rethinking Temporal Representation for Video Temporal Grounding DisTime: Distribution-based Time Representation for Video Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-31T23:32:54.788326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:32:54.788326Z digest=sha256:7c1669b96e521d0c4e1b024362d5e4022b1356a7971c84bdf13dff57b96d12b3