Pith. sign in

Paper Citation Record · LEDGER

On the Consistency of Video Large Language Models in Temporal Comprehension

As of 14 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 0 inbound Pith citation observations for arXiv:2411.12951.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.12951 v2

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T17:07:00.966961Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

59 of 59 outbound references displayed

  • verified exact0
  • verified fuzzy26
  • unresolved30
  • parse uncertain0
  • malformed identifier3
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 08c990da-1e8e-4903-88b8-653ca79ca3cd · outbound

This paper cites GPT-4 Technical Report.

On the Consistency of Video Large Language Models in Temporal Comprehension GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.665661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.665661Z digest=sha256:84872c7b27bc5322a042506b1a086cd5fbdc33aed850f1e05870041778d755a0

Observation b955f145-fd79-42bf-b5b4-2da3da0f60c4 · outbound

This paper cites The surprising effectiveness of multimodal large language models for video moment retrieval.

On the Consistency of Video Large Language Models in Temporal Comprehension The surprising effectiveness of multimodal large language models for video moment retrieval

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.670834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.670834Z digest=sha256:b53a4ecc772e7df9893224b643ba851b7b2c3966018c8f6e2c7fe74bf9731f66

Observation b3c56589-7156-44e0-8e30-add47b684379 · outbound

This paper cites End- to-end object detection with transformers.

On the Consistency of Video Large Language Models in Temporal Comprehension End- to-end object detection with transformers

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:07:02.257703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:07:00.675443Z digest=sha256:66bdd8c7c21516ef373128ce92be538d518b34b60c6ad7a4626f5c5417fd4214

Observation 7f1de3a9-2cdc-4bca-ae3d-de878c1e28fa · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

On the Consistency of Video Large Language Models in Temporal Comprehension VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.680015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.680015Z digest=sha256:f1fb80dc42205e2ae04162d2eaad1662e22f9961ed0ed61905edd7cc10d768ec

Observation ad1d2f73-aa96-4fe9-a1c5-a424690178f3 · outbound

This paper cites Instructblip: Towards general-purpose vision- language models with instruction tuning, 2023.

On the Consistency of Video Large Language Models in Temporal Comprehension Instructblip: Towards general-purpose vision- language models with instruction tuning, 2023

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:07:02.239855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:07:00.685381Z digest=sha256:00135c45a858d94d781e9cead4434a7ecdb48414f06158efc8f5cce4a74cc1ba

Observation 6a659422-42fe-42db-b2f5-a86f256dd7a9 · outbound

This paper cites Measuring and improving consistency in pretrained language models.

On the Consistency of Video Large Language Models in Temporal Comprehension Measuring and improving consistency in pretrained language models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:07:02.222390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:07:00.690358Z digest=sha256:27850833c045a31c16b10b936cc4d3702dee4576c13714bdc2670b97bdfc3ead

Observation 628ecb47-2ece-4484-bbbe-89b43659ec8e · outbound

This paper cites Tall: Temporal activity localization via language query.

On the Consistency of Video Large Language Models in Temporal Comprehension Tall: Temporal activity localization via language query

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.695685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.695685Z digest=sha256:2c28b06f6042585b0da6d0be9badf90347fe2a4ccba70b0db23f09f0251d0d8f

Observation d1c10303-a580-4b2c-85bb-671a12b0ab79 · outbound

This paper cites VTG-LLM: Integrating Timestamp Knowledge into Video LLMs for Enhanced Video Temporal Grounding.

On the Consistency of Video Large Language Models in Temporal Comprehension VTG-LLM: Integrating Timestamp Knowledge into Video LLMs for Enhanced Video Temporal Grounding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.700588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.700588Z digest=sha256:12f207b0c4177ba3cd5c4c27f93874069fdf2b0a42909f8d3aa1186f90f39557

Observation c50191b1-4e91-419b-94f6-ceb545b407d0 · outbound

This paper cites Vtimellm: Empower llm to grasp video moments.

On the Consistency of Video Large Language Models in Temporal Comprehension Vtimellm: Empower llm to grasp video moments

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:07:02.192540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:07:00.705927Z digest=sha256:818501d3f723123df96e01326ac667eaf30ce87795b0f40c9308fc2110123228

Observation 83e26ae3-5819-42ba-a3f9-862648089c34 · outbound

This paper cites LITA: Language Instructed Temporal-Localization Assistant.

On the Consistency of Video Large Language Models in Temporal Comprehension LITA: Language Instructed Temporal-Localization Assistant

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.716282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.716282Z digest=sha256:4c7757c72f297daad462e23f367306023c8cde562dd1eadba9dc32d7fc1d87f2

Observation f0cca322-944d-419b-a950-650c1a0b051e · outbound

This paper cites Modal-specific pseudo query genera- tion for video corpus moment retrieval.

On the Consistency of Video Large Language Models in Temporal Comprehension Modal-specific pseudo query genera- tion for video corpus moment retrieval

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:07:02.173036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:07:00.722101Z digest=sha256:7c8bf1d1389e6c25de304f88e39df49792dcea621ad3b19b7f070a3d169fae66

Observation 68388e89-0a61-4557-ba7c-0b2692f18aef · outbound

This paper cites Background-aware moment detection for video moment retrieval.

On the Consistency of Video Large Language Models in Temporal Comprehension Background-aware moment detection for video moment retrieval

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:07:02.156093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:07:00.726946Z digest=sha256:bc46f3d4e451e7c88d730fb3887409b9b84892b797bf35302248f6c21b5c695b

Observation 129af2f7-9a6b-4b7e-bc82-8b43005c2458 · outbound

This paper cites Language Repository for Long Video Understanding.

On the Consistency of Video Large Language Models in Temporal Comprehension Language Repository for Long Video Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.731830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.731830Z digest=sha256:a7db0d8773cc15ef113feed2a06384aa98d3e0ad30135d061e99fa6d33904583

Observation a9a11936-c5ec-4787-9ec8-97e76b1767f8 · outbound

This paper cites Dense-captioning events in videos.

On the Consistency of Video Large Language Models in Temporal Comprehension Dense-captioning events in videos

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:07:02.137308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:07:00.737259Z digest=sha256:1bb3bb460c378dd650388572c5ce82a26b732963a4057de938ffa024b515a24b

Observation 6ff56cb3-4f18-41ec-ae05-55a96c0bb826 · outbound

This paper cites Detecting mo- ments and highlights in videos via natural language queries.

On the Consistency of Video Large Language Models in Temporal Comprehension Detecting mo- ments and highlights in videos via natural language queries

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:07:02.116551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:07:00.742401Z digest=sha256:de04ea795398b187af79a1f73f0482c1c952c9f6437bf8b564872bbef97582a1

Observation 3b98d44e-4679-4ccc-9437-5f7e00f4ee6b · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

On the Consistency of Video Large Language Models in Temporal Comprehension VideoChat: Chat-Centric Video Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.747917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.747917Z digest=sha256:91c08137976467a57b560a46c8373b6bb752eb4be5d6f4c82a7f29f250f8980b

Observation daf696bf-3bbd-4c77-a87b-b3a180fd56cd · outbound

This paper cites Mvbench: A comprehensive multi-modal video understand- ing benchmark.

On the Consistency of Video Large Language Models in Temporal Comprehension Mvbench: A comprehensive multi-modal video understand- ing benchmark

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:07:02.096278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:07:00.753337Z digest=sha256:4b9fde24a50946d629878edf12855e754c0877c9ad79d18297423593326363df

Observation 6eb5061f-cf7e-4ea8-b19a-63b639977b23 · outbound

This paper cites VITATECS: A Diagnostic Dataset for Temporal Concept Understanding of Video-Language Models.

On the Consistency of Video Large Language Models in Temporal Comprehension VITATECS: A Diagnostic Dataset for Temporal Concept Understanding of Video-Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.758146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.758146Z digest=sha256:f0d6af42afe5e767f20d7b84f70380d3ee4bb67bac71598434c3f7dbe1a0038f

Observation cb8b437f-aed0-497b-b36b-f42d083d8a54 · outbound

This paper cites Benchmarking and Improving Generator-Validator Consistency of Language Models.

On the Consistency of Video Large Language Models in Temporal Comprehension Benchmarking and Improving Generator-Validator Consistency of Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.763059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.763059Z digest=sha256:057756730016f94666664f2f43d3fa04324fdc4daf312ba11ddaa520388333c2

Observation a5ef0812-feb0-4a45-9f1d-dbc5e192265f · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

On the Consistency of Video Large Language Models in Temporal Comprehension Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.768085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.768085Z digest=sha256:8926bacc13bff0be648393e7051c13dd5aa286551d1bf35ddf89b13815eeef60

Observation f8fe26f4-9732-427d-ab48-ac4823d85ce5 · outbound

This paper cites Visual instruction tuning, 2023.

On the Consistency of Video Large Language Models in Temporal Comprehension Visual instruction tuning, 2023

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.773379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.773379Z digest=sha256:1266a6f5be8f601ade864ce467e4d4bb822ba7946916cadb6ee9aedfac45ee8e

Observation bcc114b1-bc9b-4935-b563-bf168f02d0ec · outbound

This paper cites TempCompass: Do Video LLMs Really Understand Videos?.

On the Consistency of Video Large Language Models in Temporal Comprehension TempCompass: Do Video LLMs Really Understand Videos?

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.778290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.778290Z digest=sha256:75c10cf8e90e23193007e89adf45535f1ea71da0167add4bd85efd90a652e4ad

Observation 6af2b771-a107-4a7a-a8d2-f466382770b7 · outbound

This paper cites Valley: Video Assistant with Large Language model Enhanced abilitY.

On the Consistency of Video Large Language Models in Temporal Comprehension Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.783631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.783631Z digest=sha256:6fbf4ec27f87b5dcf620761081b4f09f1e367f10581ab48ef385518b954a8a99

Observation fd341e08-3ec7-49fb-a242-d8f604083cb9 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

On the Consistency of Video Large Language Models in Temporal Comprehension Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:07:02.056181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:07:00.796312Z digest=sha256:2fb6d93acdb5ec389d5edf85efae9a6574b4ab1286f1ed3aa61403a19a4e6055

Observation cf2e423e-53b8-461c-a37a-a13460241155 · outbound

This paper cites Query-dependent video representation for moment retrieval and highlight detection.

On the Consistency of Video Large Language Models in Temporal Comprehension Query-dependent video representation for moment retrieval and highlight detection

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:07:02.034221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:07:00.800616Z digest=sha256:466ebc1b48fc2e913ed8626d4a5120f2e50144f6c402a941e13677f308e6c24e

Observation f10669a0-866d-4530-b045-cb19b8ef92db · outbound

This paper cites Local-global video-text interactions for temporal grounding.

On the Consistency of Video Large Language Models in Temporal Comprehension Local-global video-text interactions for temporal grounding

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:07:02.013322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:07:00.805391Z digest=sha256:18c251cb4375b6797bad27e2d05062365fbe965d2ccde8efde779b9dd88caa92

Observation 081f7dd8-e3a6-4ac6-a1ad-3459c8ad70bb · outbound

This paper cites Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models.

On the Consistency of Video Large Language Models in Temporal Comprehension Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.809662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.809662Z digest=sha256:26fa95c74c5aec24ff1006f91a6098b7db5f24380d5362430191b4e2ce2fceaa

Observation f32a4cb9-4c3a-4245-bf2f-05710e46e5d8 · outbound

This paper cites Uncovering Hidden Challenges in Query-Based Video Moment Retrieval.

On the Consistency of Video Large Language Models in Temporal Comprehension Uncovering Hidden Challenges in Query-Based Video Moment Retrieval

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.814437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.814437Z digest=sha256:6c3f21255ec4e89427b2e6e303c9d34755f0510726c4b61c90399d4f10fdab29

Observation 52ba0e91-0b05-483b-a6b9-dbba00470632 · outbound

This paper cites Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning.

On the Consistency of Video Large Language Models in Temporal Comprehension Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.819490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.819490Z digest=sha256:79545d0ce687f4014224f08ead8b0deb391670173a8692d18b547f86fea4016d

Observation d7fd9c77-a4ce-4769-99bc-27b13b521333 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

On the Consistency of Video Large Language Models in Temporal Comprehension Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.824434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.824434Z digest=sha256:f0221322f7ba8420a34ab5df2bd70c78acf4d10b52cd52e7a5a87b54d325e3c5

Observation 7525b355-3d6c-414c-b3e2-48a18b4b3b0f · outbound

This paper cites Timechat: A time-sensitive multimodal large lan- guage model for long video understanding.

On the Consistency of Video Large Language Models in Temporal Comprehension Timechat: A time-sensitive multimodal large lan- guage model for long video understanding

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:07:01.996307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:07:00.829905Z digest=sha256:10173cce74d8afe9a9ca5d03abfa09ea0940ba7def1c83eb77801ce5f3a76f02

Observation edd3e7d5-3664-4011-9117-43be9ee7776d · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

On the Consistency of Video Large Language Models in Temporal Comprehension Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.835798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.835798Z digest=sha256:00531fce1df88d989c3c94610f9dc2dac1713005f59ffd961e6f7be34c231540

Observation 8ad9f57c-8738-44b4-9e96-6089277393f9 · outbound

This paper cites HawkEye: Training Video-Text LLMs for Grounding Text in Videos.

On the Consistency of Video Large Language Models in Temporal Comprehension HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.840663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.840663Z digest=sha256:9b26914daf1380de960b06a45e9f4b369d086abb76df9744d02b1c3dd13e033c

Observation 3c426dd8-6ad6-4685-a14d-c9b05c6a5b37 · outbound

This paper cites Negative sample matters: A renaissance of metric learning for temporal grounding.

On the Consistency of Video Large Language Models in Temporal Comprehension Negative sample matters: A renaissance of metric learning for temporal grounding

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:07:01.978865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:07:00.845487Z digest=sha256:17f7bc8ebb2c578a35c3c3d82516e8b977a9e2bdf39354b67dae55f02b872d04

Observation c0706820-e09c-44c9-b3e5-afab73f553b6 · outbound

This paper cites Chain-of- thought prompting elicits reasoning in large language models.

On the Consistency of Video Large Language Models in Temporal Comprehension Chain-of- thought prompting elicits reasoning in large language models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:07:01.961610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:07:00.850430Z digest=sha256:3ea6ab242ca5c5eda55b2ecd6f2993712594947d908cc45f6025b7a3bcab31fe

Observation 97b95638-3e78-4d88-8d31-996da5028e42 · outbound

This paper cites VideoQA in the Era of LLMs: An Empirical Study.

On the Consistency of Video Large Language Models in Temporal Comprehension VideoQA in the Era of LLMs: An Empirical Study

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.854915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.854915Z digest=sha256:422121062e733dbc233b7e66d1125683c8360cbb8967c36bb90a86597dd07010

Observation 574da5c1-99cd-4299-8ab0-9072c58b7e7d · outbound

This paper cites Can i trust your answer? visually grounded video question answering.

On the Consistency of Video Large Language Models in Temporal Comprehension Can i trust your answer? visually grounded video question answering

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:07:01.943589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:07:00.860000Z digest=sha256:4b960b143198cab2f02fae87f5a97ed0e67fe059107193bac609ef9d64b38457

Observation f9ec1c82-fc35-44e0-b10d-bd132f8e91b3 · outbound

This paper cites A closer look at temporal sentence ground- ing in videos: Dataset and metric.

On the Consistency of Video Large Language Models in Temporal Comprehension A closer look at temporal sentence ground- ing in videos: Dataset and metric

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:07:01.924923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:07:00.864538Z digest=sha256:5a342964abf6e73732db94b75dfb7cf62725e6c9319075d2071480e6964fd878

Observation 7101815f-3ff4-4f8e-80eb-738370f61abe · outbound

This paper cites Sc-tune: Unleashing self-consistent referential compre- hension in large vision language models.

On the Consistency of Video Large Language Models in Temporal Comprehension Sc-tune: Unleashing self-consistent referential compre- hension in large vision language models

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:07:01.908456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:07:00.869110Z digest=sha256:9f88aa6b19f1e0475ac5049de5a5b0fb5ed512fdd30b0bc4263619dd7e70b8f2

Observation 5bdc71a0-763f-44ab-8f84-70530c3739b9 · outbound

This paper cites Span-based Localizing Network for Natural Language Video Localization.

On the Consistency of Video Large Language Models in Temporal Comprehension Span-based Localizing Network for Natural Language Video Localization

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.874181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.874181Z digest=sha256:24de6ec94612abe7b0776aa860af7ca0bcc2942f37247fb024f58176a694b8bd

Observation 27c0a6f0-4963-4354-98a4-d5c6d0850ad0 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

On the Consistency of Video Large Language Models in Temporal Comprehension Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.879008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.879008Z digest=sha256:5d56179a61f3bad22b4f6e899f9110ab2fd69ba725825f8c50f6dd371d88c42b

Observation 6964b318-5963-4117-aa1a-4d7bd5816eca · outbound

This paper cites Learning 2d temporal adjacent networks for moment local- ization with natural language.

On the Consistency of Video Large Language Models in Temporal Comprehension Learning 2d temporal adjacent networks for moment local- ization with natural language

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.884824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.884824Z digest=sha256:7467b044f00b722db5cd047ae786fbe6f6472c052eeb8698f47754e33e5d5b88

Observation 5f458d21-c5b7-4dbf-a0ca-533ea3693a1b · outbound

This paper cites Unveiling the Tapestry of Consistency in Large Vision-Language Models.

On the Consistency of Video Large Language Models in Temporal Comprehension Unveiling the Tapestry of Consistency in Large Vision-Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.889826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.889826Z digest=sha256:cc10ba337ede7d6e56c0037f329a12f2a42dc6dbc59dbc0e3179f706cd1ea60b

Observation ab874dcf-3866-4f0b-80a6-0a7d39a40a66 · outbound

This paper cites Prompt Consistency for Zero-Shot Task Generalization.

On the Consistency of Video Large Language Models in Temporal Comprehension Prompt Consistency for Zero-Shot Task Generalization

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.895247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.895247Z digest=sha256:d9f1096c99c93e52cc5adcd4086ecd957acbc51e02a71f9ca1b019a4e011057b

Observation 6af2e0c8-f895-43dd-854b-1d6da6277b2e · outbound

This paper cites Towards auto- matic learning of procedures from web instructional videos.

On the Consistency of Video Large Language Models in Temporal Comprehension Towards auto- matic learning of procedures from web instructional videos

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.901267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.901267Z digest=sha256:3c8d0631aa3a41eaf95fc856ac092d4b8fe69da0ebfe52a10ac61e99b2f5705a

Observation 07b15573-2b1f-4fb0-b96f-1eb1ac93e7a7 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

On the Consistency of Video Large Language Models in Temporal Comprehension MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.905873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.905873Z digest=sha256:41beb9679508cdca21bc4dfc849c01c94e7a7d770a7664a0b0db6b6c93af74e5

Observation fff885c0-7168-462e-bd0c-cd1a399c9171 · outbound

This paper cites It shows a remarkable zero-shot audio understanding capability and also generates responses to the visual and audio information presented in the videos.

On the Consistency of Video Large Language Models in Temporal Comprehension It shows a remarkable zero-shot audio understanding capability and also generates responses to the visual and audio information presented in the videos

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:07:01.865606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:07:00.910597Z digest=sha256:ffd88c6eec45f301a3da68de5c48ed31b651f36ac95fc20960d784a349fccb5d

Observation b40b6262-0a19-4173-b722-e609557a1753 · outbound

This paper cites To do this, Video- LLaV A collects both image and video-text datasets and incorporates them in its instruction tuning.

On the Consistency of Video Large Language Models in Temporal Comprehension To do this, Video- LLaV A collects both image and video-text datasets and incorporates them in its instruction tuning

Reference 48

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T17:07:01.843878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:07:00.915214Z digest=sha256:2f9f0725f5bffa18a518ed14619cc554eaa6e42b978c15cf34fce38e625ba741

Observation a35b7f22-f46b-48f1-ae24-83cb385aaf2d · outbound

This paper cites It introduces a new dataset for video instruction tuning, containing 100,000 high-quality video-instruction pairs.

On the Consistency of Video Large Language Models in Temporal Comprehension It introduces a new dataset for video instruction tuning, containing 100,000 high-quality video-instruction pairs

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:07:01.827566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:07:00.919502Z digest=sha256:ac026677a72ede2dddfab00c53e4a888f5d4ce1c942d086e156c367d908a9476

Observation 17f59f9c-44dc-49ae-827c-ec549a3eddba · outbound

This paper cites an unresolved cited work.

On the Consistency of Video Large Language Models in Temporal Comprehension Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:07:01.807893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:07:00.924171Z digest=sha256:589442fe23c2c02ea56ce26abd95cb57bf142dcacfea6083162a3e7721f47b59

Observation 3996aac8-ead6-49ad-879d-829a2fe2e155 · outbound

This paper cites The format should be: ’start time - end seconds’.

On the Consistency of Video Large Language Models in Temporal Comprehension The format should be: ’start time - end seconds’

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:07:01.789263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:07:00.928772Z digest=sha256:78c15149d44fe7a422d8c74cd879b78cbaf659b863f770b2e72455f85b5b164d

Observation a304e0bb-3c63-451d-a8fa-d2311d2200bc · outbound

This paper cites The output format should be: ’start - end seconds’.

On the Consistency of Video Large Language Models in Temporal Comprehension The output format should be: ’start - end seconds’

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:07:01.771741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:07:00.932889Z digest=sha256:73d7acec2b408e43a2991b1cc6f2585dcff19724cdb84a12cd06738451b928c4

Observation 5cc8dc9c-75cd-4577-bfff-c5c5fa42642f · outbound

This paper cites Specifically, they aim to align vision and text in the first stage and then generate captions from various image-text pairs.

On the Consistency of Video Large Language Models in Temporal Comprehension Specifically, they aim to align vision and text in the first stage and then generate captions from various image-text pairs

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:07:01.753056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:07:00.937581Z digest=sha256:9cf4ad8ddc564eac0be77b08941aec9e8b66a9e8a367bb4bdea53a4059f57120

Observation 57a7ffd1-1db5-4dbe-bc4c-4e3989e145c7 · outbound

This paper cites They seamlessly integrate both visual and audio modalities in videos and propose STC connector to understand spatiotemporal video informa- tion.

On the Consistency of Video Large Language Models in Temporal Comprehension They seamlessly integrate both visual and audio modalities in videos and propose STC connector to understand spatiotemporal video informa- tion

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:07:01.731635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:07:00.941979Z digest=sha256:cb7e590a8ee1beca0500fecc27529267a7fc656b5f03192a6ad9c2e9fbac3133

Observation 3bb54507-8220-42f6-8272-5a8da55627ff · outbound

This paper cites an unresolved cited work.

On the Consistency of Video Large Language Models in Temporal Comprehension Unresolved cited work

Reference 55

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T17:07:01.711940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:07:00.946401Z digest=sha256:2b76234fbf2dda55bf874cdb01a7ff1426ff255533da145be130cc6bf36db905

Observation 46f06485-429f-4d83-8e75-e4127754a740 · outbound

This paper cites an unresolved cited work.

On the Consistency of Video Large Language Models in Temporal Comprehension Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:07:01.694046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:07:00.951178Z digest=sha256:76efce822106f57144fa8f66f6c9dfbce7c3861ba5de27064ef2ec4ff2bbf90d

Observation 041bec98-5982-4e8c-8da6-4fb868e61657 · outbound

This paper cites I’m unable to find timestamps in the video.

On the Consistency of Video Large Language Models in Temporal Comprehension I’m unable to find timestamps in the video

Reference 57

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T17:07:01.674919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:07:00.955768Z digest=sha256:4472085a6cebf0481635011651816cbbbda7df38b7c13239e843269297ebb59d

Observation 83407866-e070-4e37-98b1-af452de9d1ae · outbound

This paper cites Experiments on Charades-CON with TimeChat.

On the Consistency of Video Large Language Models in Temporal Comprehension Experiments on Charades-CON with TimeChat

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:07:01.657803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:07:00.961927Z digest=sha256:ca960c930da54c71f4f2c1159fcd97434b9ff151971cd761d14e09828cdedd43

Observation 64614f41-df1e-4010-a7fd-abc914e6d17b · outbound

This paper cites 16 Figure 11.

On the Consistency of Video Large Language Models in Temporal Comprehension 16 Figure 11

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:07:01.639697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:07:00.966961Z digest=sha256:7846f129a76bc1aef437c2d879a320188a59f7a36dabda2a20c7d8f83ddf7e58

Pith citing papers

No inbound Pith citation observations are available.