Pith. sign in

Paper Citation Record · LEDGER

GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them?

As of 9 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 3 inbound Pith citation observations for arXiv:2507.09491.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.09491 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:58:28.781308Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T12:21:30.774303Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T12:53:05.467023Z

Reference resolution

21 of 21 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d34b630c-e942-4ef4-8a8a-0d5a9c2c5151 · outbound

This paper cites AutoEval-Video: An Automatic Benchmark for Assessing Large Vision Language Models in Open-Ended Video Question Answering.

GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? AutoEval-Video: An Automatic Benchmark for Assessing Large Vision Language Models in Open-Ended Video Question Answering

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T17:58:27.085682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:58:27.085682Z digest=sha256:b04be5d9d84136a40427b5b409fed24664e1f2bda21e7411193b13ad4aacaac1

Observation 2ceb502e-6b49-4ab2-a2ca-701b69ed2a18 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T17:58:27.181609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:58:27.181609Z digest=sha256:c95b8cea9ed9cd7d624ea6f9be86b8138e79e578f5a1fc18e04f8b7aa2bd6ccc

Observation b4de45de-7812-4775-9d3e-8ec6a1d423cb · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T17:58:27.320461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:58:27.320461Z digest=sha256:3bb20f7adc670516be7651eca165cdb92c50bf31ab646541725a833bcdde73ae

Observation 429e1c97-5d48-4974-bca6-377506aab120 · outbound

This paper cites Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding.

GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T17:58:27.390172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:58:27.390172Z digest=sha256:59fbf08a80872165d221ead10df802fb963daaace834eb9d688025e32127bb74

Observation 5a7676cf-84d0-49ea-9aba-f61972b1194c · outbound

This paper cites ViLMA: A Zero-Shot Benchmark for Linguistic and Temporal Grounding in Video-Language Models.

GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? ViLMA: A Zero-Shot Benchmark for Linguistic and Temporal Grounding in Video-Language Models

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T17:58:29.172776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:58:27.448450Z digest=sha256:147e881f883f31f67d65bfb7ce1261b92098f6b53170d2878fe84c7ca739af9a

Observation 2e9fe91c-3242-4a7e-84de-3a088856acb2 · outbound

This paper cites VLM-Eval: A General Evaluation on Video Large Language Models.

GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? VLM-Eval: A General Evaluation on Video Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T17:58:27.497281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:58:27.497281Z digest=sha256:e5b69f477bcce875f91fb7db7ce34bb6c6fd448b270e7c8472478af1338d6438

Observation 12d3c181-6b2b-48f5-9a19-2e74dab9d8b7 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T17:58:27.574010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:58:27.574010Z digest=sha256:8926b20b68bdd01c22aa8a802df5e692558115c86442b74852e753dc07551c50

Observation 29e4b54c-e510-4e57-a49e-06d18a014e81 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T17:58:27.664978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:58:27.664978Z digest=sha256:acb512d6e917473d3c6d656b4314c1cecd122c2405a3883b04c260443e9e184d

Observation e3afeb3f-c265-4265-b6d6-c827bc88199f · outbound

This paper cites Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models.

GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T17:58:27.742728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:58:27.742728Z digest=sha256:da678512537116f8be77f62fb28b4de044639e8e03e6813b4b1ff615bb5ad4b3

Observation 6bae6f8d-83e3-4ed8-bd95-2e4eb9a51a04 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T17:58:27.869001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:58:27.869001Z digest=sha256:0e5a75dcec2efcad9234ef234738ece2dabe9c4b2f3773a12c17597d59666c64

Observation 6a1d268b-9279-4f76-b196-2ea10c60c1ac · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? Gemini: A Family of Highly Capable Multimodal Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T17:58:28.048573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:58:28.048573Z digest=sha256:7912bd336ac9fb0ebc1f70cba3893f044f69f1bcd87fcdfd2e216d821bfcf5cf

Observation 0df38b9a-421b-4c41-bce0-512fefeb66ad · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T17:58:28.123501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:58:28.123501Z digest=sha256:44c4101729cbe048b7747b913177bea8b6849dbc437ba032f9486654cf2e08aa

Observation 9690be43-8fa6-4a16-8d82-8cfc3cee34b6 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? LLaMA: Open and Efficient Foundation Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T17:58:28.210161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:58:28.210161Z digest=sha256:e1f2f03a49de3d678ff3faca5cffe8fb5eb503ffbc9965fef0d78efe81dd96ef

Observation 5bc9de2a-5ae8-4fba-a9dd-d6a1ad4f60d7 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T17:58:28.290683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:58:28.290683Z digest=sha256:3734ed24826fc74b336f61dd45b907dce045a32e6b92c5e2cdaeebfd5a9a6786

Observation f38c72b6-ab00-48f1-9357-6ad9026aae26 · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T17:58:28.410060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:58:28.410060Z digest=sha256:0e0c0a9492e89cf9074a791916b04e5d03dc65bbc98e562c5bbe1d6c642ce603

Observation c90621e8-d91b-4dfa-bc18-bf43b47f0302 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:58:28.607448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:58:28.607448Z digest=sha256:1b54a3dc447d196545b70b8bfa3da055f7ff064f76ebc4734b138805ab795e45

Observation cee639af-b059-46b5-a330-3c58a5c54f4c · outbound

This paper cites Calibrated Self-Rewarding Vision Language Models.

GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? Calibrated Self-Rewarding Vision Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T17:58:28.693651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:58:28.693651Z digest=sha256:a00986e6d15fd9da70aaa8ca821cb280e05e8ddf8e31433323d29c1573ed121a

Observation 7185eb2e-ce12-4469-a11e-d8a45750ae99 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T17:58:28.781308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:58:28.781308Z digest=sha256:3dabe33f831e181421646fbd7f32958131e35af03bdaeb5b282d5cd893eb317c

Observation 51d4fdf8-37b1-40a6-99de-acddef1db62f · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T17:58:26.922661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:58:26.922661Z digest=sha256:635e7139122e171c6ee5af269fad73a8067640f60f90c0fbfb51e0317187a973

Observation bbe643d1-35b2-4d2e-9e1a-e0b64b67fc02 · outbound

This paper cites TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models.

GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T17:58:26.999928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:58:26.999928Z digest=sha256:4d51d9e8eaad3c3a5889e83a34194e15a4aed76ab2c4aa785d4188e9c1c62aeb

Observation 1ab8663f-6b5a-4105-a085-03e570d9dc04 · outbound

This paper cites Unhackable Temporal Rewarding for Scalable Video MLLMs.

GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? Unhackable Temporal Rewarding for Scalable Video MLLMs

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T17:58:28.549081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:58:28.549081Z digest=sha256:f6ca784c6bb6a39b7e930ebf6e36fe068b59b5281d2d9cb60abf6c29668d3eab

Pith citing papers

Observation 237a465b-3f7e-41f7-992a-de9b69301f24 · inbound

Low-Cost Test-Time Adaptation for Robust Video Editing cites this paper.

Low-Cost Test-Time Adaptation for Robust Video Editing GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them?

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T12:21:30.774303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:21:30.774303Z digest=sha256:a1a24348de42b4e5ea762d20545317c7caccf1026f459e80ea155494745f193c

Observation 68f4f59a-9d09-44bc-bba5-561a89b54990 · inbound

Multi-Modal Machine Learning Framework for Predicting Early Recurrence of Brain Tumors Using MRI and Clinical Biomarkers cites this paper.

Multi-Modal Machine Learning Framework for Predicting Early Recurrence of Brain Tumors Using MRI and Clinical Biomarkers GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them?

Reference 85

Resolution
verified exact
local_arxiv, observed 2026-08-05T12:53:05.471912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T12:53:04.416698Z digest=sha256:6f2a70b08c85372c68119899057e9ffa69ea6305f9283b79d42ddfc6e8729c01

Observation f14a26fc-4da5-4428-b4a9-2c93fcae3bf0 · inbound

A Multimodal Deep Learning Framework for Early Diagnosis of Liver Cancer via Optimized BiLSTM-AM-VMD Architecture cites this paper.

A Multimodal Deep Learning Framework for Early Diagnosis of Liver Cancer via Optimized BiLSTM-AM-VMD Architecture GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them?

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-05T12:53:15.118127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:53:15.118127Z digest=sha256:0808f88929da213180724af734fa6500048212eeb8293e39c977cd7daad40d4c