Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T17:58:28.781308Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 3 inbound Pith citation observations for arXiv:2507.09491.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T17:58:28.781308Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T12:21:30.774303Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T12:53:05.467023Z
21 of 21 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d34b630c-e942-4ef4-8a8a-0d5a9c2c5151 · outbound
GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? AutoEval-Video: An Automatic Benchmark for Assessing Large Vision Language Models in Open-Ended Video Question Answering
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ceb502e-6b49-4ab2-a2ca-701b69ed2a18 · outbound
GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4de45de-7812-4775-9d3e-8ec6a1d423cb · outbound
GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 429e1c97-5d48-4974-bca6-377506aab120 · outbound
GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a7676cf-84d0-49ea-9aba-f61972b1194c · outbound
GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? ViLMA: A Zero-Shot Benchmark for Linguistic and Temporal Grounding in Video-Language Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2e9fe91c-3242-4a7e-84de-3a088856acb2 · outbound
GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? VLM-Eval: A General Evaluation on Video Large Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12d3c181-6b2b-48f5-9a19-2e74dab9d8b7 · outbound
GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29e4b54c-e510-4e57-a49e-06d18a014e81 · outbound
GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3afeb3f-c265-4265-b6d6-c827bc88199f · outbound
GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bae6f8d-83e3-4ed8-bd95-2e4eb9a51a04 · outbound
GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a1d268b-9279-4f76-b196-2ea10c60c1ac · outbound
GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? Gemini: A Family of Highly Capable Multimodal Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0df38b9a-421b-4c41-bce0-512fefeb66ad · outbound
GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9690be43-8fa6-4a16-8d82-8cfc3cee34b6 · outbound
GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? LLaMA: Open and Efficient Foundation Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bc9de2a-5ae8-4fba-a9dd-d6a1ad4f60d7 · outbound
GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f38c72b6-ab00-48f1-9357-6ad9026aae26 · outbound
GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c90621e8-d91b-4dfa-bc18-bf43b47f0302 · outbound
GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cee639af-b059-46b5-a330-3c58a5c54f4c · outbound
GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? Calibrated Self-Rewarding Vision Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7185eb2e-ce12-4469-a11e-d8a45750ae99 · outbound
GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51d4fdf8-37b1-40a6-99de-acddef1db62f · outbound
GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbe643d1-35b2-4d2e-9e1a-e0b64b67fc02 · outbound
GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ab8663f-6b5a-4105-a085-03e570d9dc04 · outbound
GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? Unhackable Temporal Rewarding for Scalable Video MLLMs
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 237a465b-3f7e-41f7-992a-de9b69301f24 · inbound
Low-Cost Test-Time Adaptation for Robust Video Editing GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them?
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68f4f59a-9d09-44bc-bba5-561a89b54990 · inbound
Multi-Modal Machine Learning Framework for Predicting Early Recurrence of Brain Tumors Using MRI and Clinical Biomarkers GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them?
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f14a26fc-4da5-4428-b4a9-2c93fcae3bf0 · inbound
A Multimodal Deep Learning Framework for Early Diagnosis of Liver Cancer via Optimized BiLSTM-AM-VMD Architecture GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them?
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.