Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:30:34.249618Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 1 inbound Pith citation observation for arXiv:2505.24371.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:30:34.249618Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:30:31.083510Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T12:30:34.809388Z
31 of 31 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c65a519e-1029-4de4-bda0-8b02e1b64b5a · outbound
Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 352a4101-0848-419e-83ce-c1e6b08b7fb8 · outbound
Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering Further research [5] extends the image VLM to support multi-frame and video data modalities
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d97db90e-1f28-4999-8b0c-b3b39bbfd3d3 · outbound
Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering The first phase is the tran- scription generation phase
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 74f755b7-84d5-41e9-a367-80dcf7a6f2de · outbound
Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering Baseline models We leveraged the LLaV A-1.6-7B VLM [23] to generate local and global transcriptions and the Llama-3.1-8B LLM [24] for the VideoQA task
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 797c5fb9-80c5-4173-a6bf-a4b17cc92020 · outbound
Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering The system is inherently modular and built on open-source VLM and LLM
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 53e15e52-b447-4329-b1ab-f728b312f753 · outbound
Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering MOKA: Open- vocabulary robotic manipulation through mark-based visual prompting,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 382d85ff-e26b-49d7-9343-02f316d43980 · outbound
Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering VLAAD: Vision and language assistant for au- tonomous driving,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 59c1cdc9-db3a-4ac2-9831-49adbccd3cbe · outbound
Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering Smart customer service in unmanned retail store enhanced by large language model,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c0d8171a-5c54-4539-9534-0024ca4953cf · outbound
Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b8ea0af-1122-47fb-9588-f63d5c0fe1f3 · outbound
Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering LLaV A-OneVision: Easy visual task transfer,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ac08881b-48a1-4659-a175-f4a53b64c0d8 · outbound
Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering VideoChat: Chat-Centric Video Understanding
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31d3d6c0-03e8-4cb2-803f-7631c67d5c9d · outbound
Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering A simple LLM framework for long- range video question-answering,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0c71c004-3791-4536-a46f-72046256c69e · outbound
Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering GPT-4 Technical Report
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbe3c9c6-fc2b-408d-bb41-042cbc11112f · outbound
Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering Gemini: A Family of Highly Capable Multimodal Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d6d3dc4-2dc4-41f1-ac3d-f8b7950efd09 · outbound
Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering Visual instruction tuning,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c49c3092-be31-4fe1-936b-8c2c65624229 · outbound
Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering An image grid can be worth a video: Zero-shot video question answering using a vlm,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c0414aa9-916d-4f00-9240-84c5999a2165 · outbound
Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering Self-chained image- language model for video localization and question answer- ing,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3768e06e-5716-4740-8abc-34f9f9b9459e · outbound
Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2e5dfae-a11a-49e4-bcbc-1eedfffda2b2 · outbound
Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering GeoChat: Grounded large vision-language model for remote sensing,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 21b7a483-93aa-4aa1-8821-9ec4b97b7df8 · outbound
Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering EarthMarker: A Visual Prompting Multi-modal Large Language Model for Remote Sensing
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 609b1966-0f06-4ac3-8f54-bb72cfdb3bb7 · outbound
Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering PIVOT: iterative visual prompting elicits actionable knowl- edge for vlms,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4190cbcb-4f1e-4621-a6e7-d0333ec88c03 · outbound
Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering Open-Vocabulary Action Localization with Iterative Visual Prompting
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2579bd6f-136b-45e3-97c8-57fd2d7e1a2c · outbound
Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering Video graph trans- former for video question answering,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a19f1bcc-2c7e-4e11-9527-3b1c0426aa55 · outbound
Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a61dbaa-5afa-4bdb-af4a-d952a8e3dd2f · outbound
Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering MVBench: A comprehensive multi- modal video understanding benchmark,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e5458f8c-7aa0-41de-b65b-35e2535e1d72 · outbound
Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering VISTA- LLAMA: Reducing hallucination in video language models via equal distance to visual tokens,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3c87cb0a-7f1d-46f9-a926-58e4b0f43823 · outbound
Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering CogAgent: A visual lan- guage model for gui agents,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a087d91c-8ca7-4179-9248-8d249309ce08 · outbound
Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering LLaV A-NeXT: Improved reasoning, OCR, and world knowledge,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3bda2006-30c5-4bb7-a8ee-daf3b6c3f689 · outbound
Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering The Llama 3 Herd of Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16576960-2833-4818-a5b1-2a2608474e3c · outbound
Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering NExT-QA: Next phase of question-answering to explaining temporal actions,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 22ba6029-c29b-4815-abb5-4aeecc966736 · outbound
Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering STAR: A benchmark for situated reasoning in real-world videos,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c65a519e-1029-4de4-bda0-8b02e1b64b5a · inbound
Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.