Pith. sign in

Paper Citation Record · LEDGER

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering

As of 8 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 1 inbound Pith citation observation for arXiv:2505.24371.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24371 v3

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:30:34.249618Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:30:31.083510Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T12:30:34.809388Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact2
  • verified fuzzy20
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c65a519e-1029-4de4-bda0-8b02e1b64b5a · outbound

This paper cites Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:30:34.865426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:30:31.083510Z digest=sha256:9f5754cadea630f22c7886696af61384641c90da6745215564f140a6e2a3a11a

Observation 352a4101-0848-419e-83ce-c1e6b08b7fb8 · outbound

This paper cites Further research [5] extends the image VLM to support multi-frame and video data modalities.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering Further research [5] extends the image VLM to support multi-frame and video data modalities

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:38.424721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:30:31.146651Z digest=sha256:ce5bd489e045bf7bacbfb1cc354cabd681850e0e0e55a96848c388da93273cdf

Observation d97db90e-1f28-4999-8b0c-b3b39bbfd3d3 · outbound

This paper cites The first phase is the tran- scription generation phase.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering The first phase is the tran- scription generation phase

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:38.266635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:30:31.239277Z digest=sha256:2f0259746ad46e14605764714c3fca40cea22c4d838501286b90ee913c1ddc6c

Observation 74f755b7-84d5-41e9-a367-80dcf7a6f2de · outbound

This paper cites Baseline models We leveraged the LLaV A-1.6-7B VLM [23] to generate local and global transcriptions and the Llama-3.1-8B LLM [24] for the VideoQA task.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering Baseline models We leveraged the LLaV A-1.6-7B VLM [23] to generate local and global transcriptions and the Llama-3.1-8B LLM [24] for the VideoQA task

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:38.179602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:30:31.336355Z digest=sha256:d0db92b97587b57eb60cd027d8c257b2b510d65fca3daeccb130bbcd76778a8c

Observation 797c5fb9-80c5-4173-a6bf-a4b17cc92020 · outbound

This paper cites The system is inherently modular and built on open-source VLM and LLM.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering The system is inherently modular and built on open-source VLM and LLM

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:38.058928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:30:31.423370Z digest=sha256:33460eaad82c3ab577caa038542d1b4210013089d08916c3b4b825cb5ad4a7aa

Observation 53e15e52-b447-4329-b1ab-f728b312f753 · outbound

This paper cites MOKA: Open- vocabulary robotic manipulation through mark-based visual prompting,.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering MOKA: Open- vocabulary robotic manipulation through mark-based visual prompting,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:37.875398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:30:31.499625Z digest=sha256:8bedd1e8597a8fdba88881f8608a4b2845c5167dcc2b77b3640776176811aa47

Observation 382d85ff-e26b-49d7-9343-02f316d43980 · outbound

This paper cites VLAAD: Vision and language assistant for au- tonomous driving,.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering VLAAD: Vision and language assistant for au- tonomous driving,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:37.669183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:30:31.568370Z digest=sha256:3f220b684ef75170cd86acdcc9206206b03dd5c2090a0da2c3c403534d16f4bf

Observation 59c1cdc9-db3a-4ac2-9831-49adbccd3cbe · outbound

This paper cites Smart customer service in unmanned retail store enhanced by large language model,.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering Smart customer service in unmanned retail store enhanced by large language model,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:37.490277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:30:31.641448Z digest=sha256:61fbba9ab66e8f9aa47118d6daf15e94c3482694f9dbecae6a75b0fdae867d2f

Observation c0d8171a-5c54-4539-9534-0024ca4953cf · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:31.714734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:31.714734Z digest=sha256:770203b634982f67ba3094ac2a3461b9fd38a4f7cf84aa90c42a7da013347b09

Observation 0b8ea0af-1122-47fb-9588-f63d5c0fe1f3 · outbound

This paper cites LLaV A-OneVision: Easy visual task transfer,.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering LLaV A-OneVision: Easy visual task transfer,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:37.325359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:30:31.782618Z digest=sha256:7e4880a232db264860ed6d1447fb5b73319e57fa2d9fd1eae39289f87eb9f41c

Observation ac08881b-48a1-4659-a175-f4a53b64c0d8 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering VideoChat: Chat-Centric Video Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:31.859767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:31.859767Z digest=sha256:305945a978cdda86ccbc88e37ab8c35c793123918b91dac0a28eb6f4668bcf9e

Observation 31d3d6c0-03e8-4cb2-803f-7631c67d5c9d · outbound

This paper cites A simple LLM framework for long- range video question-answering,.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering A simple LLM framework for long- range video question-answering,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:37.158371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:30:31.933492Z digest=sha256:a0497a0b7e702ebce2a3b8a38a973c7bfe28d516efc281595e620fa4ec654153

Observation 0c71c004-3791-4536-a46f-72046256c69e · outbound

This paper cites GPT-4 Technical Report.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering GPT-4 Technical Report

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:32.012193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:32.012193Z digest=sha256:b3e0da21330d3bf4046a3ea4c2825b4d84df7734fa2fe19426ca7e8ce941a643

Observation cbe3c9c6-fc2b-408d-bb41-042cbc11112f · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering Gemini: A Family of Highly Capable Multimodal Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:32.096412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:32.096412Z digest=sha256:20a174cd358acf773414e078fdd0d603e92e07751e2198c81e9e39ba449318f7

Observation 5d6d3dc4-2dc4-41f1-ac3d-f8b7950efd09 · outbound

This paper cites Visual instruction tuning,.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering Visual instruction tuning,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:32.202556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:32.202556Z digest=sha256:632476ab0268c277056842397dd1f316d1a07a2fa24f711310e711474702fc4c

Observation c49c3092-be31-4fe1-936b-8c2c65624229 · outbound

This paper cites An image grid can be worth a video: Zero-shot video question answering using a vlm,.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering An image grid can be worth a video: Zero-shot video question answering using a vlm,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:36.936620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:30:32.275288Z digest=sha256:a3eddfcd2c5763f1c9f25a88fd51483eccaf8b12d480187eb4c80a289f4cfc0c

Observation c0414aa9-916d-4f00-9240-84c5999a2165 · outbound

This paper cites Self-chained image- language model for video localization and question answer- ing,.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering Self-chained image- language model for video localization and question answer- ing,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:36.788952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:30:32.380318Z digest=sha256:6070720c61f8deb84249c6e8f45f0664b9c7073fa1534e61e6a9326268626bb8

Observation 3768e06e-5716-4740-8abc-34f9f9b9459e · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:32.456684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:32.456684Z digest=sha256:103c1891bba2730391aa17f391c53d1f83c6e50c965ba405665270fc99c976e0

Observation b2e5dfae-a11a-49e4-bcbc-1eedfffda2b2 · outbound

This paper cites GeoChat: Grounded large vision-language model for remote sensing,.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering GeoChat: Grounded large vision-language model for remote sensing,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:36.594439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:30:32.558947Z digest=sha256:7aeb36cc5febe543f1372c7562408edd3c5d31b0daec119afc68fadfbddf60fa

Observation 21b7a483-93aa-4aa1-8821-9ec4b97b7df8 · outbound

This paper cites EarthMarker: A Visual Prompting Multi-modal Large Language Model for Remote Sensing.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering EarthMarker: A Visual Prompting Multi-modal Large Language Model for Remote Sensing

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:32.685636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:32.685636Z digest=sha256:e978c3162de4b3c3e0e504c8b94d30943a4d01186d56709fd9562e5b46cc2523

Observation 609b1966-0f06-4ac3-8f54-bb72cfdb3bb7 · outbound

This paper cites PIVOT: iterative visual prompting elicits actionable knowl- edge for vlms,.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering PIVOT: iterative visual prompting elicits actionable knowl- edge for vlms,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:36.385702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:30:32.807866Z digest=sha256:23708285117aaa067949eca52f67d0a8342e7bf998eb5585acf8b9a036ade4fd

Observation 4190cbcb-4f1e-4621-a6e7-d0333ec88c03 · outbound

This paper cites Open-Vocabulary Action Localization with Iterative Visual Prompting.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering Open-Vocabulary Action Localization with Iterative Visual Prompting

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:30:34.604723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:30:32.904432Z digest=sha256:8872a8f145909377c4dc155bc0a9889455d0602967df7e939c0af7df939f6d2d

Observation 2579bd6f-136b-45e3-97c8-57fd2d7e1a2c · outbound

This paper cites Video graph trans- former for video question answering,.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering Video graph trans- former for video question answering,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:36.199823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:30:33.038946Z digest=sha256:453ea47213b847b0e04653710b7db5e7d37623f4073099c19c8ca1ce634b7f43

Observation a19f1bcc-2c7e-4e11-9527-3b1c0426aa55 · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:33.191576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:33.191576Z digest=sha256:fbdb79dd1310d23e07ad439291be8667c906cc66563796510c54cb8f4ee21355

Observation 5a61dbaa-5afa-4bdb-af4a-d952a8e3dd2f · outbound

This paper cites MVBench: A comprehensive multi- modal video understanding benchmark,.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering MVBench: A comprehensive multi- modal video understanding benchmark,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:35.967511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:30:33.328725Z digest=sha256:2697f48c0a37eea005ef4f76591857a8565f5def953cb7b4b6996a719dae036a

Observation e5458f8c-7aa0-41de-b65b-35e2535e1d72 · outbound

This paper cites VISTA- LLAMA: Reducing hallucination in video language models via equal distance to visual tokens,.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering VISTA- LLAMA: Reducing hallucination in video language models via equal distance to visual tokens,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:35.767758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:30:33.421014Z digest=sha256:2840a07fa0f2a6fde1dfa75a6c43b60bb6079e3f12a71387a3e3587427186c37

Observation 3c87cb0a-7f1d-46f9-a926-58e4b0f43823 · outbound

This paper cites CogAgent: A visual lan- guage model for gui agents,.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering CogAgent: A visual lan- guage model for gui agents,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:35.515474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:30:33.562294Z digest=sha256:444a7b0d5ded1710360619f1d150a4efb043a1f85db14c4005915a362c80bd19

Observation a087d91c-8ca7-4179-9248-8d249309ce08 · outbound

This paper cites LLaV A-NeXT: Improved reasoning, OCR, and world knowledge,.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering LLaV A-NeXT: Improved reasoning, OCR, and world knowledge,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:35.344142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:30:33.653791Z digest=sha256:6689644b27ebf942ae838723bd7d3e9a695716d400619c5567246b4aebfda486

Observation 3bda2006-30c5-4bb7-a8ee-daf3b6c3f689 · outbound

This paper cites The Llama 3 Herd of Models.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering The Llama 3 Herd of Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:33.760176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:33.760176Z digest=sha256:b3d0bc2a118b620046b0f485e001bf3fb2a80c74abe15b8c2515aefa0e628607

Observation 16576960-2833-4818-a5b1-2a2608474e3c · outbound

This paper cites NExT-QA: Next phase of question-answering to explaining temporal actions,.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering NExT-QA: Next phase of question-answering to explaining temporal actions,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:35.229409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:30:34.055988Z digest=sha256:7fe78fbfb56ddd8284e5fffa8d1d5d4ed76c861328a351c09a35cacce65caf41

Observation 22ba6029-c29b-4815-abb5-4aeecc966736 · outbound

This paper cites STAR: A benchmark for situated reasoning in real-world videos,.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering STAR: A benchmark for situated reasoning in real-world videos,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:35.085113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:30:34.249618Z digest=sha256:39184fe4efd2d76fa62e191df19a45dbfb1cf182acc99a2d8375ebaf5151293f

Pith citing papers

Observation c65a519e-1029-4de4-bda0-8b02e1b64b5a · inbound

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering cites this paper.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:30:34.865426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:30:31.083510Z digest=sha256:9f5754cadea630f22c7886696af61384641c90da6745215564f140a6e2a3a11a