Pith. sign in

Paper Citation Record · LEDGER

Benchmarking Vision-Language Models on Optical Character Recognition in Dynamic Video Environments

As of 8 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 2 inbound Pith citation observations for arXiv:2502.06445.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06445 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T15:26:35.545652Z

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T10:52:13.947049Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T12:51:02.625990Z

Reference resolution

13 of 13 outbound references displayed

  • verified exact1
  • verified fuzzy10
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0a3b1be0-97aa-43ec-80f3-e5d331418088 · outbound

This paper cites Claude 3.5 Sonnet: Advancements in multimodal AI, 2024.

Benchmarking Vision-Language Models on Optical Character Recognition in Dynamic Video Environments Claude 3.5 Sonnet: Advancements in multimodal AI, 2024

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:26:35.704519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:26:35.497443Z digest=sha256:430a3986b654cee191397bdcabc14d2e22ad7107f25aa33655cc3a32639e4c68

Observation 1046109f-197b-4a05-8f89-d21b56d4f058 · outbound

This paper cites Character region awareness for text detection.

Benchmarking Vision-Language Models on Optical Character Recognition in Dynamic Video Environments Character region awareness for text detection

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:26:35.693817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:26:35.501446Z digest=sha256:e51096787316a82b5ed945d6c9b4a12514d11f42e4a21ac7afbb579a405284ec

Observation c2ee93df-3d7b-4fad-812f-465de9584b6b · outbound

This paper cites Video Summarisation with Incident and Context Information using Generative AI.

Benchmarking Vision-Language Models on Optical Character Recognition in Dynamic Video Environments Video Summarisation with Incident and Context Information using Generative AI

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-08T15:26:35.600757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:26:35.504918Z digest=sha256:56a3f717d04caff7d70a3977e05829c9139f31ca895f4f3dbd2e1fa42b6dffff

Observation eb881b5c-62a1-4c78-ac48-9438e8c4ecc6 · outbound

This paper cites Gemini 1.5 Pro: Pushing the boundaries of multimodal learning, 2024.

Benchmarking Vision-Language Models on Optical Character Recognition in Dynamic Video Environments Gemini 1.5 Pro: Pushing the boundaries of multimodal learning, 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:26:35.683494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:26:35.509583Z digest=sha256:229196bb043a39a573110f8f2774b7b128cd29e5866a1b3121b1068dfd772ba9

Observation 3f030ddb-1ba6-4207-a14d-5a2623b60c58 · outbound

This paper cites Connectionist temporal classification.Supervised sequence labelling with recurrent neural networks, pages 61–93, 2012.

Benchmarking Vision-Language Models on Optical Character Recognition in Dynamic Video Environments Connectionist temporal classification.Supervised sequence labelling with recurrent neural networks, pages 61–93, 2012

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:26:35.673053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:26:35.514088Z digest=sha256:dbc612dcbb7075926535aee0e77c0d38cf19189603be06b64f4336ba300f718f

Observation a6ab84df-6e34-4b81-976a-17b880e340ba · outbound

This paper cites EasyOCR: Ready-to-use OCR with 80+ supported languages, 2024.

Benchmarking Vision-Language Models on Optical Character Recognition in Dynamic Video Environments EasyOCR: Ready-to-use OCR with 80+ supported languages, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:26:35.662730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:26:35.518254Z digest=sha256:44e3e584d7cfa24147e44d6272f30307b7aecbc408d0bf90a065d511d058d9dc

Observation b51d91ee-3fde-434a-b2b1-f991fcf894e0 · outbound

This paper cites VHELM: A Holistic Evaluation of Vision Language Models.

Benchmarking Vision-Language Models on Optical Character Recognition in Dynamic Video Environments VHELM: A Holistic Evaluation of Vision Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T15:26:35.522585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:26:35.522585Z digest=sha256:2201a4c578bdadabe3f07638978acbd64ba53dfb9e10d0c00fe8ad2a026e6440

Observation d3e7a6cb-a8d5-4fa7-9907-7ae932fecda9 · outbound

This paper cites TVQA: Localized, Compositional Video Question Answering.

Benchmarking Vision-Language Models on Optical Character Recognition in Dynamic Video Environments TVQA: Localized, Compositional Video Question Answering

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T15:26:35.526768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:26:35.526768Z digest=sha256:d71fe67d572df27a6ed406dea07338fdf294e2d1dfbabbc676c2f1551dc929d3

Observation 516c8911-650f-42e8-8ff9-98f89b5a6e0a · outbound

This paper cites GPT-4o: Omni-modal language model, 2024.

Benchmarking Vision-Language Models on Optical Character Recognition in Dynamic Video Environments GPT-4o: Omni-modal language model, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:26:35.651666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:26:35.530884Z digest=sha256:ddcb36d0b48ceefcf9fdbe6e5ea3c5776e8d766e7effd48d98f699a77808f460

Observation 43496147-274f-406b-9bc9-9ecfa6285dcb · outbound

This paper cites PaddleOCR: An OCR toolset based on PaddlePaddle, 2023.

Benchmarking Vision-Language Models on Optical Character Recognition in Dynamic Video Environments PaddleOCR: An OCR toolset based on PaddlePaddle, 2023

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:26:35.641320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:26:35.534603Z digest=sha256:9bba46b09ddae15c1a8fad98f8d9231c986f54645ef7ef60a58a84ff503ec22b

Observation 214282ec-339f-4b66-a61f-8a46125e10e7 · outbound

This paper cites RapidOCR: A lightweight OCR framework, 2021.

Benchmarking Vision-Language Models on Optical Character Recognition in Dynamic Video Environments RapidOCR: A lightweight OCR framework, 2021

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:26:35.631189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:26:35.538170Z digest=sha256:a0973ea82f7efba56aa57399eff21006b0d4be5f107d3d604381bd44f5ef6d94

Observation bf814452-f886-4d36-b80e-f50f86e588f2 · outbound

This paper cites VideoDB: Video infrastructure for the AI first world, 2024.

Benchmarking Vision-Language Models on Optical Character Recognition in Dynamic Video Environments VideoDB: Video infrastructure for the AI first world, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:26:35.621872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:26:35.542032Z digest=sha256:624708b58aabe5d6a3a4dd4628c4d26f1e1f5bdb50fe47cce0a69f86f1ff52e2

Observation baacb20c-5206-4603-bbe7-40ebcaaf3f32 · outbound

This paper cites CONTEXT"? OF.

Benchmarking Vision-Language Models on Optical Character Recognition in Dynamic Video Environments CONTEXT"? OF

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:26:35.612278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:26:35.545652Z digest=sha256:18a762a7be1cb51fb6819f24d26c3e9809669984abda2f213514aa4aa8cb4e10

Pith citing papers

Observation 0e34a902-f194-46ac-b30e-8a3c051e7050 · inbound

E-ARMOR: Edge case Assessment and Review of Multilingual Optical Character Recognition cites this paper.

E-ARMOR: Edge case Assessment and Review of Multilingual Optical Character Recognition Benchmarking Vision-Language Models on Optical Character Recognition in Dynamic Video Environments

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T10:52:13.947049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:52:13.947049Z digest=sha256:fd2566098a413e9503d739e74f4ca695f2ac2d2fffb9a4bcd73e5eef78198147

Observation 4723749e-4596-4773-9479-2819028e83d0 · inbound

If you're waiting for a sign... that might not be it! Mitigating Trust Boundary Confusion from Visual Injections on Vision-Language Agentic Systems cites this paper.

If you're waiting for a sign... that might not be it! Mitigating Trust Boundary Confusion from Visual Injections on Vision-Language Agentic Systems Benchmarking Vision-Language Models on Optical Character Recognition in Dynamic Video Environments

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:51:02.632753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T02:54:35.336251Z digest=sha256:c742c72b0354a43490390d54284d336f2bbe6fd8b1695da0b7ee444df0906c2c