Pith. sign in

Paper Citation Record · LEDGER

Benchmarking Vision-Language Models on Optical Character Recognition in Dynamic Video Environments

As of 9 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 2 inbound Pith citation observations for arXiv:2502.06445.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06445 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T15:26:35.545652Z

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T10:52:13.947049Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T12:51:02.625990Z

Reference resolution

13 of 13 outbound references displayed

  • verified exact1
  • verified fuzzy10
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0a3b1be0-97aa-43ec-80f3-e5d331418088 · outbound

This paper cites Claude 3.5 Sonnet: Advancements in multimodal AI, 2024.

Benchmarking Vision-Language Models on Optical Character Recognition in Dynamic Video Environments Claude 3.5 Sonnet: Advancements in multimodal AI, 2024

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:26:35.704519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:26:35.497443Z digest=sha256:f4fe19ba0bfc92e43c421a84d8db543f87e921cf4e590cec89cbcbc158e7ec4a

Observation 1046109f-197b-4a05-8f89-d21b56d4f058 · outbound

This paper cites Character region awareness for text detection.

Benchmarking Vision-Language Models on Optical Character Recognition in Dynamic Video Environments Character region awareness for text detection

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:26:35.693817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:26:35.501446Z digest=sha256:d175f4aa4f112e91810867afc95d7f3e34a9445c44ab0f73177febb6327ae1cd

Observation c2ee93df-3d7b-4fad-812f-465de9584b6b · outbound

This paper cites Video Summarisation with Incident and Context Information using Generative AI.

Benchmarking Vision-Language Models on Optical Character Recognition in Dynamic Video Environments Video Summarisation with Incident and Context Information using Generative AI

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-08T15:26:35.600757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:26:35.504918Z digest=sha256:fa64a58944dd5a35b6bb07a04118cd1f6d43b197199de4d899a1d0d31e9f7c4a

Observation eb881b5c-62a1-4c78-ac48-9438e8c4ecc6 · outbound

This paper cites Gemini 1.5 Pro: Pushing the boundaries of multimodal learning, 2024.

Benchmarking Vision-Language Models on Optical Character Recognition in Dynamic Video Environments Gemini 1.5 Pro: Pushing the boundaries of multimodal learning, 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:26:35.683494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:26:35.509583Z digest=sha256:3e6f98c64a44ebc868a66d29212a69cecc98858d18f9d914ab1050972d6fdce9

Observation 3f030ddb-1ba6-4207-a14d-5a2623b60c58 · outbound

This paper cites Connectionist temporal classification.Supervised sequence labelling with recurrent neural networks, pages 61–93, 2012.

Benchmarking Vision-Language Models on Optical Character Recognition in Dynamic Video Environments Connectionist temporal classification.Supervised sequence labelling with recurrent neural networks, pages 61–93, 2012

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:26:35.673053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:26:35.514088Z digest=sha256:92ac466fc91fe0a4cdf58bdc98a6adc4fe0193bb68d04e71b18b633ab9899028

Observation a6ab84df-6e34-4b81-976a-17b880e340ba · outbound

This paper cites EasyOCR: Ready-to-use OCR with 80+ supported languages, 2024.

Benchmarking Vision-Language Models on Optical Character Recognition in Dynamic Video Environments EasyOCR: Ready-to-use OCR with 80+ supported languages, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:26:35.662730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:26:35.518254Z digest=sha256:e34b7ca49f39cc707517dc355457e7dd59ed8f9e83641ffa639fc25f6a4c6aac

Observation b51d91ee-3fde-434a-b2b1-f991fcf894e0 · outbound

This paper cites VHELM: A Holistic Evaluation of Vision Language Models.

Benchmarking Vision-Language Models on Optical Character Recognition in Dynamic Video Environments VHELM: A Holistic Evaluation of Vision Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T15:26:35.522585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:26:35.522585Z digest=sha256:ea373555bbbe37aaf8940ec373e0384b3a79fcac1026afa0b2bc23d18d4dc6b2

Observation d3e7a6cb-a8d5-4fa7-9907-7ae932fecda9 · outbound

This paper cites TVQA: Localized, Compositional Video Question Answering.

Benchmarking Vision-Language Models on Optical Character Recognition in Dynamic Video Environments TVQA: Localized, Compositional Video Question Answering

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T15:26:35.526768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:26:35.526768Z digest=sha256:2ab6e2ee6fa0b6e17f66c9bb9651645775a1f7696c2a861687f3719bdcb09393

Observation 516c8911-650f-42e8-8ff9-98f89b5a6e0a · outbound

This paper cites GPT-4o: Omni-modal language model, 2024.

Benchmarking Vision-Language Models on Optical Character Recognition in Dynamic Video Environments GPT-4o: Omni-modal language model, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:26:35.651666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:26:35.530884Z digest=sha256:019a0fdcac5400f1d3ee282d82cae512a89fc38ca8d5267f643b18979b121e25

Observation 43496147-274f-406b-9bc9-9ecfa6285dcb · outbound

This paper cites PaddleOCR: An OCR toolset based on PaddlePaddle, 2023.

Benchmarking Vision-Language Models on Optical Character Recognition in Dynamic Video Environments PaddleOCR: An OCR toolset based on PaddlePaddle, 2023

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:26:35.641320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:26:35.534603Z digest=sha256:f60507ed0bb474a88eda4054cf2b86bd2338eb875dd0dee3d4797f11a15ed8d8

Observation 214282ec-339f-4b66-a61f-8a46125e10e7 · outbound

This paper cites RapidOCR: A lightweight OCR framework, 2021.

Benchmarking Vision-Language Models on Optical Character Recognition in Dynamic Video Environments RapidOCR: A lightweight OCR framework, 2021

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:26:35.631189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:26:35.538170Z digest=sha256:ccfea8db8a8ca261d90aabeb6dc92edfc9e4af4521536501e540e4a97e87a489

Observation bf814452-f886-4d36-b80e-f50f86e588f2 · outbound

This paper cites VideoDB: Video infrastructure for the AI first world, 2024.

Benchmarking Vision-Language Models on Optical Character Recognition in Dynamic Video Environments VideoDB: Video infrastructure for the AI first world, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:26:35.621872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:26:35.542032Z digest=sha256:3006e10f9bf6e482ef4d0605185be9df7bad8203894d0d56309762c3113a41ac

Observation baacb20c-5206-4603-bbe7-40ebcaaf3f32 · outbound

This paper cites CONTEXT"? OF.

Benchmarking Vision-Language Models on Optical Character Recognition in Dynamic Video Environments CONTEXT"? OF

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:26:35.612278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:26:35.545652Z digest=sha256:f6a8b7411d9fec2622330acf1aa75bc9991ad483b193e44cb676510eac9122ed

Pith citing papers

Observation 0e34a902-f194-46ac-b30e-8a3c051e7050 · inbound

E-ARMOR: Edge case Assessment and Review of Multilingual Optical Character Recognition cites this paper.

E-ARMOR: Edge case Assessment and Review of Multilingual Optical Character Recognition Benchmarking Vision-Language Models on Optical Character Recognition in Dynamic Video Environments

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T10:52:13.947049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:52:13.947049Z digest=sha256:d03d64554d01853684d7cedcec6cbddb695f9cd72aa1ddfcc43e33a3d9354d80

Observation 4723749e-4596-4773-9479-2819028e83d0 · inbound

If you're waiting for a sign... that might not be it! Mitigating Trust Boundary Confusion from Visual Injections on Vision-Language Agentic Systems cites this paper.

If you're waiting for a sign... that might not be it! Mitigating Trust Boundary Confusion from Visual Injections on Vision-Language Agentic Systems Benchmarking Vision-Language Models on Optical Character Recognition in Dynamic Video Environments

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:51:02.632753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T02:54:35.336251Z digest=sha256:2ea1bbf54845eb5a28fe9ad90e501ec331c54298d0be15c485d157c312e1a881