Pith. sign in

Paper Citation Record · LEDGER

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

As of 18 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 18 inbound Pith citation observations for arXiv:2412.02210.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.02210 v3

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T23:47:10.303241Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:31:36.645575Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T05:47:41.904748Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact0
  • verified fuzzy30
  • unresolved18
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4d312f58-a6a5-4c9c-9810-649d1a422ac7 · outbound

This paper cites an unresolved cited work.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-11T23:47:11.388706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:47:10.019976Z digest=sha256:1264ebc7eff57709b18fc7dedd1b406c953371a85df43f9661671355c12211d0

Observation 31aa5d48-02f9-4e3f-842d-dd6422a2d22f · outbound

This paper cites GPT-4 Technical Report.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T23:47:10.026585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:47:10.026585Z digest=sha256:b06a958894cc7c640e4be06704d60e91840443aa4de314ce5222ff37d51c8fbf

Observation 2e9d4d9c-600d-42f8-9dad-dcdbd68c47c7 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T23:47:10.033823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:47:10.033823Z digest=sha256:c04e7d48907f6e03e36089ad1698d7995a2e86efe2c54a205db66cf880cf715e

Observation 1ee5e262-a344-4d7b-9077-902b4abb54ce · outbound

This paper cites Nougat: Neural Optical Understanding for Academic Documents.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Nougat: Neural Optical Understanding for Academic Documents

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T23:47:10.040694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:47:10.040694Z digest=sha256:556489779e542431ffc0638b1147b61a845799311364e512be9755c4262bb7a8

Observation 0ccdb98b-af15-4edf-b688-0495748d2e3c · outbound

This paper cites Onechart: Purify the chart structural extraction via one auxiliary token.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Onechart: Purify the chart structural extraction via one auxiliary token

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:11.368604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:47:10.047209Z digest=sha256:7ab43212e19561689ef219c11a687266e407b33369faed4959f81a9e8c3757a1

Observation 85f1cf64-7422-4120-b920-959fd9266711 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:11.350282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:47:10.053297Z digest=sha256:17920145ab88863f0aef28f7def23576b8577ff65529a19436812ee3bcdf62c8

Observation 1b189e01-9fe0-4e59-9b2d-6623439d76d3 · outbound

This paper cites Total-text: A com- prehensive dataset for scene text detection and recognition.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Total-text: A com- prehensive dataset for scene text detection and recognition

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:11.330529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:47:10.060223Z digest=sha256:cc5675a3f965a13366a243c3696b25635f90a1b7ceb499bf1cfe212617b18073

Observation 8700ec30-408a-4a54-928e-c9ac28f40d6e · outbound

This paper cites Icpr2018 contest on robust reading for multi- type web images.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Icpr2018 contest on robust reading for multi- type web images

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:11.308637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:47:10.066507Z digest=sha256:ae25a96e4f0134d65b56da1b9f4eecc8d2d65c8709882d8d6fd5280f0dd0b91d

Observation 6134128f-727b-48ea-b3de-13e1898d02a4 · outbound

This paper cites mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T23:47:10.073378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:47:10.073378Z digest=sha256:f354b87cf5e8e5cd35e1b23d373b15bd1f8648f82022b98d4b88f1b8d63bb6fa

Observation 00f847f6-4254-4297-9a4c-3ec2ff245a1e · outbound

This paper cites an unresolved cited work.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-11T23:47:11.290550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:47:10.079322Z digest=sha256:296dd4190e3abfd68f20552ae06cfd99a59a5fd459d599dda8a29f17e4831a37

Observation 9a410f80-8faf-42e5-909b-6e6226005c6d · outbound

This paper cites Post-ocr parsing: building simple and robust parser via bio tagging.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Post-ocr parsing: building simple and robust parser via bio tagging

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:11.271230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:47:10.085071Z digest=sha256:a7dddfad9dd1d2137598be41d3d89887c2af7d4e577ff4e5898d646678d1ed5b

Observation 15702156-a23a-436d-971a-6e4dfcfcccaf · outbound

This paper cites Funsd: A dataset for form understanding in noisy scanned documents.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Funsd: A dataset for form understanding in noisy scanned documents

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:11.253184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:47:10.090339Z digest=sha256:a6dd31c9c515670f062fd6dac7e047cbef33fae15e6ab0083c99935a49125eaa

Observation a83c0578-8d9f-49fc-a64c-e6378e62c762 · outbound

This paper cites Icdar 2015 competition on robust reading.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Icdar 2015 competition on robust reading

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:11.236087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:47:10.095690Z digest=sha256:558083ab3e73fd578dd139903e953aff8c99599af0cd239cf235ab711f233490

Observation 23ba1a9b-30ac-4852-beab-aa6af8349b5c · outbound

This paper cites Ocr-free document understanding transformer.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Ocr-free document understanding transformer

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:11.216001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:47:10.100963Z digest=sha256:96f6b59531ec69f985ce77043c11b7d6dd9add1a34b4e7eaf6758bb099e922b0

Observation 77421711-b0f8-4c55-b65b-f8a221ac9727 · outbound

This paper cites Visual information extraction in the wild: practical dataset and end-to-end solu- tion.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Visual information extraction in the wild: practical dataset and end-to-end solu- tion

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:11.196174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:47:10.106119Z digest=sha256:40ad300bc2a7d8b4dcfd37972097dbae064582f6f99ac9baaf926353bc128195

Observation 0fa47a8c-203b-4719-a464-ed83920f3e01 · outbound

This paper cites Binary coors capable or ‘correcting deletions, insertions, and reversals.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Binary coors capable or ‘correcting deletions, insertions, and reversals

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:11.176016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:47:10.111356Z digest=sha256:aa572d23304263e24cb5662853a20ef102054784de95c930e9c80a1d0b811735

Observation 4b8f8499-f103-4abe-b669-406e7323761b · outbound

This paper cites TableBank: Table benchmark for image- based table detection and recognition.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy TableBank: Table benchmark for image- based table detection and recognition

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:11.154803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:47:10.116371Z digest=sha256:4e2c89ce456d1f7110ff328fba8a134e101dadacb7c1801ab17ed8a01fd6185b

Observation 87624657-47b0-4cdc-9cf4-d1c3dea777f0 · outbound

This paper cites Focus Anywhere for Fine-grained Multi-page Document Understanding.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Focus Anywhere for Fine-grained Multi-page Document Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T23:47:10.121280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:47:10.121280Z digest=sha256:cf6706a47f15defde12eb1f67c482ba28381c6b93683b531b5fa8e827cbcdd5c

Observation 19a303e8-bf54-415e-a658-ed2a8f03fdf1 · outbound

This paper cites OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T23:47:10.126191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:47:10.126191Z digest=sha256:89b7780e914a0f005bead0fb73bd38dedefe3fd638fa93cdc2d6cec42c4480a6

Observation e5bd558f-d825-4799-b185-fdde556e1d39 · outbound

This paper cites Spts v2: single-point scene text spotting.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Spts v2: single-point scene text spotting

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:11.136303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:47:10.131612Z digest=sha256:4f169044644409add7137ae5ba54e6bf8fe42f07333f61625b2a92caf20dc33e

Observation 76f39e67-cb53-499a-8984-cc1e90f7ec3f · outbound

This paper cites TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T23:47:10.136168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:47:10.136168Z digest=sha256:487796714a51fc9100d27f82ce4e12d8d0f1f8d8e3ef6ea69f0e5b9aad24be3b

Observation 2a484f35-69bc-460f-9ce5-d3e44ed11fe5 · outbound

This paper cites Parsing table structures in the wild.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Parsing table structures in the wild

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:11.114773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:47:10.141463Z digest=sha256:a4ac44719b8e894df3d73982a2067974f571cdcd42dd2c48bf63540cc2f8a7b1

Observation fa1ea388-b6cb-47d1-8f54-40660dbb9143 · outbound

This paper cites Towards end-to-end unified scene text detection and layout analysis.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Towards end-to-end unified scene text detection and layout analysis

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:11.095493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:47:10.147473Z digest=sha256:c3a775fd9e4780d8228daef540077864ec58092fd6dd4eed508ee4307685a5d2

Observation fd6e2449-2c2f-44fb-93cc-a3bb9f6f235b · outbound

This paper cites Layoutllm: Layout instruction tuning with large language models for document understanding.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Layoutllm: Layout instruction tuning with large language models for document understanding

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:11.075697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:47:10.152557Z digest=sha256:d65cf6ce4bd5d486c11ea300612be21e84166c94941330ef0e88c81ce2e98593

Observation 92df3bc4-0195-4ce9-bdc4-f6412729c6ed · outbound

This paper cites KOSMOS-2.5: A Multimodal Literate Model.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy KOSMOS-2.5: A Multimodal Literate Model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T23:47:10.158358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:47:10.158358Z digest=sha256:78c182428e8d7c6b4b2780aae46389916fc009711efaea1129a30c80d0371a0c

Observation 18dc72a7-fac3-4fa8-853d-ea813e396e4c · outbound

This paper cites Icdar 2019 crohme+ tfd: Competition on recognition of handwritten mathematical expressions and typeset formula detection.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Icdar 2019 crohme+ tfd: Competition on recognition of handwritten mathematical expressions and typeset formula detection

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:11.056054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:47:10.164365Z digest=sha256:cea45183bc57057e1b0e92149aabac2ea835e0f25a359851cc871b23035363ea

Observation b149ca7d-833d-4d4d-a03b-f1f8b2fd010c · outbound

This paper cites The iam-database: an english sentence database for offline handwriting recognition.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy The iam-database: an english sentence database for offline handwriting recognition

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:11.037443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:47:10.170323Z digest=sha256:37b073d7f57a2882d3671322a9ace9fa8ab754711c6b0eb8070c9d25b6df2985

Observation 8abdbd23-e6b1-4282-9347-94fe9b313cca · outbound

This paper cites Docvqa: A dataset for vqa on document images.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Docvqa: A dataset for vqa on document images

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T23:47:10.176206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:47:10.176206Z digest=sha256:c919d7c54201ec7ec464c41d754570c5565693fcb8cf6d0682cfffb9e07156c8

Observation 769ae03a-6242-4162-92e2-935b755750e9 · outbound

This paper cites Scene text recognition using higher order language priors.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Scene text recognition using higher order language priors

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:10.999313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:47:10.181921Z digest=sha256:655d2df92d05a0a1363b17844899cb1d1acbb759f7c93f3be0bd976afade0dcb

Observation a766abb9-13c1-4085-8164-5ac46b60b34b · outbound

This paper cites Icdar2019 robust reading challenge on multi-lingual scene text detection and recognition—rrc-mlt-2019.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Icdar2019 robust reading challenge on multi-lingual scene text detection and recognition—rrc-mlt-2019

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:10.970598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:47:10.187350Z digest=sha256:f0545ed911cb1ee3c6fd219bb8cf2427a689deed14fbb8387cda3157e0b49a89

Observation fe6323e2-5fdb-4174-a558-21e2c660c109 · outbound

This paper cites Cord: a consol- idated receipt dataset for post-ocr parsing.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Cord: a consol- idated receipt dataset for post-ocr parsing

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:10.944130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:47:10.192956Z digest=sha256:3fe3fdef939d2a3c7ce726f1a2476922f96334c86a359dfffcb87103fe8bea43

Observation 1b6993c0-e907-459c-9b38-e425299e7a76 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next gen- eration image-text models.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Laion-5b: An open large-scale dataset for training next gen- eration image-text models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T23:47:10.198930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:47:10.198930Z digest=sha256:d6c5321e0489ad7cf6a33421a3d205ef51ede2bc8e2b3b6601d1f634bb10372b

Observation 0cb5d40d-e1cc-4fcc-ae6a-e02835439dff · outbound

This paper cites Icdar2017 competition on reading chinese text in the wild (rctw-17).

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Icdar2017 competition on reading chinese text in the wild (rctw-17)

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:10.893043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:47:10.204545Z digest=sha256:e02765710f4f99d5310fa047d4e37a9a8315373b04c5904fc7550858195bf5ac

Observation 5f0525b2-d90d-4968-ab5d-d7a0784162fd · outbound

This paper cites Towards vqa models that can read.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Towards vqa models that can read

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:10.873328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:47:10.211029Z digest=sha256:2dbb2874184cd20e4ae8f7cc8f22d9a596f4fba818e237306efda849ece12cdd

Observation 1ce07450-1c4a-4083-8ba2-6ee9a0e841f8 · outbound

This paper cites Seglink++: Detecting dense and arbitrary- shaped scene text by instance-aware component grouping.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Seglink++: Detecting dense and arbitrary- shaped scene text by instance-aware component grouping

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:10.850688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:47:10.216475Z digest=sha256:07536f7c4fdd2647d92a9cb4ad37efea4177c65d7a953c88465a56baf9797009

Observation 9378978c-3f0c-47be-9c47-877ffb7c6e69 · outbound

This paper cites Mtvqa: Benchmarking multilingual text-centric visual question answering, 2024.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Mtvqa: Benchmarking multilingual text-centric visual question answering, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:10.831095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:47:10.224158Z digest=sha256:c72efd402bcbba74e27dd4d074e793a4281cf07df2bb99606f8256c32da0aee3

Observation 22ba9c20-d406-4224-91c4-1a0f4c486bd3 · outbound

This paper cites Unifying vision, text, and layout for universal document processing.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Unifying vision, text, and layout for universal document processing

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:10.810874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:47:10.230406Z digest=sha256:ae5b0100edbd8d829a636e08c14ec9e11979e6080e81fbdb07d33bb976f69163

Observation 9908f796-e438-499c-aaa8-a800c6c338f2 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Gemini: A Family of Highly Capable Multimodal Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T23:47:10.236564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:47:10.236564Z digest=sha256:acfd8279b469644a6388789e44a893beecf95a688d6704657837a01d276eb7e4

Observation 3feefc3a-554d-41bd-b8ad-2f0155313e24 · outbound

This paper cites Towards Robust Visual Information Extraction in Real World: New Dataset and Novel Solution.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Towards Robust Visual Information Extraction in Real World: New Dataset and Novel Solution

Reference 39

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T23:47:10.480277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:47:10.243109Z digest=sha256:595d31ab9e4120b8d5db9a909a3a8bca3d06b4b544f3319b42b1ea182572e9ec

Observation 857d57d6-d5bb-44be-91f4-ced6d5d97b6b · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T23:47:10.248636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:47:10.248636Z digest=sha256:8f86e5391d0361fdd7ee322eb565090a3640a28a54aa27ecd8545ef7241bd2ce

Observation f1ddf2ba-4fa9-4ea4-a4de-3b59de155505 · outbound

This paper cites General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T23:47:10.254073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:47:10.254073Z digest=sha256:09c500961218bf735684cbd8e6d210ba30321065cabbb6736749574588661a0e

Observation 1a4e9fe0-786e-4466-94ef-e64971c3a63e · outbound

This paper cites Florence-2: Advancing a unified representation for a variety of vision tasks.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Florence-2: Advancing a unified representation for a variety of vision tasks

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:10.792916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:47:10.259956Z digest=sha256:fcd8a7e68862cc64a781fb7dc8456b0d4f7a18479ee41c1d93f07f44a4e7bb97

Observation 13e1739b-7b3f-4fa0-a0e7-d06b2bf93856 · outbound

This paper cites Modeling entities as semantic points for visual information extraction in the wild.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Modeling entities as semantic points for visual information extraction in the wild

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:10.772300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:47:10.264455Z digest=sha256:78bfe4dc2f30f635ab137846f869e8d95ec95b3718fc7255017f83b9f71e3a43

Observation 21a08f2e-8c81-40b5-a8dc-1ed1bb00eca0 · outbound

This paper cites Dptext-detr: Towards better scene text detection with dynamic points in transformer.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Dptext-detr: Towards better scene text detection with dynamic points in transformer

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:10.753010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:47:10.270597Z digest=sha256:f2373fcb7890368399a9c23d0c9447927189c33a74f07e78069ad9da85ddd618

Observation 97bd656e-09d2-49df-b599-6dab9842d4cd · outbound

This paper cites Icdar 2023 competition on structured text extraction from visually-rich document images.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Icdar 2023 competition on structured text extraction from visually-rich document images

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:10.735243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:47:10.275337Z digest=sha256:5f4d8495b9e882f80e4195b9de4358744f694fdfd82fc12db286716d909fd5b8

Observation f460d0ff-ae55-4d5e-9bba-2f682d51c8b9 · outbound

This paper cites Syntax-aware network for handwritten mathematical expression recognition.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Syntax-aware network for handwritten mathematical expression recognition

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:10.716552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:47:10.280683Z digest=sha256:27b2c3f3e44d455afa59b0894edda7ab666f789e7f57a01ee2aae3899f88ff09

Observation c139026f-6b05-4428-96e3-01332c86666b · outbound

This paper cites Detecting Curve Text in the Wild: New Dataset and New Solution.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Detecting Curve Text in the Wild: New Dataset and New Solution

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T23:47:10.286185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:47:10.286185Z digest=sha256:03190ff9deb0c485c20e6de555a11f1af5e2c34c3f6f98a9e1156a7005e70d7e

Observation 22305401-e979-41f9-b576-29c2070a894e · outbound

This paper cites TabPedia: Towards Comprehensive Visual Table Understanding with Concept Synergy.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy TabPedia: Towards Comprehensive Visual Table Understanding with Concept Synergy

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T23:47:10.291965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:47:10.291965Z digest=sha256:b4e2d1a763b6ba53dd5fc950114373b7db92ce5615213c47924f153bb74ace2a

Observation 9018851f-08b9-4ffd-ac40-e4d3fc27a634 · outbound

This paper cites DocLayout-YOLO: Enhancing Document Layout Analysis through Diverse Synthetic Data and Global-to-Local Adaptive Perception.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy DocLayout-YOLO: Enhancing Document Layout Analysis through Diverse Synthetic Data and Global-to-Local Adaptive Perception

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T23:47:10.297668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:47:10.297668Z digest=sha256:aff765ed2a519d5a333425402858e4eab2ffd22c2d9a3c8b2865fc485e6d8df0

Observation 9edd0b2c-ef4c-4f1d-8e9b-c4918a5ce15b · outbound

This paper cites 3.5 kg”, the true value is “3.5kg.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy 3.5 kg”, the true value is “3.5kg

Reference 50

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T23:47:10.696214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:47:10.303241Z digest=sha256:d536bfbae0122f1f217cd0c49504ab310bca1350666f64116711dcd0dd531dd7

Pith citing papers

Observation 44a67d08-3b66-41b7-a987-03b28e41f3aa · inbound

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning cites this paper.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:33:26.764403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:a1ba77a1e6c54aa424135219e88ff1843304900cace09f6101a55416b3d9d7d8

Observation 1e8e66b1-1f2e-478f-b609-8a51f5b3f20c · inbound

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? cites this paper.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.645575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.645575Z digest=sha256:b52c9356877e2ff574a4ee62a5a5985dbf32ed665387402e06f91a3ba49bddfa

Observation 736416f8-5d2a-4a1a-8b65-af190c10210b · inbound

OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning cites this paper.

OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T14:57:21.254510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:57:21.254510Z digest=sha256:6a2aa13010cc27adeeebceeb065c267d8e50ba74f1fac7d57f9856d77f546915

Observation 2f7b4b42-248e-43ad-ab69-7b0e71a560c5 · inbound

MMTABREAL: Real-World Benchmark for Multimodal Table Understanding cites this paper.

MMTABREAL: Real-World Benchmark for Multimodal Table Understanding CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:21.170204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:21.170204Z digest=sha256:1fb7bee37c0f846a1cb5933ba7cf78778b7727f5beabf326efb23bd6ad5e2c48

Observation c4292ea1-2568-4750-ae9c-1109fad454ab · inbound

Rethinking Multilingual Vision-Language Translation: Dataset, Evaluation, and Adaptation cites this paper.

Rethinking Multilingual Vision-Language Translation: Dataset, Evaluation, and Adaptation CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:38.470496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:09:38.470496Z digest=sha256:8d73d4015f95c80b32ed2b1b370fdf139b7343260f9cfd278a6e2580f1c98fbf

Observation 964a464a-ceb5-4326-8e21-cf34afe55792 · inbound

Position: Reasoning After Perception Means Reasoning Without Vision cites this paper.

Position: Reasoning After Perception Means Reasoning Without Vision CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:02.315456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:02.315456Z digest=sha256:0d260dfb5ce73fd393a6de58379067ac872e4a4a886b27d0f6eae2588c2e74c2

Observation 0b2729ac-0c6d-42a0-a2a7-401f20bcb0e0 · inbound

MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models cites this paper.

MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-05T23:03:10.162023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:03:10.162023Z digest=sha256:afed9aa3011db91d13f842649e08229187d7e58173cfa62bbc601364041dc8fc

Observation 79d6a524-fbf5-4e88-97de-bc2ff4c5d8ac · inbound

Training-Free Multimodal Large Language Model Orchestration cites this paper.

Training-Free Multimodal Large Language Model Orchestration CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-19T00:12:54.173395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T00:12:39.834892Z digest=sha256:900ef260b9f5f7343e266474bac3c7a77a21dcf71d8aa2d341aa5651531aea76

Observation 4a312208-d36e-4512-bc46-8e0e4077c625 · inbound

Training-Free Multimodal Large Language Model Orchestration cites this paper.

Training-Free Multimodal Large Language Model Orchestration CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-25T08:05:30.767369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-25T08:02:15.950975Z digest=sha256:2b11596bbe9fdb38de059f53f0546b757323251833360e374407fe23564e6f88

Observation 8875ee5a-2409-47cc-b173-7ab2f4a3f511 · inbound

E-ARMOR: Edge case Assessment and Review of Multilingual Optical Character Recognition cites this paper.

E-ARMOR: Edge case Assessment and Review of Multilingual Optical Character Recognition CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T10:52:13.943512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:52:13.943512Z digest=sha256:b0c806d27a7915b485143707a30eb436b32a7f6ed25a2b81ad158caa60a0c328

Observation 5740e854-0131-47fa-9469-3996b278d76e · inbound

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing cites this paper.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:25:32.078836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:2ee021e4442c89495d8333d75415d16762709b765e9f63373c1bc482945e3c1e

Observation 5ddc59e4-c439-4678-8284-13ebfcc68c12 · inbound

Physical Plausibility Reasoning via HCM-GRPO: Empowering Compact Model for Superior Performance cites this paper.

Physical Plausibility Reasoning via HCM-GRPO: Empowering Compact Model for Superior Performance CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T22:36:09.992374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:36:09.992374Z digest=sha256:67118b1902eb8592491ba418b370ef3640efca93805b355815a4fe0a498d9d11

Observation cfe939af-8412-4dfb-ab86-9f248e99c88e · inbound

FinCriticalED: A Visual Benchmark for Financial Fact-Level OCR cites this paper.

FinCriticalED: A Visual Benchmark for Financial Fact-Level OCR CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:20:11.506437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-17T20:19:52.701262Z digest=sha256:599d2aad97e27fb3f68d3c14c070b5dcfad2ba7f8afd36c10f696fd5741d4a4b

Observation a18699c5-30c6-40c8-910e-6ffe2189aff8 · inbound

Towards Real-World Document Parsing via Realistic Scene Synthesis and Document-Aware Training cites this paper.

Towards Real-World Document Parsing via Realistic Scene Synthesis and Document-Aware Training CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:18:26.614884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T01:15:26.757215Z digest=sha256:f385dfeb7b71aac7ce2e7381bd7745ac726d1b92b6ca219db490361329a0cc03

Observation 947b5532-e2b7-4993-a8c3-a44d4155f191 · inbound

LatentRouter: Can We Choose the Right Multimodal Model Before Seeing Its Answer? cites this paper.

LatentRouter: Can We Choose the Right Multimodal Model Before Seeing Its Answer? CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:47:04.327801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T01:42:54.802658Z digest=sha256:1adb360e956716d5ca6981dc0e1dfb467ffbb6e19365fac721c2d572447391e1

Observation 5a521e15-f43c-4ae3-b776-ccd1164a5765 · inbound

METATR: A Multilingual, Evolving Benchmark for Automatic Text Recognition cites this paper.

METATR: A Multilingual, Evolving Benchmark for Automatic Text Recognition CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T18:53:51.525916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T18:48:45.623530Z digest=sha256:760ba8f895f26f79ef704989248c588aec443db02f6e82b5e0e933b5ac48f2c5

Observation ef1ae605-bc50-4308-b92b-e56086257121 · inbound

Towards Fully Automated Exam Grading: Fairness-Aware Recognition of Handwritten Answers with Foundation Models cites this paper.

Towards Fully Automated Exam Grading: Fairness-Aware Recognition of Handwritten Answers with Foundation Models CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T05:47:41.906627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-27T13:03:07.941337Z digest=sha256:942292f09ae8768b545b966749afa7ac0e6f59057a5336b68e68bc7d3b1f3da9

Observation 1aa6dd25-2830-4ce1-8020-ba4edd373b63 · inbound

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model cites this paper.

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 131

Resolution
unresolved
no resolver link, observed 2026-07-31T06:20:14.141939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:20:14.141939Z digest=sha256:65b30ba22fc233a05d33da64defae8cf25e7c5392c72062f86bc64e7a044eefc