Pith. sign in

Paper Citation Record · LEDGER

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment

As of 21 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2412.12902.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.12902 v2

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T13:40:39.718350Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

54 of 54 outbound references displayed

  • verified exact1
  • verified fuzzy30
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 66ec3495-180b-4d7b-a51c-83e6a5ce15bc · outbound

This paper cites Docformer: End-to-end transformer for document understanding.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Docformer: End-to-end transformer for document understanding

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T13:40:39.484541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:40:39.484541Z digest=sha256:b9c16ce9c563d84f183421a8c23826b5a024659aeb30a36009c19fca5c3c5ba9

Observation 8fd7490b-cab5-4011-a31a-31d79b7f9cbf · outbound

This paper cites Visual and textual deep feature fusion for document image classification.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Visual and textual deep feature fusion for document image classification

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:40.434122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T13:40:39.490094Z digest=sha256:8f8017d59cfe9636878b962f000768ea7ec8b009b601f6a24dffe28e506c7228

Observation 9bea97bc-1f37-4130-a7d9-2aaf2f2a2339 · outbound

This paper cites Eaml: Ensemble self-attention-based mu- tual learning network for document image classification,.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Eaml: Ensemble self-attention-based mu- tual learning network for document image classification,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:40.420940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T13:40:39.494541Z digest=sha256:0f44047be0c59a8f413a32f3e4a2b88f504b00d4d606af6c911dcab75b4d8835

Observation 0333c493-8e97-4c62-9425-584319afcd0e · outbound

This paper cites BEiT: BERT Pre-Training of Image Transformers.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment BEiT: BERT Pre-Training of Image Transformers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T13:40:39.499670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:40:39.499670Z digest=sha256:422149a03babe1d1a167b298c5794bccb5f1c31a5e30b0a705012da3396324c1

Observation e9c3414a-0dbb-4003-9c61-65d09ea10892 · outbound

This paper cites Gritsenko, Matthias Minderer, Charles Blundell, Razvan Pascanu, and Jovana Mitrovi´c.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Gritsenko, Matthias Minderer, Charles Blundell, Razvan Pascanu, and Jovana Mitrovi´c

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:40.407110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T13:40:39.505431Z digest=sha256:3693aabfac3509400780a7783d9067840d339c368b64bbf167d95654d6bde03c

Observation b65dd56a-4296-4227-854b-d3697cd3fa76 · outbound

This paper cites Cascade r-cnn: High quality object detection and instance segmentation.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Cascade r-cnn: High quality object detection and instance segmentation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:40.391229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T13:40:39.510451Z digest=sha256:b59c1cdb2d61608986457ada5e55c6690d8bf28dae97d5228bd74c215a3ad2f3

Observation 682cc618-ae98-41f9-bca2-cc528aa83daf · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Emerg- ing properties in self-supervised vision transformers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T13:40:39.515084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:40:39.515084Z digest=sha256:b9623cab7d67625753856e7d5baea90fa3f8415e9ba978f79590158d52b84665

Observation 8400b773-7954-4b58-8269-14422081365a · outbound

This paper cites Conceptual 12m: Pushing web-scale image-text pre- training to recognize long-tail visual concepts.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Conceptual 12m: Pushing web-scale image-text pre- training to recognize long-tail visual concepts

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:40.361518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T13:40:39.520754Z digest=sha256:c419f8021af7aeb889ab40a3c7f75e91d8588c6875dbe5175251f7efc14c73e0

Observation 49f50b91-07a4-4063-9c8f-30f4706b3949 · outbound

This paper cites A simple framework for contrastive learning of visual representations.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment A simple framework for contrastive learning of visual representations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T13:40:39.525119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:40:39.525119Z digest=sha256:6560d8794e6f17fcfab687853aa181fa4f1fa458caa6e39c08e507e28f1fcff4

Observation 89747ae7-cac3-433a-8592-12cfd1b1d7e9 · outbound

This paper cites Big self-supervised mod- els are strong semi-supervised learners.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Big self-supervised mod- els are strong semi-supervised learners

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:40.327008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T13:40:39.529747Z digest=sha256:2edf93ccdb1f6336d409afc6b323426806718302a974d1c66457772b2c57b9eb

Observation f32e35b1-f074-45e7-a5b6-6eb998f3ce3a · outbound

This paper cites Uniter: Universal image-text representation learning.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Uniter: Universal image-text representation learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T13:40:39.534111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:40:39.534111Z digest=sha256:7a73b4a3fc012ecec301926ed41b6aeb394698fe5b1be1fc053c2518bfcb06ab

Observation 67578e09-632f-4fae-a273-f183d55d4f1f · outbound

This paper cites Vision grid transformer for document layout analysis.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Vision grid transformer for document layout analysis

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:40.304685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T13:40:39.538610Z digest=sha256:a75e4583989905684eae3b81ac74fcfeb52e5ec84aaa6e6fd2316a3a96858b6a

Observation 70680e1c-a978-45ea-84b2-2887f04cb98f · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment An image is worth 16x16 words: Transformers for image recognition at scale

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:40.291726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T13:40:39.542811Z digest=sha256:8ea011cedb89416539efc8c59ff8ac68124ce18ccb69546f6e87acae9614bdc0

Observation b16bcab3-423f-4d11-bfad-252ed5f3bbe1 · outbound

This paper cites Bootstrap your own latent-a new approach to self-supervised learning.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Bootstrap your own latent-a new approach to self-supervised learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T13:40:39.546905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:40:39.546905Z digest=sha256:733b28aa0e4ff8eff5c517c476fe4c8cf8359808b8c572a7c3188fecc8ecc952

Observation e633ae88-9220-4280-8de8-f86f6947bfe1 · outbound

This paper cites Evaluation of deep convolutional nets for document image classification and retrieval.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Evaluation of deep convolutional nets for document image classification and retrieval

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:40.263374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T13:40:39.551384Z digest=sha256:a96cbd07a97a6c35d3137d173cd107f0c557669fa472047ab095bf640efd269c

Observation 8d25c0fe-45d6-41a7-afdd-f20c1c32357d · outbound

This paper cites Momentum contrast for unsupervised visual rep- resentation learning.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Momentum contrast for unsupervised visual rep- resentation learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T13:40:39.556434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:40:39.556434Z digest=sha256:1cc436ce61e2e5498ffc21314bbff7edbbebc3868e67d3d6b23ede7262305aad

Observation a02407b9-2880-48ab-b9de-61c0d382e819 · outbound

This paper cites Masked autoencoders are scalable vision learners.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Masked autoencoders are scalable vision learners

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:40.236336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T13:40:39.560702Z digest=sha256:c48a9ab30c7bba1698e4050e1bd02091235a3f734283a66b676c5decb2e13b45

Observation 3383098b-8d62-40c2-869e-743b5fe97593 · outbound

This paper cites Bros: A pre-trained lan- guage model focusing on text and layout for better key infor- mation extraction from documents.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Bros: A pre-trained lan- guage model focusing on text and layout for better key infor- mation extraction from documents

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:40.222480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T13:40:39.564700Z digest=sha256:48ee830495e5735d721b1ee7414b6aa654f3bf4ceb2a7a563740be50fd07dff6

Observation 60f325e4-3b2f-490f-a045-690fcb2b1e8b · outbound

This paper cites Layoutlmv3: Pre-training for document ai with unified text and image masking.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Layoutlmv3: Pre-training for document ai with unified text and image masking

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T13:40:39.569129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:40:39.569129Z digest=sha256:34b2d45c2fa704222990a2b6d63f7d419ee68bfeaf74ee5691cab2aadd5994d3

Observation 323b5738-e0c2-494a-bf73-211acd795769 · outbound

This paper cites Icdar2019 compe- tition on scanned receipt ocr and information extraction.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Icdar2019 compe- tition on scanned receipt ocr and information extraction

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:40.200010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T13:40:39.573650Z digest=sha256:a24d9e2514829ccf115b7304a6856d30e5e0c690bd942a5757292644a14e9b7c

Observation 4b133606-b0da-42ab-a007-0f1a1399d918 · outbound

This paper cites an unresolved cited work.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:40:40.184396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T13:40:39.579360Z digest=sha256:fe25786025f701fd285884c2e297bd208b01d88031b91d92c5eee40a21294f17

Observation 444b683c-04af-42ed-a478-d1265eb77b62 · outbound

This paper cites Funsd: A dataset for form understanding in noisy scanned documents, 2019.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Funsd: A dataset for form understanding in noisy scanned documents, 2019

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:40.170443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T13:40:39.584087Z digest=sha256:da6a630fc5f755e987cbe22bee2ec072cdd0173555c1a3ccbf14167b57e289dd

Observation 4d4173ea-df47-4d37-9d8f-d67018a45527 · outbound

This paper cites Scaling up visual and vision-language representa- tion learning with noisy text supervision.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Scaling up visual and vision-language representa- tion learning with noisy text supervision

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T13:40:39.588594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:40:39.588594Z digest=sha256:55a3601d87cc6fd5a9ba084a7a5f71850090f1154ce32cb5ecff7c1bb014dca3

Observation ea22a01c-c224-4ff8-8aa3-aaa0034c5661 · outbound

This paper cites Ocr-free document understanding transformer.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Ocr-free document understanding transformer

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:40.145285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T13:40:39.592918Z digest=sha256:66e8b30626c75e9dbc3b65c87f80efd2355ffe7aafbcf848384e45416b35ac54

Observation ea3e7d8c-81de-4470-aa51-c34996a3bc19 · outbound

This paper cites Dit: Self-supervised pre-training for docu- ment image transformer.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Dit: Self-supervised pre-training for docu- ment image transformer

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:40.130892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T13:40:39.597159Z digest=sha256:9dfded34989a617d565f29692439fd21308002a3351d078ec21f56020ddcb4c8

Observation c8d01da3-b4e8-4be9-ab56-c66f1ed1f1f8 · outbound

This paper cites Grounded language-image pre-training, 2022.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Grounded language-image pre-training, 2022

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T13:40:39.601576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:40:39.601576Z digest=sha256:cc7ddb26bc01a1d81bd79024b3017227ce03dc0d6c83d8c3900cefab8846926b

Observation bfc282d1-3098-4457-bef1-d54bcbc8b982 · outbound

This paper cites Grounded language-image pre-training.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Grounded language-image pre-training

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:40.108412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T13:40:39.605247Z digest=sha256:7ebad728579e41e5252cff888e57b8ea307820246c88b198b670969a26bbbc90

Observation c4734ac9-db7e-4647-8924-73316f72819b · outbound

This paper cites DocBank: A Benchmark Dataset for Document Layout Analysis.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment DocBank: A Benchmark Dataset for Document Layout Analysis

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T13:40:39.609154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:40:39.609154Z digest=sha256:a21a7072884d4d94268fd9f1ef2b826ab932076912667d33f7270e5e8079cdae

Observation 43b1ff01-eb7e-4580-8647-fe82aa0ad910 · outbound

This paper cites Bi-VLDoc: Bidirectional Vision-Language Modeling for Visually-Rich Document Understanding.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Bi-VLDoc: Bidirectional Vision-Language Modeling for Visually-Rich Document Understanding

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-11T13:40:39.806354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T13:40:39.613456Z digest=sha256:e0ac74a971f156b9942d13f09b171721cbb87e27de7d40feef751f83d59d265d

Observation 74d9812c-b55e-4de7-ab11-fc769ecdbf2e · outbound

This paper cites Docvqa: A dataset for vqa on document images.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Docvqa: A dataset for vqa on document images

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:40.092914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T13:40:39.617708Z digest=sha256:ce273dd2f5fc723b6854d8e4ea814eb8969a9eb17aa86b5f0ccdae8a34620098

Observation 560bf376-4f24-4416-a1b2-1b777202dd5b · outbound

This paper cites Infographicvqa.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Infographicvqa

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:40.080406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T13:40:39.621716Z digest=sha256:a91227267f0d484f4d8e6ac1f843083ff802dfa682dba2fdab20e7686290943b

Observation 953a4027-a9fa-4bcc-8324-f6f675b6ec19 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment DINOv2: Learning Robust Visual Features without Supervision

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T13:40:39.625439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:40:39.625439Z digest=sha256:1c91c7c81a183f6d4d0af164308a60a0958b0e55b66483ba868a4a84d762cc2f

Observation d869e613-2f33-43ad-9ea8-26a9094cacf4 · outbound

This paper cites {CORD}: A consolidated receipt dataset for post-{ocr} parsing.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment {CORD}: A consolidated receipt dataset for post-{ocr} parsing

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:40.066575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T13:40:39.630141Z digest=sha256:ecaca99191c17f02d57ae0d3818f5bb3fff230f96bf7a374a96b95f20e647e7d

Observation 93011028-67ea-4c9f-a3cc-353573530b1f · outbound

This paper cites Doclaynet: A large human- annotated dataset for document-layout segmentation.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Doclaynet: A large human- annotated dataset for document-layout segmentation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:40.053568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T13:40:39.633975Z digest=sha256:e429ef741eb3850b850b68295584d565b67a4910183c311e83b2c46f955a5ee5

Observation 47f192b0-31d2-4a0b-9bac-37cddd362576 · outbound

This paper cites Going full-tilt boogie on document understanding with text-image-layout transformer.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Going full-tilt boogie on document understanding with text-image-layout transformer

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:40.040932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T13:40:39.637850Z digest=sha256:caca07c535b226165817fde4ef85ce2a5602dac7bc6a31ec0867fb1f1b3e2d60

Observation fa9145ef-c848-49f8-9306-68dbc95b562f · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Learning transferable visual models from natural language supervi- sion

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T13:40:39.641900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:40:39.641900Z digest=sha256:be12e4ed2b71bcf09ae8b86d39aec4d553b075c8ba9bff5a3b8d1a4860549114

Observation 62979f90-35aa-4ffc-9a99-a2bb6acb48b9 · outbound

This paper cites Imagenet large scale visual recognition challenge.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Imagenet large scale visual recognition challenge

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T13:40:39.646043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:40:39.646043Z digest=sha256:6a3f5b3e6fb68dd3eba623d2bba00bec46952562a6e87c755b415a307d45a71d

Observation 1e2f8194-9003-49d7-8273-f33e93b20574 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Laion-5b: An open large-scale dataset for training next generation image-text models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T13:40:39.650138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:40:39.650138Z digest=sha256:86b769badabc71003ba35fa57e3d2d9288d20ff9884ede7b70b93e8ea5a4611e

Observation 65695760-db70-443f-aee2-426a6b7ce92f · outbound

This paper cites Complex document information processing (cdip) dataset, 2022.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Complex document information processing (cdip) dataset, 2022

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:40.004602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T13:40:39.654274Z digest=sha256:a3e09c87f0a1b5e686a456b471ec212f1bcf6047b9253546a7d58c779127f686

Observation 345b4aa3-8ecc-4782-8dd1-45a40c9bd442 · outbound

This paper cites Kleister: key in- formation extraction datasets involving long documents with complex layouts.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Kleister: key in- formation extraction datasets involving long documents with complex layouts

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:39.990257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T13:40:39.658478Z digest=sha256:dceb97d491e896cd20535eaf2a0b0b50252cdc73c2ffbf9ea863097f2ed9e2db

Observation 5cabec59-b2ba-4b11-8a65-3bbbc638f022 · outbound

This paper cites Vl-bert: Pre-training of generic visual- linguistic representations.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Vl-bert: Pre-training of generic visual- linguistic representations

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:39.977499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T13:40:39.662701Z digest=sha256:de53078c957d0c3846b11b56822274f6ecf9bcea96ff73b23f889315b68910d2

Observation 2028df8f-6ae4-4a69-aa05-8d38aa65529e · outbound

This paper cites Revisiting unreasonable effectiveness of data in deep learning era.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Revisiting unreasonable effectiveness of data in deep learning era

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:39.964964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T13:40:39.667525Z digest=sha256:0d2f68387d531c4658341ac635bad31a323b10048f67faf3db14a453db8e83c6

Observation 8e70fc14-c584-4731-89ec-6bc277c3b6d3 · outbound

This paper cites Unifying vision, text, and layout for universal document processing.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Unifying vision, text, and layout for universal document processing

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:39.952752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T13:40:39.671585Z digest=sha256:bdfb538d3d7ac6dd43dc4e50be185881ccb774c9101862460587cf2320b52c2a

Observation 7bb084c8-b4ae-4cd3-9a6d-86c1bb614771 · outbound

This paper cites Yfcc100m: The new data in multimedia research.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Yfcc100m: The new data in multimedia research

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T13:40:39.675226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:40:39.675226Z digest=sha256:0a8aad8aedeb06462bef47f7d4f60144796aa31fb66a897b64e2a08b46e3083e

Observation c4e96163-fed6-436d-98d9-d8d38625d646 · outbound

This paper cites Training data-efficient image transformers & distillation through at- tention, 2021.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Training data-efficient image transformers & distillation through at- tention, 2021

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T13:40:39.679383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:40:39.679383Z digest=sha256:4de781e0b048dec1c1a4e7debb855b76b6e16b69462cd3853f740b6e250c1523

Observation 91bc9e1f-cdd3-455c-81da-290aef275597 · outbound

This paper cites Detectron2.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Detectron2

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T13:40:39.683469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:40:39.683469Z digest=sha256:fdee0abe529175a80b31971771f0c033b14150f0549af00b0c11be5c371c236c

Observation b5a48ba7-8102-4142-8dab-dd75a01539d3 · outbound

This paper cites Aggregated residual transformations for deep neural networks, 2017.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Aggregated residual transformations for deep neural networks, 2017

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:39.917601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T13:40:39.687287Z digest=sha256:c4aa307a0214cc9eca29ee6217e49c4bee25b5bb1418ae41c30e4fa858a4324e

Observation 90ce9152-23a7-469c-bfe6-609d473f6bb1 · outbound

This paper cites Layoutlm: Pre-training of text and layout for document image understanding.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Layoutlm: Pre-training of text and layout for document image understanding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T13:40:39.691284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:40:39.691284Z digest=sha256:0d6e46f53a817b536774ffcce7613ccd63867b3a2dc53832274bd0762d2bec2c

Observation d8d7e7f7-c9b3-4c1b-a072-821c6f2242fb · outbound

This paper cites LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T13:40:39.695400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:40:39.695400Z digest=sha256:b45aac99c6a37460d150782bb423be051c241269c4c6b03bac96965764b98c92

Observation b6ede840-0a43-439c-b526-01ee219b47c5 · outbound

This paper cites FILIP: Fine-grained interactive language- image pre-training.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment FILIP: Fine-grained interactive language- image pre-training

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:39.895575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T13:40:39.700181Z digest=sha256:5df708eda150f2ac09b9140e08937ce383126269d6b680de2457cbbcd94b6553

Observation 3c7e747a-4390-4ab8-b9ed-b9d9b70a1ea0 · outbound

This paper cites StrucTexTv2: Masked Visual-Textual Prediction for Document Image Pre-training.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment StrucTexTv2: Masked Visual-Textual Prediction for Document Image Pre-training

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T13:40:39.704209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:40:39.704209Z digest=sha256:5da613a583ca0390554703f16eb66485a5c4c9b67da3e0cb74fd97317cea94e5

Observation 97cf0576-2402-4cb0-9cd7-0e7c19cab44c · outbound

This paper cites Sigmoid loss for language image pre-training,.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Sigmoid loss for language image pre-training,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T13:40:39.709000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:40:39.709000Z digest=sha256:16b64a573160ca8f8ec5d40428b6c8b73c56afbea39e65bfad2da76973fc17f2

Observation 2c154ed6-f2a5-4984-8f95-5f5575dfe960 · outbound

This paper cites Zhang, H.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Zhang, H

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:39.868641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T13:40:39.713836Z digest=sha256:0c1f31bed5a3d9d93aad195d3a1fa68e0d49c1be75f6fedcb9f7d6cddc798b97

Observation 58d87be3-8112-40c9-9d50-d2d0c2364bd7 · outbound

This paper cites Pub- laynet: largest dataset ever for document layout analysis.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Pub- laynet: largest dataset ever for document layout analysis

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:39.850776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T13:40:39.718350Z digest=sha256:9588943f45c0a5b3baaa060b024d6fdc9e20ee4ace036e3b16eae0b05691ce0b

Pith citing papers

No inbound Pith citation observations are available.