Pith. sign in

REVIEW 2 cited by

PICK: Processing Key Information Extraction from Documents using Improved Graph Learning-Convolutional Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2004.07464 v3 pith:5QCUYYJF submitted 2020-04-16 cs.CV

classification cs.CV
keywords documentsfeaturesgraphtextualvisualbeenextractioninformation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Computer vision with state-of-the-art deep learning models has achieved huge success in the field of Optical Character Recognition (OCR) including text detection and recognition tasks recently. However, Key Information Extraction (KIE) from documents as the downstream task of OCR, having a large number of use scenarios in real-world, remains a challenge because documents not only have textual features extracting from OCR systems but also have semantic visual features that are not fully exploited and play a critical role in KIE. Too little work has been devoted to efficiently make full use of both textual and visual features of the documents. In this paper, we introduce PICK, a framework that is effective and robust in handling complex documents layout for KIE by combining graph learning with graph convolution operation, yielding a richer semantic representation containing the textual and visual features and global layout without ambiguity. Extensive experiments on real-world datasets have been conducted to show that our method outperforms baselines methods by significant margins. Our code is available at https://github.com/wenwenyu/PICK-pytorch.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. WordVIS: A Color Worth A Thousand Words

    cs.CV 2024-12 conditional novelty 5.0 of 10

    Recoloring words with a hand-crafted letter-to-color scheme improves image-only document classifiers on Tobacco-3482 by 3-5%, reaching a reported 91.14%.

  2. Patchfinder: Leveraging Visual Language Models for Accurate Information Retrieval using Model Uncertainty

    cs.CV 2024-12 conditional novelty 5.0 of 10

    PatchFinder uses VLM token confidence to select patch size and the most confident patch, achieving 94% field-extraction accuracy on 190 noisy scanned well documents.

Pith tools