REVIEW 3 major objections 5 minor 2 cited by
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends
T0 review · 3 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read This survey of visually rich document question answering concludes that explicit 2D layout encoding, not image resolution, separates the top-scoring models.
desk verdict A genuinely useful survey map of VRD question answering, but the Section 6 ranking claim contradicts its own Table 3 and should be softened. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The object that carries the argument is the document representation itself, decomposed into three modalities: text tokens, bounding boxes (layout), and the page image. The survey's comparison is organised around how models fuse these modalities: absolute 2D positional embeddings, relative 2D attention biases, disentangled attention that separates a token's semantic meaning from its horizontal and vertical distance to other tokens, cross-attention between visual and textual tokens, and page-level compression tokens. The evaluation machinery is ANLS (Average Normalized Levenshtein Similarity), the metric used in Table 3 to rank single- and multi-page VQA systems. The taxonomy's pivot is the contrast between structured encoders that consume layout explicitly and vision-only LVLMs that see the whole page as an image; the authors argue that this distinction, not resolution, separates the top performers.
What would settle it
Run a matched experiment: take one backbone, train a vision-only variant that sees page images at high resolution and a layout-aware variant that consumes text tokens with 2D positional biases, on the same VQA datasets, and compare ANLS on DocVQA, DUDE, and MMLongBench-Doc; if the vision-only variant matches or beats the layout-aware one on multi-page questions, the survey's central claim that text and layout are essential would be falsified.
Extended reading notes
Core claim
The central claim is that, in current document VQA, how a model represents the spatial layout of a page matters more than how many pixels it can see. The authors read Table 3 as showing that models making extensive use of positional features—ERNIE-Layout's disentangled attention over sequential, horizontal, and vertical relative distances, and Arctic-TILT's blockwise attention with a role bias for text tokens—have the best results, and they infer from this that text and layout information are essential for answering questions, even in complex charts and figures. They argue that structured multimodal approaches combining text, layout, and vision are more efficient for multi-page understanding than vision-only LVLMs, which either compress each page so heavily that performance degrades or must depend on a retriever to select relevant pages. They therefore recommend that the community prioritise layout handling and explicit 2D position encoding, and that visual features be injected through cross-attention with text tokens as queries rather than through self-attention over concatenated visual and textual tokens.
Load-bearing premise
The load-bearing premise is that the ANLS scores in Table 3 can be compared across models even though each model was trained and evaluated by a different team under different protocols; the authors themselves flag this in Section 7, noting that it is challenging to draw definitive conclusions.
Editorial extensions
If this is right
- Architecture choices should favour explicit 2D position encoding, whether absolute embeddings, relative biases, or 2D rotary positions, over treating document pages as generic images.
- For multi-page QA, sparse-attention designs such as global-local or blockwise attention are the paper's recommended direction, ahead of page-by-page compression and retrieval-dependent pipelines.
- Pretraining on document parsing tasks that turn page screenshots into structured text (HTML, Markdown, CSV/JSON) should be a standard step, since it aligns text, layout, and vision and makes visual features partially redundant.
- Cross-attention visual injection with text tokens as queries is preferred over self-attention over concatenated token lists, because it keeps visual features separate while letting the LLM interrogate them.
- Because the paper's score table is not a controlled comparison, the rankings should be re-validated with uniform training protocols before architectural conclusions are treated as settled.
Reading between the lines
- If the correlation between positional-feature use and top scores is causal, a matched ablation should show a larger layout advantage on table-heavy and multi-page benchmarks than on plain-text documents; the survey's cross-paper table cannot demonstrate this directly.
- The paper's evidence implies that for text-dense pages the visual modality is largely redundant; a concrete extension would route text-heavy pages through layout-aware text encoders and reserve vision encoders for figures, cutting compute with little accuracy loss.
- The recommended text-guided cross-attention fusion could be made adaptive by switching the query modality (text vs. vision) based on detected document type, generalising the paper's binary suggestion into a testable mechanism.
- Because the survey restricts itself to transformers, its conclusion that layout encoding is essential is scoped; graph-based layout models would be the natural comparison to see whether the finding survives outside attention architectures.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript is a survey of question answering over visually rich documents. It organizes recent work into three encoding families: structured multi-modal encoders that combine text, layout bounding boxes, and visual features (Section 2); vision-only large vision-language models that treat pages as images (Section 3); and multi-page strategies based on retrieval, per-page compression/query tokens, or sparse/recurrent attention (Section 4). Section 5 compares self-attention and cross-attention mechanisms for injecting visual features into an LLM decoder. The survey includes comparative tables of models and datasets (Tables 1-4) and concludes in Section 6 that models making extensive use of positional features, such as ERNIE-Layout and Arctic-TILT, achieve the best results and that text and layout are essential, while Section 7 acknowledges that cross-paper evaluation conditions differ.
Significance. The paper is a useful structured map of a fast-moving area, with broad coverage of encoder designs, multi-page strategies, and dataset characteristics. Its principal value is taxonomic: it collects a large set of recent models into a clear two-step pipeline and provides an appendix of VQA datasets. To the authors' credit, the limitations of the cross-paper comparison are explicitly acknowledged in Section 7. However, the central comparative conclusion in Section 6 is not established by the evidence in Table 3, and as written the conclusion is internally inconsistent with the caveat in Section 7. The significance of the paper as a guide for architecture choice therefore depends on a revision that either presents the conclusion as a hypothesis or supplies a controlled comparison.
major comments (3)
- [Section 6, Table 3] Section 6 states that 'models that make extensive use of positional features—such as ERNIE-Layout and Arctic-TILT—have the best results,' but the DocVQA column of Table 3 shows GPT-4o at 92.8 and InternLMXComposer2-4KHD at 90.0, both above ERNIE-Layout's 88.4 and with InternLMXComposer2-4KHD essentially tied with Arctic-TILT's 90.2. GPT-4o is a vision-only commercial LVLM in the survey's own taxonomy, and InternLMXComposer2-4KHD is also a vision-only model; the claimed ranking is therefore contradicted by the table's own numbers.
- [Section 6 vs Section 7] The causal conclusion in Section 6 ('This indicates that text and layout information are essential') directly conflicts with the limitation stated in Section 7: because methods are 'evaluated in their original experimental setups, which differ in terms of model architecture, training protocols, and datasets,' the authors themselves say it is 'challenging to draw definitive conclusions.' Observational cross-paper ANLS values cannot establish that layout information is essential; that would require a controlled ablation in which the same base model, training data, and protocol are evaluated with and without layout features.
- [Table 3] Table 3 mixes incomparable conditions: rows marked * use retrievers (e.g., InternLMXComposer2-4KHD with PDF-Wukong, Pix2Struct with Naidu et al., QwenVL and Idefics2 with M3DocRAG), rows marked ² concatenate page representations rather than performing true multi-page reasoning, and several cells are empty. Rankings based on such heterogeneous scores are not robust. The top-3 bold marking should be disclosed per column and restricted to comparable settings, or the table should be relabeled as a compilation of reported scores without ranking claims.
minor comments (5)
- [References] References Huang et al. 2024a and 2024b are identical ('From detection to application...'), and Xu et al. 2024a and 2024b are identical (LLaVA-UHD); please merge or disambiguate them.
- [Table 3] Table 3 model names are inconsistent: 'mPLUGDoc', 'mPLUGDoc1.5', and 'ILMXC24KHD' should be spelled as in the main text (mPLUG-DocOwl, mPLUG-DocOwl1.5, InternLMXComposer2-4KHD).
- [Table 4] Table 4 header says '#Pages' per document but BoundingDocs reports 237k, which appears to be a total rather than a per-document average; please clarify the units in that column.
- [Section 2.1] The notation \hat{V} = V ∪ [BBOX] should use a set of special tokens, e.g., V ∪ {[BBOX]}, and clarify how the marker is tokenized and inserted into the sequence.
- [Table 3] The bold top-3 markers are not visible in the text and the basis for choosing top-3 across a heterogeneous score matrix should be stated explicitly in the caption.
Circularity Check
No significant circularity: the survey's claims are interpretive summaries of external cited results, not derivations from its own scaffolding.
full rationale
This is a survey paper, so its claims are literature-synthesis statements rather than results derived from equations or fitted parameters. The main load-bearing claim in Section 6, that layout-aware models such as ERNIE-Layout and Arctic-TILT perform best, is presented as a reading of Table 3, which reports ANLS scores taken from external papers. The survey does not fit any parameter, define a quantity in terms of its own conclusion, or invoke a uniqueness theorem; the Table 3 numbers are independent evidence, albeit gathered under heterogeneous setups. The authors' own Section 7 explicitly disclaims the comparability of these numbers, which weakens the strength of the Section 6 inference but does not make it circular. The claim is a contestable interpretation of external observations, not a reduction of the survey's conclusion to its own inputs. Self-citation is not used as load-bearing support, and no cited prior work by the authors is invoked to forbid alternatives or supply an unverified premise. The correct verdict is therefore no significant circularity, with a score of 0.
Assumptions & free parameters
assumptions (2)
- domain assumption The reported benchmark scores in the cited papers are accurate and directly comparable across models.
- domain assumption The selection of papers is representative of the field.
Cite this review
Pith. "Pith review of Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends." pith.science (2026). https://pith.science/paper/BNUODZGZ
@misc{pith2026250102235,
author = {Pith},
title = {Pith review of: Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends},
year = {2026},
howpublished = {\url{https://pith.science/paper/BNUODZGZ}},
note = {Machine review of arXiv:2501.02235}
}
read the original abstract
The field of visually-rich document understanding, which involves interacting with visually-rich documents (whether scanned or born-digital), is rapidly evolving and still lacks consensus on several key aspects of the processing pipeline. In this work, we provide a comprehensive overview of state-of-the-art approaches, emphasizing their strengths and limitations, pointing out the main challenges in the field, and proposing promising research directions.
Figures
Forward citations
Cited by 2 Pith papers
-
Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis
Scene-aware multi-agent document synthesis plus error-driven hard-example expansion improves compact Qwen3-VL models on constrained and open-category KIE, topping reported on-device baselines.
-
DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth
OCR tools can be ranked without ground-truth labels by measuring how much a multimodal LLM must correct each tool's output.
Reference graph
Works this paper leans on
-
[1]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Joshua Ainslie, Santiago Onta \ n \'o n, Chris Alberti, Philip Pham, Anirudh Ravula, Sumit Sanghai, Zhuyun Meng, and Lana Hou. 2020. https://arxiv.org/abs/2004.08483 Etc: Encoding long and structured data in transformers . arXiv preprint arXiv:2004.08483
arXiv 2020
-
[4]
Mirna Al-Shetairy, Hanan Hindy, Dina Khattab, and Mostafa M. Aref. 2024. https://arxiv.org/abs/2410.13883 Transformers utilization in chart understanding: A review of recent advances and future trends . Preprint, arXiv:2410.13883
arXiv 2024
-
[5]
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob Menick, Sebastian Borgeaud, Andrew Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski,...
arXiv 2022
- [6]
- [7]
-
[8]
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. 2023. https://arxiv.org/abs/2308.12966 Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond . Preprint, arXiv:2308.12966
arXiv 2023
Show all 110 references
-
[9]
Tsachi Blau, Sharon Fogel, Roi Ronen, Alona Golts, Roy Ganz, Elad Ben Avraham, Aviad Aberdam, Shahar Tsiper, and Ron Litman. 2024. https://arxiv.org/abs/2401.03411 Gram: Global reasoning for multi-page vqa . Preprint, arXiv:2401.03411
2024 arXiv
-
[10]
Lukas Blecher, Guillem Cucurull, Thomas Scialom, and Robert Stojnic. 2023. https://arxiv.org/abs/2308.13418 Nougat: Neural optical understanding for academic documents . Preprint, arXiv:2308.13418
2023 arXiv
-
[11]
Borchmann, Michal Pietruszka, Wojciech Ja'skowski, Dawid Jurkiewicz, Piotr Halama, Pawel J'oziak, Lukasz Garncarek, Pawel Liskowski, Karolina Szyndler, Andrzej Gretkowski, Julita Oltusek, Gabriela Nowakowska, Artur Zawlocki, Lukasz Duhr, Pawel Dyda, and Michal Turski. 2024. ht...
2024 arXiv
-
[12]
Haoyu Cao, Changcun Bao, Chaohu Liu, Huang Chen, Kun Yin, Hao Liu, Yinsong Liu, Deqiang Jiang, and Xing Sun. 2023. https://arxiv.org/abs/2309.01131 Attention where it matters: Rethinking visual document understanding with selective region concentration . Preprint, arXiv:2309.01131
2023 arXiv
-
[13]
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. 2020. https://arxiv.org/abs/2005.12872 End-to-end object detection with transformers . Preprint, arXiv:2005.12872
2020 arXiv
-
[14]
Junbum Cha, Wooyoung Kang, Jonghwan Mun, and Byungseok Roh. 2024. https://arxiv.org/abs/2312.06742 Honeybee: Locality-enhanced projector for multimodal llm . Preprint, arXiv:2312.06742
2024 arXiv
-
[15]
Liang Chen, Haozhe Zhao, Tianyu Liu, Shuai Bai, Junyang Lin, Chang Zhou, and Baobao Chang. 2024. https://arxiv.org/abs/2403.06764 An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models . Preprint, arXiv:2403.06764
2024 arXiv
-
[16]
Pu-Chin Chen, Henry Tsai, Srinadh Bhojanapalli, Hyung Won Chung, Yin-Wen Chang, and Chun-Sung Ferng. 2021. https://arxiv.org/abs/2104.08698 A simple and effective positional encoding for transformers . Preprint, arXiv:2104.08698
2021 arXiv
-
[17]
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. 2020. https://arxiv.org/abs/1909.11740 Uniter: Universal image-text representation learning . Preprint, arXiv:1909.11740
2020 arXiv
-
[18]
Yew Ken Chia, Liying Cheng, Hou Pong Chan, Chaoqun Liu, Maojia Song, Sharifah Mahani Aljunied, Soujanya Poria, and Lidong Bing. 2024. https://arxiv.org/abs/2411.06176 M-longdoc: A benchmark for multimodal super-long document understanding and a retrieval-aware tuning framework...
2024 arXiv
-
[19]
Jaemin Cho, Debanjan Mahata, Ozan Irsoy, Yujie He, and Mohit Bansal. 2024. https://arxiv.org/abs/2411.04952 M3docrag: Multi-modal retrieval is what you need for multi-page multi-document understanding . Preprint, arXiv:2411.04952
2024 arXiv
-
[20]
He-Sen Dai, Xiao-Hui Li, Fei Yin, Xudong Yan, Shuqi Mei, and Cheng-Lin Liu. 2024. https://api.semanticscholar.org/CorpusID:272694741 Graphmllm: A graph-based multi-level layout language-independent model for document understanding . In IEEE International Conference on Document...
2024
-
[21]
Brian Davis, Bryan Morse, Bryan Price, Chris Tensmeyer, Curtis Wigington, and Vlad Morariu. 2022. https://arxiv.org/abs/2203.16618 End-to-end document recognition and understanding with dessurt . Preprint, arXiv:2203.16618
2022 arXiv
-
[22]
Xiang Deng, Huan Sun, Alyssa Lees, You Wu, and Cong Yu. 2020. https://arxiv.org/abs/2006.14806 Turl: Table understanding through representation learning . Preprint, arXiv:2006.14806
2020 arXiv
-
[23]
Mohamed Dhouib, Ghassen Bettaieb, and Aymen Shabou. 2023. https://arxiv.org/abs/2304.12484 Docparser: End-to-end ocr-free information extraction from visually rich documents . Preprint, arXiv:2304.12484
2023 arXiv
-
[24]
Yihao Ding, Jean Lee, and Soyeon Caren Han. 2024. https://arxiv.org/abs/2408.01287 Deep learning based visually rich document content understanding: A survey . Preprint, arXiv:2408.01287
2024 arXiv
-
[25]
Qi Dong, Lei Kang, and Dimosthenis Karatzas. 2024 a . https://api.semanticscholar.org/CorpusID:272701570 Multi-page document vqa with recurrent memory transformer . In International Workshop on Document Analysis Systems
2024
-
[26]
Xiaoyi Dong, Pan Zhang, Yuhang Zang, Yuhang Cao, Bin Wang, Linke Ouyang, Songyang Zhang, Haodong Duan, Wenwei Zhang, Yining Li, Hang Yan, Yang Gao, Zhe Chen, Xinyue Zhang, Wei Li, Jingwen Li, Wenhai Wang, Kai Chen, Conghui He, Xingcheng Zhang, Jifeng Dai, Yu Qiao, Dahua Lin, a...
2024 arXiv
-
[27]
Manuel Faysse, Hugues Sibille, Tony Wu, Bilel Omrani, Gautier Viaud, Céline Hudelot, and Pierre Colombo. 2024. https://arxiv.org/abs/2407.01449 Colpali: Efficient document retrieval with vision language models . Preprint, arXiv:2407.01449
2024 arXiv
-
[28]
Hao Feng, Qi Liu, Hao Liu, Jingqun Tang, Wengang Zhou, Houqiang Li, and Can Huang. 2024. https://arxiv.org/abs/2311.11810 Docpedia: Unleashing the power of large multimodal model in the frequency domain for versatile document understanding . Preprint, arXiv:2311.11810
2024 arXiv
-
[29]
Hao Feng, Zijian Wang, Jingqun Tang, Jinghui Lu, Wengang Zhou, Houqiang Li, and Can Huang. 2023. https://arxiv.org/abs/2308.11592 Unidoc: A universal large multimodal model for simultaneous text detection, recognition, spotting and understanding . Preprint, arXiv:2308.11592
2023 arXiv
-
[30]
Masato Fujitake. 2024. https://arxiv.org/abs/2403.14252 Layoutllm: Large language model instruction tuning for visually rich document understanding . Preprint, arXiv:2403.14252
2024 arXiv
-
[31]
Lukasz Garncarek, Rafal Powalski, Tomasz Stanislawek, Bartosz Topolski, Piotr Halama, Michał Turski, and Filip Grali'nski. 2020. https://api.semanticscholar.org/CorpusID:235262539 Lambert: Layout-aware language modeling for information extraction . In IEEE International Confer...
2020
-
[32]
Simone Giovannini, Fabio Coppini, Andrea Gemelli, and Simone Marinai. 2025. https://arxiv.org/abs/2501.03403 Boundingdocs: a unified dataset for document question answering with spatial annotations . Preprint, arXiv:2501.03403
2025
-
[33]
Prashant Gupta, Daniel Borchmann, Alvaro Dossantos, and Umapada Pal. 2022. https://arxiv.org/abs/2207.06881 Recurrent memory transformer . arXiv preprint arXiv:2207.06881
2022 arXiv
-
[34]
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2021. https://arxiv.org/abs/2006.03654 Deberta: Decoding-enhanced bert with disentangled attention . Preprint, arXiv:2006.03654
2021 arXiv
-
[35]
Teakgyu Hong, Donghyun Kim, Mingi Ji, Wonseok Hwang, Daehyun Nam, and Sungrae Park. 2022. https://arxiv.org/abs/2108.04539 Bros: A pre-trained language model focusing on text and layout for better key information extraction from documents . Preprint, arXiv:2108.04539
2022 arXiv
-
[36]
Wenyi Hong, Weihan Wang, Qingsong Lv, Jiazheng Xu, Wenmeng Yu, Junhui Ji, Yan Wang, Zihan Wang, Yuxuan Zhang, Juanzi Li, Bin Xu, Yuxiao Dong, Ming Ding, and Jie Tang. 2024. https://arxiv.org/abs/2312.08914 Cogagent: A visual language model for gui agents . Preprint, arXiv:2312.08914
2024 arXiv
-
[37]
Anwen Hu, Haiyang Xu, Jiabo Ye, Ming Yan, Liang Zhang, Bo Zhang, Chen Li, Ji Zhang, Qin Jin, Fei Huang, and Jingren Zhou. 2024 a . https://arxiv.org/abs/2403.12895 mplug-docowl 1.5: Unified structure learning for ocr-free document understanding . Preprint, arXiv:2403.12895
2024 arXiv
-
[38]
Anwen Hu, Haiyang Xu, Liang Zhang, Jiabo Ye, Ming Yan, Ji Zhang, Qin Jin, Fei Huang, and Jingren Zhou. 2024 b . https://arxiv.org/abs/2409.03420 mplug-docowl2: High-resolution compressing for ocr-free multi-page document understanding . Preprint, arXiv:2409.03420
2024 arXiv
-
[39]
Pengfei Hu, Zhenrong Zhang, Jiefeng Ma, Shuhang Liu, Jun Du, and Jianshu Zhang. 2025. https://arxiv.org/abs/2409.11887 Docmamba: Efficient document pre-training with state space model . Preprint, arXiv:2409.11887
2025 arXiv
-
[41]
Jiani Huang, Haihua Chen, Fengchang Yu, and Wei Lu. 2024 b . https://doi.org/10.1145/3657285 From detection to application: Recent advances in understanding scientific tables and figures . ACM Comput. Surv., 56(10)
2024 doi
-
[42]
Yupan Huang, Tengchao Lv, Lei Cui, Yutong Lu, and Furu Wei. 2022. https://arxiv.org/abs/2204.08387 Layoutlmv3: Pre-training for document ai with unified text and image masking . Preprint, arXiv:2204.08387
2022 arXiv
-
[43]
Geewook Kim, Teakgyu Hong, Moonbin Yim, Jeongyeon Nam, Jinyoung Park, Jinyeong Yim, Wonseok Hwang, Sangdoo Yun, Dongyoon Han, and Seunghyun Park. 2022. https://arxiv.org/abs/2111.15664 Ocr-free document understanding transformer . Preprint, arXiv:2111.15664
2022 arXiv
-
[44]
Jordy Van Landeghem, Rubén Tito, Łukasz Borchmann, Michał Pietruszka, Paweł Józiak, Rafał Powalski, Dawid Jurkiewicz, Mickaël Coustaty, Bertrand Ackaert, Ernest Valveny, Matthew Blaschko, Sien Moens, and Tomasz Stanisławek. 2023. https://arxiv.org/abs/2305.08455 Document under...
2023 arXiv
-
[45]
Hugo Laurençon, Léo Tronchon, Matthieu Cord, and Victor Sanh. 2024. https://arxiv.org/abs/2405.02246 What matters when building vision-language models? Preprint, arXiv:2405.02246
2024 arXiv
-
[46]
Junlong Lee, Yiheng Xu, Yang Xiao, Huan Wang, Jinlong Zhao, Pengchuan Xie, Miao Xu, Baolin Shi, and Lei Xu. 2022. https://arxiv.org/abs/2203.08411 Formnet: Structural encoding beyond sequential modeling in form document information extraction . arXiv preprint arXiv:2203.08411
2022 arXiv
-
[47]
Kenton Lee, Mandar Joshi, Iulia Turc, Hexiang Hu, Fangyu Liu, Julian Eisenschlos, Urvashi Khandelwal, Peter Shaw, Ming-Wei Chang, and Kristina Toutanova. 2023. https://arxiv.org/abs/2210.03347 Pix2struct: Screenshot parsing as pretraining for visual language understanding . Pr...
2023 arXiv
-
[48]
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2021. https://arxiv.org/abs/2005.11401 Retrieval-augmented generation for knowledge-int...
2021 arXiv
-
[49]
Chenliang Li, Bin Bi, Ming Yan, Wei Wang, Songfang Huang, Fei Huang, and Luo Si. 2021 a . https://arxiv.org/abs/2105.11210 Structurallm: Structural pre-training for form understanding . Preprint, arXiv:2105.11210
2021 arXiv
-
[50]
Jia-Nan Li, Jian Guan, Wei Wu, Zhengtao Yu, and Rui Yan. 2024 a . https://arxiv.org/abs/2409.19700 2d-tpe: Two-dimensional positional encoding enhances table understanding for large language models . Preprint, arXiv:2409.19700
2024 arXiv
-
[51]
Junlong Li, Yiheng Xu, Tengchao Lv, Lei Cui, Cha Zhang, and Furu Wei. 2022. https://arxiv.org/abs/2203.02378 Dit: Self-supervised pre-training for document image transformer . Preprint, arXiv:2203.02378
2022 arXiv
-
[52]
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023 a . https://arxiv.org/abs/2301.12597 Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models . Preprint, arXiv:2301.12597
2023 arXiv
-
[53]
Morariu, Handong Zhao, Rajiv Jain, Varun Manjunatha, and Hongfu Liu
Peizhao Li, Jiuxiang Gu, Jason Kuen, Vlad I. Morariu, Handong Zhao, Rajiv Jain, Varun Manjunatha, and Hongfu Liu. 2021 b . https://arxiv.org/abs/2106.03331 Selfdoc: Self-supervised document representation learning . Preprint, arXiv:2106.03331
2021 arXiv
-
[54]
Peng Li, Xiaotang Zhao, Wei Fang, et al. 2024 b . https://arxiv.org/pdf/2407.02392 Tokenpacker: Efficient visual projector for multimodal llm . arXiv preprint arXiv:2407.02392
2024 arXiv
-
[55]
Li, Xiantao Cai, Bo Du, and Hai Zhao
Qiwei Li, Z. Li, Xiantao Cai, Bo Du, and Hai Zhao. 2023 b . https://api.semanticscholar.org/CorpusID:260899841 Enhancing visually-rich document understanding via layout structure modeling . Proceedings of the 31st ACM International Conference on Multimedia
2023
-
[56]
Yanwei Li, Yuechen Zhang, Chengyao Wang, Zhisheng Zhong, Yixin Chen, Ruihang Chu, Shaoteng Liu, and Jiaya Jia. 2024 c . https://arxiv.org/abs/2403.18814 Mini-gemini: Mining the potential of multi-modality vision language models . Preprint, arXiv:2403.18814
2024 arXiv
-
[57]
Zhang Li, Biao Yang, Qiang Liu, Zhiyin Ma, Shuo Zhang, Jingxu Yang, Yabo Sun, Yuliang Liu, and Xiang Bai. 2024 d . https://arxiv.org/abs/2311.06607 Monkey: Image resolution and text label are important things for large multi-modal models . Preprint, arXiv:2311.06607
2024 arXiv
-
[58]
Manmatha, and Vijay Mahadevan
Haofu Liao, Aruni RoyChowdhury, Weijian Li, Ankan Bansal, Yuting Zhang, Zhuowen Tu, Ravi Kumar Satzoda, R. Manmatha, and Vijay Mahadevan. 2023. https://arxiv.org/abs/2307.07929 Doctr: Document transformer for structured information extraction in documents . Preprint, arXiv:2307.07929
2023 arXiv
-
[59]
Wenhui Liao, Jiapeng Wang, Hongliang Li, Chengyu Wang, Jun Huang, and Lianwen Jin. 2024. https://arxiv.org/abs/2408.15045 Doclayllm: An efficient and effective multi-modal extension of large language models for text-rich document understanding . Preprint, arXiv:2408.15045
2024 arXiv
-
[60]
Ziyi Lin, Chris Liu, Renrui Zhang, Peng Gao, Longtian Qiu, Han Xiao, Han Qiu, Chen Lin, Wenqi Shao, Keqin Chen, Jiaming Han, Siyuan Huang, Yichi Zhang, Xuming He, Hongsheng Li, and Yu Qiao. 2023. https://arxiv.org/abs/2311.07575 Sphinx: The joint mixing of weights, tasks, and ...
2023 arXiv
-
[61]
Chaohu Liu, Kun Yin, Haoyu Cao, Xinghua Jiang, Xin Li, Yinsong Liu, Deqiang Jiang, Xing Sun, and Linli Xu. 2024 a . https://arxiv.org/abs/2404.06918 Hrvda: High-resolution visual document assistant . Preprint, arXiv:2404.06918
2024 arXiv
-
[62]
Fangyu Liu, Julian Martin Eisenschlos, Francesco Piccinno, Syrine Krichene, Chenxi Pang, Kenton Lee, Mandar Joshi, Wenhu Chen, Nigel Collier, and Yasemin Altun. 2023. https://arxiv.org/abs/2212.10505 Deplot: One-shot visual language reasoning by plot-to-table translation . Pre...
2023 arXiv
-
[63]
Hao Liu, Xinghua Jiang, Xin Li, Antai Guo, Deqiang Jiang, and Bo Ren. 2022 a . https://arxiv.org/abs/2204.08227 The devil is in the frequency: Geminated gestalt autoencoder for self-supervised visual pre-training . Preprint, arXiv:2204.08227
2022 arXiv
-
[64]
Yuliang Liu, Biao Yang, Qiang Liu, Zhang Li, Zhiyin Ma, Shuo Zhang, and Xiang Bai. 2024 b . https://arxiv.org/abs/2403.04473 Textmonkey: An ocr-free large multimodal model for understanding document . Preprint, arXiv:2403.04473
2024 arXiv
-
[65]
Ze Liu, Han Hu, Yutong Lin, Zhuliang Yao, Zhenda Xie, Yixuan Wei, Jia Ning, Yue Cao, Zheng Zhang, Li Dong, Furu Wei, and Baining Guo. 2022 b . https://arxiv.org/abs/2111.09883 Swin transformer v2: Scaling up capacity and resolution . Preprint, arXiv:2111.09883
2022 arXiv
-
[66]
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. 2021. https://arxiv.org/abs/2103.14030 Swin transformer: Hierarchical vision transformer using shifted windows . Preprint, arXiv:2103.14030
2021 arXiv
-
[67]
Junyu Lu, Dixiang Zhang, Songxin Zhang, Zejian Xie, Zhuoyang Song, Cong Lin, Jiaxing Zhang, Bingyi Jing, and Pingjian Zhang. 2024. https://arxiv.org/abs/2312.05278 Lyrics: Boosting fine-grained language-vision alignment and comprehension via semantic-aware visual objects . Pre...
2024 arXiv
-
[68]
Gen Luo, Yiyi Zhou, Yuxin Zhang, Xiawu Zheng, Xiaoshuai Sun, and Rongrong Ji. 2024. https://arxiv.org/abs/2403.03003 Feast your eyes: Mixture-of-resolution adaptation for multimodal large language models . Preprint, arXiv:2403.03003
2024 arXiv
-
[69]
Tengchao Lv, Yupan Huang, Jingye Chen, Yuzhong Zhao, Yilin Jia, Lei Cui, Shuming Ma, Yaoyao Chang, Shaohan Huang, Wenhui Wang, Li Dong, Weiyao Luo, Shaoxiang Wu, Guoxin Wang, Cha Zhang, and Furu Wei. 2024. https://arxiv.org/abs/2309.11419 Kosmos-2.5: A multimodal literate mode...
2024 arXiv
-
[70]
Feipeng Ma, Yizhou Zhou, Hebei Li, Zilong He, Siying Wu, Fengyun Rao, Yueyi Zhang, and Xiaoyan Sun. 2024 a . https://ar5iv.labs.arxiv.org/html/2408.11795v1 Ee-mllm: A data-efficient and compute-efficient multimodal large language model . arXiv preprint arXiv:2408.11795
2024 arXiv
-
[71]
Xueguang Ma, Sheng-Chieh Lin, Minghan Li, Wenhu Chen, and Jimmy Lin. 2024 b . https://arxiv.org/abs/2406.11251 Unifying multimodal retrieval via document screenshot embedding . Preprint, arXiv:2406.11251
2024 arXiv
-
[72]
Yubo Ma, Yuhang Zang, Liangyu Chen, Meiqi Chen, Yizhu Jiao, Xinze Li, Xinyuan Lu, Ziyu Liu, Yan Ma, Xiaoyi Dong, Pan Zhang, Liangming Pan, Yu-Gang Jiang, Jiaqi Wang, Yixin Cao, and Aixin Sun. 2024 c . https://arxiv.org/abs/2407.01523 MMLongBench-Doc: Benchmarking Long-context ...
2024 arXiv
-
[73]
Zhiming Mao, Haoli Bai, Lu Hou, Jiansheng Wei, Xin Jiang, Qun Liu, and Kam-Fai Wong. 2024. https://arxiv.org/abs/2403.16516 Visually guided generative text-layout pre-training for document intelligence . Preprint, arXiv:2403.16516
2024 arXiv
-
[74]
V Jawahar
Minesh Mathew, Viraj Bagal, Rubèn Pérez Tito, Dimosthenis Karatzas, Ernest Valveny, and C. V Jawahar. 2021 a . https://arxiv.org/abs/2104.12756 Infographicvqa . Preprint, arXiv:2104.12756
2021 arXiv
-
[75]
Minesh Mathew, Dimosthenis Karatzas, and C. V. Jawahar. 2021 b . https://arxiv.org/abs/2007.00398 Docvqa: A dataset for vqa on document images . Preprint, arXiv:2007.00398
2021 arXiv
-
[76]
Chaitanya Naidu, Mohammad Khan, and C. V. Jawahar. 2024. https://arxiv.org/abs/2404.19024 Multi-page document visual question answering using self-attention scoring mechanism . arXiv preprint arXiv:2404.19024
2024 arXiv
-
[77]
Qiming Peng, Yinxu Pan, Wenjin Wang, Bin Luo, Zhenyu Zhang, Zhengjie Huang, Teng Hu, Weichong Yin, Yongfeng Chen, Yin Zhang, Shikun Feng, Yu Sun, Hao Tian, Hua Wu, and Haifeng Wang. 2022. https://arxiv.org/abs/2210.06155 Ernie-layout: Layout knowledge enhanced pre-training for...
2022 arXiv
-
[78]
Rafał Powalski, Łukasz Borchmann, Dawid Jurkiewicz, Tomasz Dwojak, Michał Pietruszka, and Gabriela Pałka. 2021. https://arxiv.org/abs/2102.09550 Going full-tilt boogie on document understanding with text-image-layout transformer . Preprint, arXiv:2102.09550
2021 arXiv
-
[79]
Subhojeet Pramanik, Shashank Mujumdar, and Hima Patel. 2022. https://arxiv.org/abs/2009.14457 Towards a multi-modal, multi-task learning based pre-training framework for document representation learning . Preprint, arXiv:2009.14457
2022 arXiv
-
[80]
Smith, and Mike Lewis
Ofir Press, Noah A. Smith, and Mike Lewis. 2022. https://arxiv.org/abs/2108.12409 Train short, test long: Attention with linear biases enables input length extrapolation . Preprint, arXiv:2108.12409
2022 arXiv
-
[81]
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2023. https://arxiv.org/abs/1910.10683 Exploring the limits of transfer learning with a unified text-to-text transformer . Preprint, arXiv:1910.10683
2023 arXiv
-
[82]
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2016. https://arxiv.org/abs/1506.01497 Faster r-cnn: Towards real-time object detection with region proposal networks . Preprint, arXiv:1506.01497
2016 arXiv
-
[83]
Imanol Schlag, Paul Smolensky, Roland Fernandez, Nebojsa Jojic, Jürgen Schmidhuber, and Jianfeng Gao. 2020. https://arxiv.org/abs/1910.06611 Enhancing the transformer with explicit relational encoding for math problem solving . Preprint, arXiv:1910.06611
2020 arXiv
-
[84]
Yuzhang Shang, Mu Cai, Bingxin Xu, Yong Jae Lee, and Yan Yan. 2024. https://arxiv.org/abs/2403.15388 Llava-prumerge: Adaptive token reduction for efficient large multimodal models . Preprint, arXiv:2403.15388
2024
-
[85]
Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu. 2023. https://arxiv.org/abs/2104.09864 Roformer: Enhanced transformer with rotary position embedding . Preprint, arXiv:2104.09864
2023 arXiv
-
[86]
Ryota Tanaka, Taichi Iki, Kyosuke Nishida, Kuniko Saito, and Jun Suzuki. 2024. https://arxiv.org/abs/2401.13313 Instructdoc: A dataset for zero-shot generalization of visual document understanding with instructions . Preprint, arXiv:2401.13313
2024 arXiv
-
[87]
Ryota Tanaka, Kyosuke Nishida, Kosuke Nishida, Taku Hasegawa, Itsumi Saito, and Kuniko Saito. 2023. https://arxiv.org/abs/2301.04883 Slidevqa: A dataset for document visual question answering on multiple images . Preprint, arXiv:2301.04883
2023 arXiv
-
[88]
Ryota Tanaka, Kyosuke Nishida, and Sen Yoshida. 2021. https://arxiv.org/abs/2101.11272 Visualmrc: Machine reading comprehension on document images . Preprint, arXiv:2101.11272
2021 arXiv
-
[89]
Jingqun Tang, Chunhui Lin, Zhen Zhao, Shu Wei, Binghong Wu, Qi Liu, Hao Feng, Yang Li, Siqi Wang, Lei Liao, Wei Shi, Yuliang Liu, Hao Liu, Yuan Xie, Xiang Bai, and Can Huang. 2024. https://arxiv.org/abs/2404.12803 Textsquare: Scaling up text-centric visual instruction tuning ....
2024 arXiv
-
[90]
Zineng Tang, Ziyi Yang, Guoxin Wang, Yuwei Fang, Yang Liu, Chenguang Zhu, Michael Zeng, Cha Zhang, and Mohit Bansal. 2023. https://arxiv.org/abs/2212.02623 Unifying vision, text, and layout for universal document processing . Preprint, arXiv:2212.02623
2023 arXiv
-
[91]
Rubèn Tito, Dimosthenis Karatzas, and Ernest Valveny. 2023. https://arxiv.org/abs/2212.05935 Hierarchical multimodal transformers for multi-page docvqa . Preprint, arXiv:2212.05935
2023 arXiv
-
[92]
Zilong Wang, Jiuxiang Gu, Chris Tensmeyer, Nikolaos Barmpalios, Ani Nenkova, Tong Sun, Jingbo Shang, and Vlad I. Morariu. 2022. https://arxiv.org/abs/2211.14958 Mgdoc: Pre-training with multi-granular hierarchy for document image understanding . Preprint, arXiv:2211.14958
2022 arXiv
-
[93]
Haoran Wei, Lingyu Kong, Jinyue Chen, Liang Zhao, Zheng Ge, Jinrong Yang, Jianjian Sun, Chunrui Han, and Xiangyu Zhang. 2023. https://arxiv.org/abs/2312.06109 Vary: Scaling up the vision vocabulary for large vision-language models . Preprint, arXiv:2312.06109
2023 arXiv
- [94]
-
[96]
Ruyi Xu, Yuan Yao, Zonghao Guo, Junbo Cui, Zanlin Ni, Chunjiang Ge, Tat-Seng Chua, Zhiyuan Liu, Maosong Sun, and Gao Huang. 2024 b . https://arxiv.org/abs/2403.11703 Llava-uhd: an lmm perceiving any aspect ratio and high-resolution images . Preprint, arXiv:2403.11703
2024 arXiv
-
[97]
Yang Xu, Yiheng Xu, Tengchao Lv, Lei Cui, Furu Wei, Guoxin Wang, Yijuan Lu, Dinei Florencio, Cha Zhang, Wanxiang Che, Min Zhang, and Lidong Zhou. 2022. https://arxiv.org/abs/2012.14740 Layoutlmv2: Multi-modal pre-training for visually-rich document understanding . Preprint, ar...
2022 arXiv
-
[98]
Yiheng Xu, Minghao Li, Lei Cui, Shaohan Huang, Furu Wei, and Ming Zhou. 2019. https://api.semanticscholar.org/CorpusID:209515395 Layoutlm: Pre-training of text and layout for document image understanding . Proceedings of the 26th ACM SIGKDD International Conference on Knowledg...
2019
-
[99]
Yiheng Xu, Tengchao Lv, Lei Cui, Guoxin Wang, Yijuan Lu, Dinei Florencio, Cha Zhang, and Furu Wei. 2021. https://arxiv.org/abs/2104.08836 Layoutxlm: Multimodal pre-training for multilingual visually-rich document understanding . Preprint, arXiv:2104.08836
2021 arXiv
-
[100]
Jiabo Ye, Anwen Hu, Haiyang Xu, Qinghao Ye, Ming Yan, Yuhao Dan, Chenlin Zhao, Guohai Xu, Chenliang Li, Junfeng Tian, Qian Qi, Ji Zhang, and Fei Huang. 2023 a . https://arxiv.org/abs/2307.02499 mplug-docowl: Modularized multimodal large language model for document understandin...
2023 arXiv
-
[101]
Jiabo Ye, Anwen Hu, Haiyang Xu, Qinghao Ye, Ming Yan, Guohai Xu, Chenliang Li, Junfeng Tian, Qi Qian, Ji Zhang, Qin Jin, Liang He, Xin Alex Lin, and Fei Huang. 2023 b . https://arxiv.org/abs/2310.05126 Ureader: Universal ocr-free visually-situated language understanding with m...
2023 arXiv
-
[102]
Pengcheng Yin, Graham Neubig, Wen tau Yih, and Sebastian Riedel. 2020. https://arxiv.org/abs/2005.08314 Tabert: Pretraining for joint understanding of textual and tabular data . Preprint, arXiv:2005.08314
2020 arXiv
-
[103]
Ya-Qi Yu, Minghui Liao, Jihao Wu, Yongxin Liao, Xiaoyu Zheng, and Wei Zeng. 2024. https://arxiv.org/abs/2404.09204 Texthawk: Exploring efficient fine-grained perception of multimodal large language models . Preprint, arXiv:2404.09204
2024 arXiv
-
[104]
Jiarui Zhang, Mahyar Khayatkhoei, Prateek Chhikara, and Filip Ilievski. 2025. https://openreview.net/forum?id=DgaY5mDdmT MLLM s know where to look: Training-free perception of small visual details with multimodal LLM s . In The Thirteenth International Conference on Learning R...
2025
-
[105]
Jiaxin Zhang, Wentao Yang, Songxuan Lai, Zecheng Xie, and Lianwen Jin. 2024 a . https://arxiv.org/abs/2406.19101 Dockylin: A large multimodal model for visual document understanding with efficient visual slimming . arXiv preprint arXiv:2406.19101
2024 arXiv
-
[106]
Liang Zhang, Anwen Hu, Haiyang Xu, Ming Yan, Yichen Xu, Qin Jin, Ji Zhang, and Fei Huang. 2024 b . https://arxiv.org/abs/2404.16635 Tinychart: Efficient chart understanding with visual token merging and program-of-thoughts learning . Preprint, arXiv:2404.16635
2024 arXiv
-
[107]
Qintong Zhang, Victor Shea-Jay Huang, Bin Wang, Junyuan Zhang, Zhengren Wang, Hao Liang, Shawn Wang, Matthieu Lin, Conghui He, and Wentao Zhang. 2024 c . https://arxiv.org/abs/2410.21169 Document parsing unveiled: Techniques, challenges, and prospects for structured informatio...
2024 arXiv
-
[108]
Yanzhe Zhang, Ruiyi Zhang, Jiuxiang Gu, Yufan Zhou, Nedim Lipka, Diyi Yang, and Tong Sun. 2024 d . https://arxiv.org/abs/2306.17107 Llavar: Enhanced visual instruction tuning for text-rich image understanding . Preprint, arXiv:2306.17107
2024 arXiv
-
[109]
Zhenrong Zhang, Jiefeng Ma, Jun Du, Licheng Wang, and Jianshu Zhang. 2022. https://api.semanticscholar.org/CorpusID:247748605 Multimodal pre-training based on graph attention network for document understanding . IEEE Transactions on Multimedia, 25:6743--6755
2022
-
[110]
Fengbin Zhu, Wenqiang Lei, Fuli Feng, Chao Wang, Haozhou Zhang, and Tat-Seng Chua. 2022. https://doi.org/10.1145/3503161.3548422 Towards complex document understanding by discrete reasoning . In Proceedings of the 30th ACM International Conference on Multimedia, page 4857–4866. ACM
2022
-
[111]
Fengbin Zhu, Ziyang Liu, Xiang Yao Ng, Haohui Wu, Wenjie Wang, Fuli Feng, Chao Wang, Huanbo Luan, and Tat Seng Chua. 2024. https://arxiv.org/abs/2410.21311 Mmdocbench: Benchmarking large vision-language models for fine-grained visual document understanding . Preprint, arXiv:2410.21311
2024 arXiv
-
[112]
Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. 2021. https://arxiv.org/abs/2010.04159 Deformable detr: Deformable transformers for end-to-end object detection . Preprint, arXiv:2010.04159
2021 arXiv
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.