REVIEW 5 cited by
Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting
T0 review · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read A two-stage model that analyzes page layout first, then parses text, tables, and formulas in parallel, reports state-of-the-art accuracy and speed on the benchmarks it evaluates.
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Extended reading notes
Core claim
Abstract: 'Dolphin achieves state-of-the-art performance across diverse page-level and element-level settings, while ensuring superior efficiency through its lightweight architecture and parallel parsing mechanism.' Concretely, Table 1 reports edit distances of 0.0114 and 0.0131 on Fox-Page EN/ZH and 0.1028 on Dolphin-Page, at 0.1729 FPS, outperforming all listed baselines. If correct, a 322M-parameter two-stage model can parse document pages more accurately and faster than much larger autoregressive VLMs.
Load-bearing premise
The evaluation assumes train/test disjointness: Section 4.1 lists PubTabNet and PubTab1M as table training data and Section 4.2 evaluates on the same benchmark families without stating that official test splits were excluded. It also introduces Dolphin-Page and Dolphin-Block as self-constructed benchmarks without stating that their pages and paragraphs are disjoint from the 30M-sample training corpus. If these overlaps exist, the SOTA page-level and table results reduce to fitting the evaluation distribution.
Editorial analysis
A structured set of objections, weighed in public.
Assumptions & free parameters
assumptions (4)
- domain assumption Official test splits of PubTabNet, PubTab1M, and formula benchmarks are excluded from the 30M-sample training set.
- domain assumption Synthetic documents rendered from HTML, LaTeX, and Markdown transfer to real-world documents.
- domain assumption Dolphin-Page and Dolphin-Block are unbiased, well-annotated benchmarks.
- domain assumption A single 322M-parameter encoder-decoder with Swin and mBart can represent all needed document parsing skills.
Cite this review
Pith. "Pith review of Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting." pith.science (2026). https://pith.science/paper/EZLI4UWJ
@misc{pith2026250514059,
author = {Pith},
title = {Pith review of: Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting},
year = {2026},
howpublished = {\url{https://pith.science/paper/EZLI4UWJ}},
note = {Machine review of arXiv:2505.14059}
}
read the original abstract
Document image parsing is challenging due to its complexly intertwined elements such as text paragraphs, figures, formulas, and tables. Current approaches either assemble specialized expert models or directly generate page-level content autoregressively, facing integration overhead, efficiency bottlenecks, and layout structure degradation despite their decent performance. To address these limitations, we present \textit{Dolphin} (\textit{\textbf{Do}cument Image \textbf{P}arsing via \textbf{H}eterogeneous Anchor Prompt\textbf{in}g}), a novel multimodal document image parsing model following an analyze-then-parse paradigm. In the first stage, Dolphin generates a sequence of layout elements in reading order. These heterogeneous elements, serving as anchors and coupled with task-specific prompts, are fed back to Dolphin for parallel content parsing in the second stage. To train Dolphin, we construct a large-scale dataset of over 30 million samples, covering multi-granularity parsing tasks. Through comprehensive evaluations on both prevalent benchmarks and self-constructed ones, Dolphin achieves state-of-the-art performance across diverse page-level and element-level settings, while ensuring superior efficiency through its lightweight architecture and parallel parsing mechanism. The code and pre-trained models are publicly available at https://github.com/ByteDance/Dolphin
Figures
Figures from the paper (9 more)
Forward citations
Cited by 5 Pith papers
-
MORE: A Multilingual Document Parsing Benchmark and Evaluation
MORE provides a 149-language, structure-aware document parsing benchmark from real PDFs and reports baselines showing specialized OCR models still fail on tables and rare scripts.
-
HSD: Training-Free Acceleration for Document Parsing Vision-Language Models with Hierarchical Speculative Decoding
A training-free, two-stage speculative decoding scheme accelerates VLM document parsers by ~2.8x end-to-end (up to 7x) while keeping parsing accuracy essentially unchanged.
-
UniRec-0.1B: Unified Text and Formula Recognition with 0.1B Parameters
A 0.1B-parameter text/formula recognition model trained on a new 40M-sample dataset matches or beats much larger OCR models and runs 2-9× faster.
-
LPCAN: Lightweight Pyramid Cross-Attention Network for Rail Surface Defect Detection Using RGB-D Data
A lightweight RGB-D cross-attention network is proposed for rail defect detection, but the SOTA accuracy and generalization claims are internally inconsistent and the implementation is not public.
-
Knowledge-Embedded and Hypernetwork-Guided Few-Shot Substation Meter Defect Image Generation Method
Fine-tuning Stable Diffusion with DreamBooth-style knowledge and hypernetwork-guided crack control maps can synthesize substation meter defect images that boost a YOLOv8 defect detector's mAP when added to the training set.
Reference graph
Works this paper leans on
-
[1]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Haoli Bai, Zhiguang Liu, Xiaojun Meng, Shuang Liu, LUO Yifeng, Rongfu Zheng, Liangwei Wang, Lu Hou, Jiansheng Wei, Xin Jiang, et al. 2023. Wukong-Reader : Multi-modal pre-training for fine-grained visual document understanding. In Proceedings of the Annual Meeting Of The Association For Computational Linguistics
work page 2023
-
[4]
Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, et al. 2025. Qwen2.5-VL technical report. arXiv preprint arXiv:2502.13923
arXiv 2025
-
[5]
Nougat: Neural optical understanding for academic documents
Lukas Blecher, Guillem Cucurull, Thomas Scialom, and Robert Stojnic. Nougat: Neural optical understanding for academic documents. In Proceedings of the International Conference on Learning Representations
-
[6]
Mingxu Chai, Ziyu Shen, Chong Zhang, Yue Zhang, Xiao Wang, Shihan Dou, Jihua Kang, Jiazheng Zhang, and Qi Zhang. 2024. DocFusion : A unified framework for document parsing tasks. arXiv preprint arXiv:2412.12505
work page Pith review arXiv 2024
-
[7]
Song Chen, Xinyu Guo, Yadong Li, Tao Zhang, Mingan Lin, Dongdong Kuang, Youwei Zhang, Lingfeng Ming, Fengyu Zhang, Yuran Wang, et al. 2025. Ocean-OCR : Towards general OCR application via a vision-language model. arXiv preprint arXiv:2501.15558
arXiv 2025
-
[8]
Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, et al. 2024. InternVL : Scaling up vision foundation models and aligning for generic visual-linguistic tasks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 24185--24198
work page 2024
Show all 57 references
-
[9]
LaTe x rainbow: Open source document layout semantic annotation framework
Changxu Duan and Sabine Bartsch. LaTe x rainbow: Open source document layout semantic annotation framework. In Proceedings of the Workshop for Natural Language Processing Open Source Software
-
[10]
Hao Feng, Qi Liu, Hao Liu, Jingqun Tang, Wengang Zhou, Houqiang Li, and Can Huang. 2024. DocPedia : Unleashing the power of large multimodal model in the frequency domain for versatile document understanding. Science China Information Sciences, 67(12):1--14
2024
-
[11]
Hao Feng, Zijian Wang, Jingqun Tang, Jinghui Lu, Wengang Zhou, Houqiang Li, and Can Huang. 2023. UniDoc : A universal large multimodal model for simultaneous text detection, recognition, spotting and understanding. arXiv preprint arXiv:2308.11592
2023 arXiv
-
[12]
Jonathan Herzig, Pawel Krzysztof Nowak, Thomas Mueller, Francesco Piccinno, and Julian Eisenschlos. 2020. TaPas : Weakly supervised table parsing via pre-training. In Proceedings of the Annual Meeting of the Association for Computational Linguistics, pages 4320--4333
2020
-
[13]
Anwen Hu, Haiyang Xu, Liang Zhang, Jiabo Ye, Ming Yan, Ji Zhang, Qin Jin, Fei Huang, and Jingren Zhou. 2024. mPLUG-DocOwl2 : High-resolution compressing for OCR -free multi-page document understanding. arXiv preprint arXiv:2409.03420
2024 arXiv
-
[14]
Yupan Huang, Tengchao Lv, Lei Cui, Yutong Lu, and Furu Wei. 2022. LayoutLMv3 : Pre-training for document ai with unified text and image masking. In Proceedings of the ACM International Conference on Multimedia, pages 4083--4091
2022
-
[15]
Donghyun Kim, Teakgyu Hong, Moonbin Yim, Yoonsik Kim, and Geewook Kim. 2023. On web-based visual corpus construction for visual document understanding. In Proceedings of the International Conference on Document Analysis and Recognition, pages 297--313
2023
-
[16]
Geewook Kim, Teakgyu Hong, Moonbin Yim, JeongYeon Nam, Jinyoung Park, Jinyeong Yim, Wonseok Hwang, Sangdoo Yun, Dongyoon Han, and Seunghyun Park. 2022. OCR -free document understanding transformer. In Proceedings of the European Conference on Computer Vision, pages 498--517
2022
-
[17]
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. 2023. Efficient memory management for large language model serving with pagedattention. In Proceedings of the Symposium on Operating Systems Principles, ...
2023
-
[18]
Mike Lewis. 2019. Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. arXiv preprint arXiv:1910.13461
2019 arXiv
-
[19]
Zhang Li, Biao Yang, Qiang Liu, Zhiyin Ma, Shuo Zhang, Jingxu Yang, Yabo Sun, Yuliang Liu, and Xiang Bai. 2024. Monkey: Image resolution and text label are important things for large multi-modal models. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recog...
2024
-
[20]
Chenglong Liu, Haoran Wei, Jinyue Chen, Lingyu Kong, Zheng Ge, Zining Zhu, Liang Zhao, Jianjian Sun, Chunrui Han, and Xiangyu Zhang. 2024 a . Focus anywhere for fine-grained multi-page document understanding. arXiv preprint arXiv:2405.14295
2024 arXiv
-
[21]
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2024 b . Visual instruction tuning. In Proceedings of the Neural Information Processing Systems, volume 36
2024
-
[22]
Yuliang Liu, Biao Yang, Qiang Liu, Zhang Li, Zhiyin Ma, Shuo Zhang, and Xiang Bai. 2024 c . TextMonkey : An OCR -free large multimodal model for understanding document. arXiv preprint arXiv:2403.04473
2024 arXiv
-
[23]
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. 2021. Swin Transformer : Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE International Conference on Computer Vision, pages 10012--10022
2021
-
[24]
Tengchao Lv, Yupan Huang, Jingye Chen, Yuzhong Zhao, Yilin Jia, Lei Cui, Shuming Ma, Yaoyao Chang, Shaohan Huang, Wenhui Wang, et al. 2023. Kosmos-2.5: A multimodal literate model. arXiv preprint arXiv:2309.11419
2023 arXiv
-
[25]
John MacFarlane. 2013. Pandoc: a universal document converter. URL: http://pandoc. org, 8
2013
-
[26]
Ahmed Nassar, Andres Marafioti, Matteo Omenetti, Maksym Lysak, Nikolaos Livathinos, Christoph Auer, Lucas Morin, Rafael Teixeira de Lima, Yusik Kim, A Said Gurbuz, et al. 2025. SmolDocling : An ultra-compact vision-language model for end-to-end multi-modal document conversion....
2025 arXiv
-
[27]
Zhiliang Peng, Wenhui Wang, Li Dong, Yaru Hao, Shaohan Huang, Shuming Ma, and Furu Wei. 2023. Kosmos-2: Grounding multimodal large language models to the world. arXiv preprint arXiv:2306.14824
2023 arXiv
-
[28]
Jake Poznanski, Jon Borchardt, Jason Dunkelberger, Regan Huff, Daniel Lin, Aman Rangapur, Christopher Wilhelm, Kyle Lo, and Luca Soldaini. 2025. olmOCR : Unlocking trillions of tokens in PDFs with vision language models. arXiv preprint arXiv:2502.18443
2025
-
[29]
Brandon Smock, Rohith Pesala, and Robin Abraham. 2022. PubTables-1M : Towards comprehensive table extraction from unstructured documents. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 4634--4642
2022
-
[30]
Jingqun Tang, Chunhui Lin, Zhen Zhao, Shu Wei, Binghong Wu, Qi Liu, Hao Feng, Yang Li, Siqi Wang, Lei Liao, et al. 2024 a . TextSquare : Scaling up text-centric visual instruction tuning. arXiv preprint arXiv:2404.12803
2024 arXiv
-
[31]
Jingqun Tang, Qi Liu, Yongjie Ye, Jinghui Lu, Shu Wei, Chunhui Lin, Wanqing Li, Mohamad Fitri Faiz Bin Mahmood, Hao Feng, Zhen Zhao, et al. 2024 b . MTVQA : Benchmarking multilingual text-centric visual question answering. arXiv preprint arXiv:2405.11985
2024 arXiv
-
[32]
Jingqun Tang, Su Qiao, Benlei Cui, Yuhang Ma, Sheng Zhang, and Dimitrios Kanoulas. 2022 a . You can even annotate text with voice: Transcription-only-supervised text spotting. In Proceedings of the ACM International Conference on Multimedia, pages 4154--4163
2022
-
[33]
Jingqun Tang, Wenqing Zhang, Hongye Liu, MingKun Yang, Bo Jiang, Guanglong Hu, and Xiang Bai. 2022 b . Few could be better than all: Feature sampling and grouping for scene text detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages ...
2022
-
[34]
Zineng Tang, Ziyi Yang, Guoxin Wang, Yuwei Fang, Yang Liu, Chenguang Zhu, Michael Zeng, Cha Zhang, and Mohit Bansal. 2023. Unifying vision, text, and layout for universal document processing. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pag...
2023
-
[35]
Gemini Team, Petko Georgiev, Ving Ian Lei, Ryan Burnell, Libin Bai, Anmol Gulati, Garrett Tanzer, Damien Vincent, Zhufeng Pan, Shibo Wang, et al. 2024. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv preprint arXiv:2403.05530
2024 arXiv
-
[36]
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Proceedings of the Neural Information Processing Systems, pages 5998--6008
2017
-
[37]
Bin Wang, Zhuangcheng Gu, Guang Liang, Chao Xu, Bo Zhang, Botian Shi, and Conghui He. 2024 a . UnimerNet : A universal network for real-world mathematical expression recognition. arXiv preprint arXiv:2404.15254
2024 arXiv
-
[38]
Bin Wang, Chao Xu, Xiaomeng Zhao, Linke Ouyang, Fan Wu, Zhiyuan Zhao, Rui Xu, Kaiwen Liu, Yuan Qu, Fukai Shang, et al. 2024 b . MinerU : An open-source solution for precise document content extraction. arXiv preprint arXiv:2409.18839
2024 arXiv
-
[39]
Dongsheng Wang, Natraj Raman, Mathieu Sibue, Zhiqiang Ma, Petr Babkin, Simerjot Kaur, Yulong Pei, Armineh Nourbakhsh, and Xiaomo Liu. 2024 c . DocLLM : A layout-aware generative language model for multimodal document understanding. In Proceedings of the Annual Meeting Of The A...
2024
-
[40]
Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, et al. 2024 d . Qwen2-VL : Enhancing vision-language model's perception of the world at any resolution. arXiv preprint arXiv:2409.12191
2024 arXiv
-
[41]
Peng Wang, Zhaohai Li, Jun Tang, Humen Zhong, Fei Huang, Zhibo Yang, and Cong Yao. 2024 e . PlatyPus : A generalized specialist model for reading text in various forms. In Proceedings of the European Conference on Computer Vision, pages 165--183
2024
-
[42]
Yonghui Wang, Wengang Zhou, Hao Feng, Keyi Zhou, and Houqiang Li. 2023. Towards improving document understanding: An exploration on text-grounding via mllms. arXiv preprint arXiv:2311.13194
2023 arXiv
-
[43]
Haoran Wei, Lingyu Kong, Jinyue Chen, Liang Zhao, Zheng Ge, Jinrong Yang, Jianjian Sun, Chunrui Han, and Xiangyu Zhang. 2024 a . Vary : Scaling up the vision vocabulary for large vision-language model. In Proceedings of the European Conference on Computer Vision, pages 408--424
2024
-
[44]
Haoran Wei, Chenglong Liu, Jinyue Chen, Jia Wang, Lingyu Kong, Yanming Xu, Zheng Ge, Liang Zhao, Jianjian Sun, Yuang Peng, et al. 2024 b . General OCR theory: Towards OCR -2.0 via a unified end-to-end model. arXiv preprint arXiv:2409.01704
2024 arXiv
-
[45]
Zhiyu Wu, Xiaokang Chen, Zizheng Pan, Xingchao Liu, Wen Liu, Damai Dai, Huazuo Gao, Yiyang Ma, Chengyue Wu, Bingxuan Wang, et al. 2024. Deepseek-VL2 : Mixture-of-experts vision-language models for advanced multimodal understanding. arXiv preprint arXiv:2412.10302
2024 arXiv
-
[46]
Yang Xu, Yiheng Xu, Tengchao Lv, Lei Cui, Furu Wei, Guoxin Wang, Yijuan Lu, Dinei Florencio, Cha Zhang, Wanxiang Che, et al. 2020 a . LayoutLMv2 : Multi-modal pre-training for visually-rich document understanding. arXiv preprint arXiv:2012.14740
2020 arXiv
-
[47]
Yiheng Xu, Minghao Li, Lei Cui, Shaohan Huang, Furu Wei, and Ming Zhou. 2020 b . LayoutLM : Pre-training of text and layout for document image understanding. In Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, pages 1192--1200
2020
-
[48]
Zhengyuan Yang, Linjie Li, Kevin Lin, Jianfeng Wang, Chung-Ching Lin, Zicheng Liu, and Lijuan Wang. 2023. The dawn of LMMs : Preliminary explorations with GPT -4 V (ision). arXiv preprint arXiv:2309.17421
2023 arXiv
-
[49]
Yuan Yao, Tianyu Yu, Ao Zhang, Chongyi Wang, Junbo Cui, Hongji Zhu, Tianchi Cai, Haoyu Li, Weilin Zhao, Zhihui He, et al. 2024. MiniCPM-V : A GPT-4V level MLLM on your phone. arXiv preprint arXiv:2408.01800
2024 arXiv
-
[50]
Jiabo Ye, Anwen Hu, Haiyang Xu, Qinghao Ye, Ming Yan, Yuhao Dan, Chenlin Zhao, Guohai Xu, Chenliang Li, Junfeng Tian, et al. 2023 a . mPLUG-DocOwl : Modularized multimodal large language model for document understanding. arXiv preprint arXiv:2307.02499
2023 arXiv
-
[51]
Jiabo Ye, Anwen Hu, Haiyang Xu, Qinghao Ye, Ming Yan, Guohai Xu, Chenliang Li, Junfeng Tian, Qi Qian, Ji Zhang, et al. 2023 b . UReader : Universal OCR -free visually-situated language understanding with multimodal large language model. arXiv preprint arXiv:2310.05126
2023 arXiv
-
[52]
Ya-Qi Yu, Minghui Liao, Jihao Wu, Yongxin Liao, Xiaoyu Zheng, and Wei Zeng. 2024 a . TextHawk : Exploring efficient fine-grained perception of multimodal large language models. arXiv preprint arXiv:2404.09204
2024 arXiv
-
[53]
Ya-Qi Yu, Minghui Liao, Jiwen Zhang, and Jihao Wu. 2024 b . TextHawk2 : A large vision-language model excels in bilingual OCR and grounding with 16x fewer tokens. arXiv preprint arXiv:2410.05261
2024 arXiv
-
[54]
Jianshu Zhang, Jun Du, Shiliang Zhang, Dan Liu, Yulong Hu, Jinshui Hu, Si Wei, and Lirong Dai. 2017. Watch, attend and parse: An end-to-end neural network based approach to handwritten mathematical expression recognition. Pattern Recognition, 71:196--206
2017
-
[55]
Weichao Zhao, Hao Feng, Qi Liu, Jingqun Tang, Binghong Wu, Lei Liao, Shu Wei, Yongjie Ye, Hao Liu, Wengang Zhou, et al. 2024 a . TabPedia : Towards comprehensive visual table understanding with concept synergy. In Proceedings of the Neural Information Processing Systems, volum...
2024
-
[56]
Zhen Zhao, Jingqun Tang, Chunhui Lin, Binghong Wu, Can Huang, Hao Liu, Xin Tan, Zhizhong Zhang, and Yuan Xie. 2024 b . Multi-modal in-context learning makes an ego-evolving scene text recognizer. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition,...
2024
-
[57]
Xu Zhong, Elaheh ShafieiBavani, and Antonio Jimeno Yepes. 2020. Image-based table recognition: data, model, and evaluation. In Proceedings of the European Conference on Computer Vision, pages 564--580
2020
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.