Pith. sign in

REVIEW 9 cited by

Ocean-OCR: Towards General OCR Application via a Vision-Language Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.15558 v1 pith:IYEUQ4ID submitted 2025-01-26 cs.CV

classification cs.CV
keywords ocean-ocrscenariosunderstandingvariousabilityacrosscapabilitiesexcelling
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multimodal large language models (MLLMs) have shown impressive capabilities across various domains, excelling in processing and understanding information from multiple modalities. Despite the rapid progress made previously, insufficient OCR ability hinders MLLMs from excelling in text-related tasks. In this paper, we present \textbf{Ocean-OCR}, a 3B MLLM with state-of-the-art performance on various OCR scenarios and comparable understanding ability on general tasks. We employ Native Resolution ViT to enable variable resolution input and utilize a substantial collection of high-quality OCR datasets to enhance the model performance. We demonstrate the superiority of Ocean-OCR through comprehensive experiments on open-source OCR benchmarks and across various OCR scenarios. These scenarios encompass document understanding, scene text recognition, and handwritten recognition, highlighting the robust OCR capabilities of Ocean-OCR. Note that Ocean-OCR is the first MLLM to outperform professional OCR models such as TextIn and PaddleOCR.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Do VLMs Read or Rewrite? On Transcription Faithfulness in Vision-Language Models

    cs.AI 2026-05 conditional novelty 7.0 of 10

    General-purpose VLMs systematically rewrite perturbed words back to the original — up to 4.5 WER points on English — with rewriting tied to representation similarity and word length.

  2. HPD-Parsing: Hierarchical Parallel Document Parsing

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Hierarchical parallel decoding — a global layout branch plus concurrent content branches with multi-token prediction — reaches 4,752 tokens/sec (≈3× a vanilla autoregressive baseline) at competitive accuracy on OmniDocBench.

  3. Real5-OmniDocBench: A Full-Scale Physical Reconstruction Benchmark for Robust Document Parsing in the Wild

    cs.CV 2026-03 conditional novelty 6.0 of 10

    A full-scale physical reconstruction of OmniDocBench with five distortion scenarios shows all document-parsing models degrade in the real world, with the authors' PaddleOCR-VL-1.5 topping the leaderboard.

  4. HSD: Training-Free Acceleration for Document Parsing Vision-Language Models with Hierarchical Speculative Decoding

    cs.CV 2026-02 conditional novelty 6.0 of 10

    A training-free, two-stage speculative decoding scheme accelerates VLM document parsers by ~2.8x end-to-end (up to 7x) while keeping parsing accuracy essentially unchanged.

  5. Toward Reliable VLM: A Fine-Grained Benchmark and Framework for Exposure, Bias, and Inference in Korean Street Views

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A fine-grained Korean street-view benchmark shows that adding captions with place names to photos lets vision-language models pinpoint locations at high rates, highlighting privacy exposure.

  6. Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A two-stage model that analyzes page layout first, then parses text, tables, and formulas in parallel, reports state-of-the-art accuracy and speed on the benchmarks it evaluates.

  7. UniRec-0.1B: Unified Text and Formula Recognition with 0.1B Parameters

    cs.CV 2025-12 conditional novelty 5.0 of 10

    A 0.1B-parameter text/formula recognition model trained on a new 40M-sample dataset matches or beats much larger OCR models and runs 2-9× faster.

  8. Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A new resolution-focused benchmark and an open-source native-resolution training framework show that preserving original image resolution improves VLM performance on fine-grained visual tasks.

  9. Multi-Agent Interactive Question Generation Framework for Long Document Understanding

    cs.CL 2025-07 conditional novelty 4.0 of 10

    A multi-agent question generation pipeline produces long-context English and Arabic QA pairs (AraEngLongBench), and top LVLMs score below 50% on the resulting benchmark.

Pith tools