Pith. sign in

REVIEW 5 cited by

Visual Text Processing: A Comprehensive Review and Unified Evaluation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.21682 v2 pith:VLX4NSKR submitted 2025-04-30 cs.CV

classification cs.CV
keywords textvisualprocessingmodelsevaluationadvancementscomprehensiveeffectively
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Visual text is a crucial component in both document and scene images, conveying rich semantic information and attracting significant attention in the computer vision community. Beyond traditional tasks such as text detection and recognition, visual text processing has witnessed rapid advancements driven by the emergence of foundation models, including text image reconstruction and text image manipulation. Despite significant progress, challenges remain due to the unique properties that differentiate text from general objects. Effectively capturing and leveraging these distinct textual characteristics is essential for developing robust visual text processing models. In this survey, we present a comprehensive, multi-perspective analysis of recent advancements in visual text processing, focusing on two key questions: (1) What textual features are most suitable for different visual text processing tasks? (2) How can these distinctive text features be effectively incorporated into processing frameworks? Furthermore, we introduce VTPBench, a new benchmark that encompasses a broad range of visual text processing datasets. Leveraging the advanced visual quality assessment capabilities of multimodal large language models (MLLMs), we propose VTPScore, a novel evaluation metric designed to ensure fair and reliable evaluation. Our empirical study with more than 20 specific models reveals substantial room for improvement in the current techniques. Our aim is to establish this work as a fundamental resource that fosters future exploration and innovation in the dynamic field of visual text processing. The relevant repository is available at https://github.com/shuyansy/Visual-Text-Processing-survey.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Vid-SME: Membership Inference Attacks against Large Video Understanding Models

    cs.CV 2025-05 reject novelty 7.0 of 10

    Vid-SME computes Sharma-Mittal entropy differences between natural and reversed video frame sequences to infer training membership in video understanding LLMs, but its effectiveness is confounded by member/non-member ...

  2. SlerpFlow: Spherical Trajectory Correction for Rectified Flow Inversion

    cs.CV 2026-07 reject novelty 6.0 of 10

    SlerpFlow replaces Euclidean solver steps with spherical-linear-interpolation (slerp) direction correction for rectified-flow inversion, reporting improved FLUX reconstruction and editing on PIE-Bench.

  3. EpiAgent: An Agent-Centric System for Ancient Inscription Restoration

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    EpiAgent is a new agent-centric system that restores degraded ancient inscriptions with better quality and generalization than prior rigid AI methods by using an LLM planner to coordinate multimodal tools and iterativ...

  4. Uni-DocDiff: A Unified Document Restoration Model Based on Diffusion

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A dual-stream diffusion model with a handcrafted prior pool and a prior fusion module unifies six document restoration tasks and matches task-specific specialists.

  5. Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A full-image, mask-guided CLIP model with two-stage multi-granularity alignment training sets a new state of the art for Chinese scene text retrieval and introduces a diverse-layout benchmark.

Pith tools