Pith. sign in

REVIEW 13 cited by

AnyText: Multilingual Visual Text Generation And Editing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.03054 v5 pith:TR27DDZC submitted 2023-11-06 cs.CV

classification cs.CV
keywords textanytextgenerationdiffusioneditingimagemultilingualvisual
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Diffusion model based Text-to-Image has achieved impressive achievements recently. Although current technology for synthesizing images is highly advanced and capable of generating images with high fidelity, it is still possible to give the show away when focusing on the text area in the generated image. To address this issue, we introduce AnyText, a diffusion-based multilingual visual text generation and editing model, that focuses on rendering accurate and coherent text in the image. AnyText comprises a diffusion pipeline with two primary elements: an auxiliary latent module and a text embedding module. The former uses inputs like text glyph, position, and masked image to generate latent features for text generation or editing. The latter employs an OCR model for encoding stroke data as embeddings, which blend with image caption embeddings from the tokenizer to generate texts that seamlessly integrate with the background. We employed text-control diffusion loss and text perceptual loss for training to further enhance writing accuracy. AnyText can write characters in multiple languages, to the best of our knowledge, this is the first work to address multilingual visual text generation. It is worth mentioning that AnyText can be plugged into existing diffusion models from the community for rendering or editing text accurately. After conducting extensive evaluation experiments, our method has outperformed all other approaches by a significant margin. Additionally, we contribute the first large-scale multilingual text images dataset, AnyWord-3M, containing 3 million image-text pairs with OCR annotations in multiple languages. Based on AnyWord-3M dataset, we propose AnyText-benchmark for the evaluation of visual text generation accuracy and quality. Our project will be open-sourced on https://github.com/tyxsspa/AnyText to improve and promote the development of text generation technology.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. WeEdit: A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing

    cs.CV 2026-03 conditional novelty 7.0 of 10

    WeEdit trains a glyph-guided, RL-optimized image editor on a 330K-pair synthetic multilingual dataset and reports open-source SOTA on its own bilingual and multilingual text-editing benchmarks.

  2. FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark

    cs.CV 2025-09 conditional novelty 7.0 of 10

    The authors build a 6M-image, 20M-caption reasoning dataset with generation chain-of-thought and a 7-track VLM-judged benchmark, then rank 19 text-to-image models.

  3. PosterHarness: Turning Scientific Poster Generation into an Auditable Instruction-Following Benchmark

    cs.CV 2026-07 conditional novelty 6.5 of 10

    A placeholder-first harness separates visual poster design from scientific figure grounding, turning poster generation into measurable instruction-following with a 12-paper pilot and failure taxonomy.

  4. InnoText: A Unified Model for Visual Text Generation and Editing

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A unified DiT model with font-size-aware modulation and region-weighted loss outperforms existing visual text generation and editing systems on bilingual benchmarks.

  5. AI-generated Images Challenge Visual Trust in High-risk Scenarios

    cs.CV 2026-07 conditional novelty 6.0 of 10

    On SafeIMG, a new safety-focused benchmark of 1,131 GPT Image 2 images, the best VLM detects 49.5% of generated images and the best specialized detector 33.1%, versus 81.7% for humans.

  6. IGD: Instructional Graphic Design with Multimodal Layer Generation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    IGD generates editable multi-layer graphic designs (posters, slides, stickers) from text instructions using an MLLM for layout and a diffusion model for image assets.

  7. ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A 91K GPT-4o-generated image and editing dataset, and a fine-tuned open model Janus-4o, report improved text-to-image scores and new editing ability.

  8. FontAdapter: Instant Font Adaptation in Visual Text Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A two-stage curriculum with synthetic paired font data enables instant adaptation of unseen fonts in text-to-image generation using one reference glyph, without test-time fine-tuning.

  9. OrienText: Surface Oriented Textual Image Generation

    cs.CV 2025-05 reject novelty 6.0 of 10

    OrienText uses surface-normal maps and a control branch on a diffusion model to render text that follows the angle of the surface it is placed on.

  10. Syn3DTxt: Embedding 3D Cues for Scene Text Generation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Adding RGB surface-normal masks to synthetic scene-text data improves perspective handling for a MOSTEL-based text editor, though several reported gains are inconsistent with the tables.

  11. Exploring In-Image Machine Translation with Real-World Background

    cs.CL 2025-05 conditional novelty 6.0 of 10

    DebackX translates text inside images by separating text from the background, translating the text-image directly, and fusing it back, outperforming prior IIMT models on a new real-background dataset.

  12. TextPixs: Glyph-Conditioned Diffusion with Character-Aware Attention and OCR-Guided Supervision

    cs.CV 2025-07 reject novelty 5.0 of 10

    The GCDA framework claims state-of-the-art text rendering in diffusion images via dual-stream encoding, attention segregation, and OCR supervision, but the paper lacks verifiable artifacts and contains internal incons...

  13. PosterCraft: Rethinking High-Quality Aesthetic Poster Generation in a Unified Framework

    cs.CV 2025-06 conditional novelty 5.0 of 10

    PosterCraft improves text-to-poster generation by cascading four stages of training (text rendering, region-weighted fine-tuning, preference optimization, and vision-language feedback), outperforming open-source basel...

Pith tools