Pith. sign in

REVIEW 3 cited by

Character-Aware Models Improve Visual Text Rendering

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2212.10562 v2 pith:NI7VAKCB submitted 2022-12-20 cs.CL cs.CV

classification cs.CLcs.CV
keywords modelsvisualcharacter-awaretextcharacter-blinddomaingainsgeneration
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Current image generation models struggle to reliably produce well-formed visual text. In this paper, we investigate a key contributing factor: popular text-to-image models lack character-level input features, making it much harder to predict a word's visual makeup as a series of glyphs. To quantify this effect, we conduct a series of experiments comparing character-aware vs. character-blind text encoders. In the text-only domain, we find that character-aware models provide large gains on a novel spelling task (WikiSpell). Applying our learnings to the visual domain, we train a suite of image generation models, and show that character-aware variants outperform their character-blind counterparts across a range of novel text rendering tasks (our DrawText benchmark). Our models set a much higher state-of-the-art on visual spelling, with 30+ point accuracy gains over competitors on rare words, despite training on far fewer examples.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation

    cs.CV 2025-06 conditional novelty 7.0 of 10

    Layer-normalized averaging of all decoder-only LLM hidden states, rather than last-layer embeddings, improves text-to-image compositional alignment and beats T5 on GenAI-Bench.

  2. IGD: Instructional Graphic Design with Multimodal Layer Generation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    IGD generates editable multi-layer graphic designs (posters, slides, stickers) from text instructions using an MLLM for layout and a diffusion model for image assets.

  3. Calligrapher: Freestyle Text Image Customization

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Calligrapher trains a style encoder and in-context inference on self-distilled FLUX outputs to redraw arbitrary text in the visual style of a reference image.

Pith tools