Pith. sign in

REVIEW 1 cited by

Zero-shot Text-guided Infinite Image Synthesis with LLM guidance

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.12642 v2 pith:ZBV4A764 submitted 2024-07-17 cs.CV cs.AI

classification cs.CVcs.AI
keywords imagegloballocalcaptiontext-guidedcontextexpandhigh-resolution
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Text-guided image editing and generation methods have diverse real-world applications. However, text-guided infinite image synthesis faces several challenges. First, there is a lack of text-image paired datasets with high-resolution and contextual diversity. Second, expanding images based on text requires global coherence and rich local context understanding. Previous studies have mainly focused on limited categories, such as natural landscapes, and also required to train on high-resolution images with paired text. To address these challenges, we propose a novel approach utilizing Large Language Models (LLMs) for both global coherence and local context understanding, without any high-resolution text-image paired training dataset. We train the diffusion model to expand an image conditioned on global and local captions generated from the LLM and visual feature. At the inference stage, given an image and a global caption, we use the LLM to generate a next local caption to expand the input image. Then, we expand the image using the global caption, generated local caption and the visual feature to consider global consistency and spatial local context. In experiments, our model outperforms the baselines both quantitatively and qualitatively. Furthermore, our model demonstrates the capability of text-guided arbitrary-sized image generation in zero-shot manner with LLM guidance.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Test-time Prompt Refinement for Text-to-Image Models

    cs.LG 2025-07 conditional novelty 5.0 of 10

    A training-free closed loop, in which a multimodal LLM rewrites a text prompt after inspecting the generated image, improves overall text-to-image alignment but degrades some attribute and spatial categories.

Pith tools