Pith. sign in

REVIEW 7 cited by

CogView3: Finer and Faster Text-to-Image Generation via Relay Diffusion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.05121 v1 pith:EPQWU72G submitted 2024-03-08 cs.CV

classification cs.CV
keywords text-to-imagediffusioncogview3inferenceonlyfirstgenerationmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advancements in text-to-image generative systems have been largely driven by diffusion models. However, single-stage text-to-image diffusion models still face challenges, in terms of computational efficiency and the refinement of image details. To tackle the issue, we propose CogView3, an innovative cascaded framework that enhances the performance of text-to-image diffusion. CogView3 is the first model implementing relay diffusion in the realm of text-to-image generation, executing the task by first creating low-resolution images and subsequently applying relay-based super-resolution. This methodology not only results in competitive text-to-image outputs but also greatly reduces both training and inference costs. Our experimental results demonstrate that CogView3 outperforms SDXL, the current state-of-the-art open-source text-to-image diffusion model, by 77.0\% in human evaluations, all while requiring only about 1/2 of the inference time. The distilled variant of CogView3 achieves comparable performance while only utilizing 1/10 of the inference time by SDXL.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

    cs.CV 2025-11 unverdicted novelty 6.0 of 10

    Z-Image is an efficient 6B-parameter foundation model for image generation that rivals larger commercial systems in photorealism and bilingual text rendering through a new single-stream diffusion transformer and strea...

  2. MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    MMIG-Bench is a unified benchmark of 4,850 prompts and 1,750 reference images with a three-level evaluation suite, including the VQA-based Aspect Matching Score that correlates with human ratings.

  3. SafeCFG: Controlling Harmful Features with Dynamic Safe Guidance for Safe Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    SafeCFG adapts classifier-free guidance with a learned feature controller so that clean prompts generate normally while harmful prompts are pushed away from unsafe content.

  4. Owl-1: Omni World Model for Consistent Long Video Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Owl-1 generates long, multi-scene videos by using a language model to maintain a latent state and predict text dynamics, then rendering each clip with a video diffusion model.

  5. Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects

    cs.CV 2024-11 conditional novelty 6.0 of 10

    Self-Cross Diffusion Guidance penalizes overlap between aggregated self-attention maps of one subject and cross-attention maps of another, reducing subject mixing in text-to-image diffusion models.

  6. Text-to-Image Synthesis: A Decade Survey

    cs.CV 2024-11 conditional novelty 1.0 of 10

    A decade-spanning survey categorizes over 440 text-to-image papers by architecture, research problem, dataset, and evaluation metric.

  7. From Noise to Nuance: Advances in Deep Generative Image Models

    cs.CV 2024-12 conditional

    A broad literature review of deep generative image models from GANs to diffusion and transformer architectures, with no new empirical results.

Pith tools