Pith. sign in

REVIEW 4 cited by

UltraPixel: Advancing Ultra-High-Resolution Image Synthesis to New Peaks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.02158 v2 pith:4KMDQKB4 submitted 2024-07-02 cs.CV

classification cs.CV
keywords high-resolutionimagestrainingultrapixelcomplexityefficiencygenerationimage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Ultra-high-resolution image generation poses great challenges, such as increased semantic planning complexity and detail synthesis difficulties, alongside substantial training resource demands. We present UltraPixel, a novel architecture utilizing cascade diffusion models to generate high-quality images at multiple resolutions (\textit{e.g.}, 1K to 6K) within a single model, while maintaining computational efficiency. UltraPixel leverages semantics-rich representations of lower-resolution images in the later denoising stage to guide the whole generation of highly detailed high-resolution images, significantly reducing complexity. Furthermore, we introduce implicit neural representations for continuous upsampling and scale-aware normalization layers adaptable to various resolutions. Notably, both low- and high-resolution processes are performed in the most compact space, sharing the majority of parameters with less than 3$\%$ additional parameters for high-resolution outputs, largely enhancing training and inference efficiency. Our model achieves fast training with reduced data requirements, producing photo-realistic high-resolution images and demonstrating state-of-the-art performance in extensive experiments.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LUVE : Latent-Cascaded Ultra-High-Resolution Video Generation with Dual Frequency Experts

    cs.CV 2026-02 conditional novelty 6.0 of 10

    A latent-cascaded video generation framework with dual frequency-split experts reports state-of-the-art 2K/4K video generation on VBench, FIDpatch, and human preference.

  2. CineScale: Free Lunch in High-Resolution Cinematic Visual Generation

    cs.CV 2025-08 conditional novelty 6.0 of 10

    CineScale extends pre-trained diffusion models to 8k image and 4k video generation with mostly tuning-free inference plus a small LoRA adaptation for video.

  3. Progressive Growing of Video Tokenizers for Temporally Compact Latent Spaces

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Progressive growing of a video tokenizer from 4x to 8x and 16x temporal compression yields higher reconstruction quality and efficient diffusion training than direct training.

  4. FreeScale: Unleashing the Resolution of Diffusion Models via Tuning-Free Scale Fusion

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A tuning-free scale-fusion method that lets frozen diffusion models generate 8k images and high-res videos by combining global and local attention through frequency filtering.

Pith tools