Pith. sign in

REVIEW 3 cited by

Kandinsky: an Improved Text-to-Image Synthesis with Image Prior and Latent Diffusion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.03502 v1 pith:6XXDE33V submitted 2023-10-05 cs.CV

classification cs.CV
keywords imagegenerationmodelmodelsdiffusionlatentpriortext-to-image
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Text-to-image generation is a significant domain in modern computer vision and has achieved substantial improvements through the evolution of generative architectures. Among these, there are diffusion-based models that have demonstrated essential quality enhancements. These models are generally split into two categories: pixel-level and latent-level approaches. We present Kandinsky1, a novel exploration of latent diffusion architecture, combining the principles of the image prior models with latent diffusion techniques. The image prior model is trained separately to map text embeddings to image embeddings of CLIP. Another distinct feature of the proposed model is the modified MoVQ implementation, which serves as the image autoencoder component. Overall, the designed model contains 3.3B parameters. We also deployed a user-friendly demo system that supports diverse generative modes such as text-to-image generation, image fusion, text and image fusion, image variations generation, and text-guided inpainting/outpainting. Additionally, we released the source code and checkpoints for the Kandinsky models. Experimental evaluations demonstrate a FID score of 8.03 on the COCO-30K dataset, marking our model as the top open-source performer in terms of measurable image generation quality.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Controllable Generation of Diverse Dermatological Imagery for Fair and Efficient Malignancy Classification

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A hybrid generative framework expands scarce dermatology data over 400× and reports 90.9% malignancy classification accuracy with improved fairness on the DDI benchmark.

  2. StyleSentinel: Reliable Artistic Copyright Verification via Stylistic Fingerprints

    cs.CV 2025-08 conditional novelty 5.0 of 10

    StyleSentinel detects style mimicry by learning a hypersphere around an artist's style fingerprint in VGG feature space and checking whether suspect images fall inside it.

  3. FBI: Learning Dexterous In-hand Manipulation with Dynamic Visuotactile Shortcut Policy

    cs.RO 2025-08 reject novelty 3.0 of 10

    The abstract claims a new visuotactile robot manipulation policy (FBI) that outperforms baselines, but the manuscript body is an unrelated paper on text-to-image synthesis, so the claimed result is absent.

Pith tools