Pith. sign in

REVIEW 5 cited by

KNN-Diffusion: Image Generation via Large-Scale Retrieval

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2204.02849 v2 pith:OQI4E6SG submitted 2022-04-06 cs.CV cs.AIcs.CLcs.GRcs.LG

classification cs.CVcs.AIcs.CLcs.GRcs.LG
keywords large-scaleresultsretrievaltext-to-imagedatadatasetsdiffusionefficient
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent text-to-image models have achieved impressive results. However, since they require large-scale datasets of text-image pairs, it is impractical to train them on new domains where data is scarce or not labeled. In this work, we propose using large-scale retrieval methods, in particular, efficient k-Nearest-Neighbors (kNN), which offers novel capabilities: (1) training a substantially small and efficient text-to-image diffusion model without any text, (2) generating out-of-distribution images by simply swapping the retrieval database at inference time, and (3) performing text-driven local semantic manipulations while preserving object identity. To demonstrate the robustness of our method, we apply our kNN approach on two state-of-the-art diffusion backbones, and show results on several different datasets. As evaluated by human studies and automatic metrics, our method achieves state-of-the-art results compared to existing approaches that train text-to-image generation models using images only (without paired text data)

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Align & Invert: Solving Inverse Problems with Diffusion and Flow-based Models via Representation Alignment

    cs.CV 2025-11 conditional novelty 6.0 of 10

    Adding DINOv2 representation alignment to diffusion/flow inverse-problem solvers, using corrupted measurements as proxies, improves LPIPS/FID and cuts required sampling steps.

  2. Benefit from Reference: Retrieval-Augmented Cross-modal Point Cloud Completion

    cs.CV 2025-07 reject novelty 6.0 of 10

    A retrieval-augmented cross-modal framework with structural shared encoding and gated reference priors is claimed to achieve state-of-the-art point cloud completion, although the evaluation may be tainted by same-obje...

  3. LayoutRAG: Retrieval-Augmented Model for Content-agnostic Conditional Layout Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Retrieval-augmented generation, where retrieved layout templates guide a flow-matching model, improves conditional layout generation on RICO and PubLayNet.

  4. IA-T2I: Internet-Augmented Text-to-Image Generation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    IA-T2I uses active retrieval, hierarchical image selection, and self-reflection to supply internet reference images to T2I models, improving generation accuracy on uncertain-knowledge prompts.

  5. Retrieval Augmented Comic Image Generation

    cs.CV 2025-06 reject novelty 5.0 of 10

    RaCig combines retrieval-based character assignment with regional IP-Adapter and ControlNet injection to generate comic panels with consistent characters and diverse gestures.

Pith tools