REVIEW 5 cited by
KNN-Diffusion: Image Generation via Large-Scale Retrieval
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Recent text-to-image models have achieved impressive results. However, since they require large-scale datasets of text-image pairs, it is impractical to train them on new domains where data is scarce or not labeled. In this work, we propose using large-scale retrieval methods, in particular, efficient k-Nearest-Neighbors (kNN), which offers novel capabilities: (1) training a substantially small and efficient text-to-image diffusion model without any text, (2) generating out-of-distribution images by simply swapping the retrieval database at inference time, and (3) performing text-driven local semantic manipulations while preserving object identity. To demonstrate the robustness of our method, we apply our kNN approach on two state-of-the-art diffusion backbones, and show results on several different datasets. As evaluated by human studies and automatic metrics, our method achieves state-of-the-art results compared to existing approaches that train text-to-image generation models using images only (without paired text data)
Forward citations
Cited by 5 Pith papers
-
Align & Invert: Solving Inverse Problems with Diffusion and Flow-based Models via Representation Alignment
Adding DINOv2 representation alignment to diffusion/flow inverse-problem solvers, using corrupted measurements as proxies, improves LPIPS/FID and cuts required sampling steps.
-
Benefit from Reference: Retrieval-Augmented Cross-modal Point Cloud Completion
A retrieval-augmented cross-modal framework with structural shared encoding and gated reference priors is claimed to achieve state-of-the-art point cloud completion, although the evaluation may be tainted by same-obje...
-
LayoutRAG: Retrieval-Augmented Model for Content-agnostic Conditional Layout Generation
Retrieval-augmented generation, where retrieved layout templates guide a flow-matching model, improves conditional layout generation on RICO and PubLayNet.
-
IA-T2I: Internet-Augmented Text-to-Image Generation
IA-T2I uses active retrieval, hierarchical image selection, and self-reflection to supply internet reference images to T2I models, improving generation accuracy on uncertain-knowledge prompts.
-
Retrieval Augmented Comic Image Generation
RaCig combines retrieval-based character assignment with regional IP-Adapter and ControlNet injection to generate comic panels with consistent characters and diverse gestures.
Discussion (0). Continue with ORCID to comment.