Pith. sign in

REVIEW 5 cited by

Group Diffusion Transformers are Unsupervised Multitask Learners

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.15027 v1 pith:TIXFN4PY submitted 2024-10-19 cs.CV

classification cs.CV
keywords generationgroupgdtsvisualdesigndiffusionimagetasks
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

While large language models (LLMs) have revolutionized natural language processing with their task-agnostic capabilities, visual generation tasks such as image translation, style transfer, and character customization still rely heavily on supervised, task-specific datasets. In this work, we introduce Group Diffusion Transformers (GDTs), a novel framework that unifies diverse visual generation tasks by redefining them as a group generation problem. In this approach, a set of related images is generated simultaneously, optionally conditioned on a subset of the group. GDTs build upon diffusion transformers with minimal architectural modifications by concatenating self-attention tokens across images. This allows the model to implicitly capture cross-image relationships (e.g., identities, styles, layouts, surroundings, and color schemes) through caption-based correlations. Our design enables scalable, unsupervised, and task-agnostic pretraining using extensive collections of image groups sourced from multimodal internet articles, image galleries, and video frames. We evaluate GDTs on a comprehensive benchmark featuring over 200 instructions across 30 distinct visual generation tasks, including picture book creation, font design, style transfer, sketching, colorization, drawing sequence generation, and character customization. Our models achieve competitive zero-shot performance without any additional fine-tuning or gradient updates. Furthermore, ablation studies confirm the effectiveness of key components such as data scaling, group size, and model design. These results demonstrate the potential of GDTs as scalable, general-purpose visual generation systems.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. IDEA-Bench: How Far are Generative Models from Professional Designing?

    cs.CV 2024-12 conditional novelty 7.0 of 10

    IDEA-Bench measures generative models on 100 professional design tasks and finds the best tested system scores only 22.48 out of 100.

  2. MUSAR: Exploring Multi-Subject Customization from Single-Subject Dataset via Attention Routing

    cs.CV 2025-05 conditional novelty 6.0 of 10

    MUSAR trains multi-subject text-to-image customization from a single-subject dataset by synthesizing diptych pairs and routing each image region's attention to the correct reference subject.

  3. ChatDiT: A Training-Free Baseline for Task-Agnostic Free-Form Chatting with Diffusion Transformers

    cs.CV 2024-12 conditional novelty 6.0 of 10

    An LLM-agent planner plus an unmodified FLUX.1-dev diffusion transformer achieves the top aggregate score on IDEA-Bench for zero-shot, freely described visual generation tasks, with no fine-tuning.

  4. Grid: Omni Visual Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    GRID shows that fine-tuning an image diffusion model on videos arranged as grid images can generate coherent video and multi-view sequences with far less data and compute than specialized video models.

  5. Multimodal Benchmarking and Recommendation of Text-to-Image Generation Models

    cs.GR 2025-05 conditional novelty 4.0 of 10

    Adding structured metadata such as fabric, sleeve length, and neckline to prompts improves composite quality and ground-truth similarity for most text-to-image models, while slightly reducing prompt-image alignment.

Pith tools