{"work":{"id":"8625279b-e85e-4c87-ba22-16d4d9e1f709","openalex_id":null,"doi":null,"arxiv_id":"2504.05741","raw_key":null,"title":"DDT: Decoupled Diffusion Transformer","authors":null,"authors_text":"S","year":2025,"venue":"cs.CV","abstract":"Diffusion transformers have demonstrated remarkable generation quality, albeit requiring longer training iterations and numerous inference steps. In each denoising step, diffusion transformers encode the noisy inputs to extract the lower-frequency semantic component and then decode the higher frequency with identical modules. This scheme creates an inherent optimization dilemma: encoding low-frequency semantics necessitates reducing high-frequency components, creating tension between semantic encoding and high-frequency decoding. To resolve this challenge, we propose a new \\textbf{\\color{ddt}D}ecoupled \\textbf{\\color{ddt}D}iffusion \\textbf{\\color{ddt}T}ransformer~(\\textbf{\\color{ddt}DDT}), with a decoupled design of a dedicated condition encoder for semantic extraction alongside a specialized velocity decoder. Our experiments reveal that a more substantial encoder yields performance improvements as model size increases. For ImageNet $256\\times256$, Our DDT-XL/2 achieves a new state-of-the-art performance of {1.31 FID}~(nearly $4\\times$ faster training convergence compared to previous diffusion transformers). For ImageNet $512\\times512$, Our DDT-XL/2 achieves a new state-of-the-art FID of 1.28. Additionally, as a beneficial by-product, our decoupled architecture enhances inference speed by enabling the sharing self-condition between adjacent denoising steps. To minimize performance degradation, we propose a novel statistical dynamic programming approach to identify optimal sharing strategies.","external_url":"https://arxiv.org/abs/2504.05741","cited_by_count":null,"metadata_source":"pith","metadata_fetched_at":"2026-07-10T18:47:31.564501+00:00","pith_arxiv_id":"2504.05741","created_at":"2026-05-10T06:31:30.460812+00:00","updated_at":"2026-07-10T18:47:31.564501+00:00","title_quality_ok":false,"display_title":"Ddt: Decoupled diffusion transformer","render_title":"Ddt: Decoupled diffusion transformer"},"hub":{"state":{"work_id":"8625279b-e85e-4c87-ba22-16d4d9e1f709","tier":"hub","tier_reason":"10+ Pith inbound or 1,000+ external citations","pith_inbound_count":23,"external_cited_by_count":null,"distinct_field_count":3,"first_pith_cited_at":"2025-11-17T18:59:57+00:00","last_pith_cited_at":"2026-07-09T11:49:57+00:00","author_build_status":"not_needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"not_needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-23T14:39:44.100784+00:00","tier_text":"hub"},"tier":"hub","role_counts":[{"context_role":"background","n":4},{"context_role":"baseline","n":3},{"context_role":"method","n":1}],"polarity_counts":[{"context_polarity":"background","n":4},{"context_polarity":"baseline","n":3},{"context_polarity":"use_method","n":1}],"runs":{},"summary":{},"graph":{},"authors":[]}}