REVIEW 4 major objections 5 minor 137 references
High-Resolution Image Synthesis via Next-Token Prediction
T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read An autoregressive next-token-prediction model using continuous tokens, flow matching, and a resolution-normalizing positional embedding matches diffusion models on text-to-image benchmarks and generates images up to 4K.
desk verdict A strong AR text-to-image system with real benchmark numbers, but the 'first SOTA' and 4K claims outrun the evidence. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing piece is VoPE (visual rotary positional embedding), a positional encoding that maps any pixel coordinate (w,h) into a normalized g×g reference grid using a resolution density ρ and a centering offset b, so the attention dot product depends only on the relative normalized distance (m−n)/ρ. That makes positional information invariant to image scale and aspect ratio, unlike RoPE, which needs base-frequency retuning and causes blurry repetitive outputs when extrapolated. Around it, the model uses the D-JEPA architecture to predict masked continuous tokens, a flow-matching loss to denoise each predicted token, and a data-feedback training loop in which a small critic model labels whether the current generator handles a sampled example well and reweights future sampling toward cases it fails.
What would settle it
Generate a batch of 2048×2048 and 4096×4096 images from the trained model and measure a global-coherence metric such as FID against a reference set or human pairwise preference against 1K outputs from the same model; if scores degrade sharply with resolution, or if human raters prefer lower-resolution versions, the claimed arbitrary-resolution capability does not transfer and the central high-resolution claim fails.
Extended reading notes
Core claim
The central claim is that D-JEPA·T2I is, for the first time, a next-token-prediction model that achieves state-of-the-art high-resolution text-to-image synthesis. Using continuous tokens encoded by a KL-VAE, a multimodal visual transformer that fuses T5 text features with visual features, and a flow-matching objective in place of a diffusion loss, the model reaches 0.66 overall on GenEval, surpassing same-scale diffusion baselines such as SDXL and SD3.0-2B and rivaling DALL·E 3 and Fluid at larger scales. It also improves over autoregressive predecessors like LlamaGen and Emu3, and human pairwise ratings put it close to Midjourney v6. The paper attributes the resolution flexibility to VoPE, which normalizes pixel coordinates into a fixed grid so that relative positions stay consistent across resolutions, and to a random token-drop training strategy that caps each iteration at 4096 tokens, allowing 4K-scale synthesis without 4K-scale memory.
Load-bearing premise
The high-resolution claim rests on the assumption that training with at most 4096 randomly dropped tokens and VoPE's normalized coordinates transfers to full 4K sampling while keeping the image globally coherent, an assumption the paper supports only with sample images, not quantitative measurements.
Editorial extensions
If this is right
- Autoregressive text-to-image can rival diffusion at similar parameter counts: 0.66 GenEval for 2.6B parameters versus 0.55 for SDXL and 0.62 for SD3.0-2B, and ahead of open autoregressive baselines of up to 8B.
- One model covers continuous resolutions and aspect ratios without per-size fine-tuning; sampling uses at most 128 autoregressive steps regardless of resolution.
- Data feedback roughly halves early training time to reach a given GenEval score and raises the human win rate against Midjourney v6 from 17.3% to 35.6% in the late training stage.
- Adjusting the positional offset b gives explicit layout control, letting the model shift off-center subjects back into view.
Reading between the lines
- Editorial inference: if VoPE transfers as claimed, the same normalized-coordinate trick could let a single autoregressive or diffusion transformer train at low resolution and sample at arbitrary high resolutions in other modalities, such as video, where absolute positional embeddings currently force interpolation.
- Editorial inference: the paper's 4K results are qualitative only; a quantitative 4K evaluation (FID or human ratings on 2048×2048 and 4096×4096 outputs) would test whether the token-drop training preserves global coherence, since the model never sees a full high-resolution image during training.
- Editorial inference: the critic-model feedback loop is a cheap online substitute for preference fine-tuning, but it assumes the critic's labels remain aligned with actual model weaknesses as training progresses; periodic re-labeling with human judgments is what keeps that assumption valid here.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents D-JEPA·T2I, a 2.6B-parameter autoregressive text-to-image model based on continuous tokens, the D-JEPA architecture, flow-matching loss, and a new visual rotary positional embedding (VoPE) for continuous-resolution learning. A data-feedback training strategy uses statistical analysis sampling plus an online critic model to re-weight training data toward underperforming cases. The model is trained on an internal 1B+ image-text dataset and evaluated on GenEval, T2I-CompBench++, and GenAI-Bench, with a human-study win rate against Midjourney v6. The central claim is that this is the first state-of-the-art high-resolution (up to 4K) image synthesis via next-token prediction.
Significance. If the central claims hold, the paper would be a meaningful step for autoregressive text-to-image generation: it demonstrates that a relatively small (2.6B) continuous-token NTP model can match or exceed several diffusion baselines on standard alignment benchmarks, and the VoPE mechanism is a plausible solution for variable-resolution and variable-aspect-ratio generation. The data-feedback training strategy is also a useful and fairly resource-efficient idea, and the paper includes ablations (Table 4) that support its contribution. The paper is less strong on the 4K and 'first SOTA' claims: the 4K evidence is qualitative only, and the GenEval comparison in Table 3 shows Fluid (10.5B) at 0.69 vs. 0.66 here, which puts the 'state-of-the-art via next-token prediction' phrasing under an unstated scope restriction.
major comments (4)
- [Abstract and §4.4 (Fig. 6)] The abstract and Section 4.4 state that D-JEPA·T2I 'performs comparably to Midjourney v6', but Fig. 6 reports a win rate of 39.37% against Midjourney v6, below the 50% baseline. A 39% win rate is not 'comparable' in the usual sense; the text should be revised to report the actual value and to qualify the comparison, e.g., as 'competitive among open models of similar size' rather than 'comparable to Midjourney v6'.
- [§9 (Scaling to 4K Resolution) and §14 (Limitation)] The headline claim of 'high-resolution image synthesis, up to 4K' rests on Fig. 11 and the random token-drop training strategy, but no quantitative evaluation at resolutions above 1K is provided. The paper's own §9 states that random token-drop 'might limit the model's ability to learn global features', and §14 admits that 4K performance is 'less than optimal'. Since the paper provides no FID, VQAScore, GenEval, or human evaluation at 2K/4K, and no comparison against any baseline at those resolutions, the 4K-capability claim is unverified. Please add quantitative results at 2K and 4K, or explicitly scope the claim to 1K in the abstract and title.
- [Table 2 and Table 3 (GenEval comparison)] The claim of 'state-of-the-art high-resolution image synthesis via next-token prediction' is not directly supported by the GenEval numbers in Table 3: Fluid, an NTP model, achieves 0.69, which is higher than the reported 0.66. The phrase 'state-of-the-art' is therefore only valid under an unstated scope restriction (e.g., models under 3B parameters, or open-source models without DPO). Please either add the scope restriction explicitly, or soften the claim to match the data (e.g., 'state-of-the-art among sub-3B NTP models').
- [§3.2 (Critic Model Sampling) and Table 4] The critic model is trained on labels derived from T2I-CompBench and GenEval-style automated metrics, and the same benchmarks are used for final evaluation. This creates a potential feedback loop where the model is explicitly optimized toward those benchmarks. The paper should discuss this circularity risk and, ideally, report results on a held-out benchmark that was not used for critic training (e.g., a human-preference benchmark like PickScore or a different compositional benchmark). Table 4's ablation is useful, but the reported GenEval gains may partly reflect overfitting to the evaluation metric rather than general improvement.
minor comments (5)
- [§2.3, Eq. (3)] The VoPE derivation in §2.3 would benefit from a note that the normalization with ρ and b assumes max(W,H) is known at inference time; for arbitrary user-specified resolutions this is fine, but the exact handling of non-integer ρ and b is not specified.
- [§9 (Inference Details)] The paper reports that the time-shifting factor was determined by grid search to be 4.5, but does not report the search range or sensitivity; a brief sensitivity analysis would improve reproducibility.
- [Table 1, GenAI-Bench 'basic' prompts] The 'Avg' column in Table 1 appears to be computed over the five categories, but the 'basic' and 'advanced' tables have different category sets; please clarify whether the average is unweighted over the displayed categories.
- [§4.1 (Training)] The description of the second training phase says resolutions 'progressively increase from 128 to 1024 pixels', but §9 and Fig. 10 describe a dynamic resolution distribution that also samples beyond 1K. Please reconcile these two descriptions.
- [References] Reference [41] (Lumina-T2X) is cited for the flow matching formulation, but the paper uses a slightly different interpolation schedule (t x_i + (1-t) epsilon); please cite the original flow matching papers (e.g., Lipman et al. and Liu et al.) directly for this specific form.
Circularity Check
No significant circularity: the central benchmark results are measured directly against external baselines; the D-JEPA self-citation is architectural inheritance rather than a load-bearing derivation, and the critic/benchmark feedback loop does not force any reported number by construction.
full rationale
The paper's central claimed contribution is an empirical T2I model, and its headline numbers (GenEval 0.66, T2I-CompBench++, GenAI-Bench, human win rates) are computed by running the trained model on public benchmarks and comparing to external systems; none of those numbers is an identity or a fitted quantity renamed as a prediction. The same-author citation to D-JEPA [20] supplies the backbone architecture and the Lpred/Lflow masking recipe, but the present model's performance is not inferred from that citation; it is measured in this paper's own experiments (e.g., Table 2, Table 4, Fig. 6). The critic model is trained on labels derived from T2I-CompBench and then used to reweight training data, and the same benchmark family appears in the final evaluation table; this is a legitimate benchmark-alignment risk, but it is not circular in the by-construction sense because the critic is a binary filter over a few thousand prompts and the reported metrics are aggregate scores on benchmark suites, so the improvement is not logically forced. VoPE's resolution invariance is obtained by defining normalized coordinates (1/rho)(m-n), which is a designed reparameterization of RoPE rather than a hidden reuse of the result it is supposed to explain; the claim that 1K training suffices for 4K generation is an empirical assertion ("we found that through dynamic resolution training... D-JEPA·T2I can quickly adapt"), not a mathematical consequence. The paper itself flags the relevant limitations: Section 9 concedes random token-drop "might limit the model's ability to learn global features," and Section 14 states 4K performance is "less than optimal"; no quantitative evaluation above 1K is provided, and Table 3 shows Fluid (a next-token-prediction model) at GenEval 0.69 versus 0.66 here, so the "first/SOTA" phrasing is under-scoped. These are evidence gaps and correctness risks, not circular derivations; the derivation chain is self-contained against external benchmarks, so the circularity score is 0.
Assumptions & free parameters
free parameters (4)
- Classifier-free guidance scale =
6.0 for benchmark evaluation, 2.0 otherwise
- Time shifting factor =
4.5
- Truncated normal resolution sampling parameters =
trunc_norm(0.512, 0.12, 0.256, 1.024) and trunc_norm(1.024, 1.0, 0.256, 4.096)
- Number of sampling AR steps =
64 for 256x256, 128 for higher
assumptions (4)
- domain assumption D-JEPA architecture and its training objective are effective for class-conditional generation, and this extends to text-to-image.
- domain assumption The KL-VAE from Esser et al. provides a good continuous latent space for images.
- standard math Flow matching loss with a linear interpolation is an appropriate replacement for diffusion loss in token prediction.
- ad hoc to paper The critic model's rejection probability identifies samples the T2I model will struggle on.
Cite this review
Pith. "Pith review of High-Resolution Image Synthesis via Next-Token Prediction." pith.science (2026). https://pith.science/paper/TYB2U2RT
@misc{pith2026241114808,
author = {Pith},
title = {Pith review of: High-Resolution Image Synthesis via Next-Token Prediction},
year = {2026},
howpublished = {\url{https://pith.science/paper/TYB2U2RT}},
note = {Machine review of arXiv:2411.14808}
}
abstract
Recently, autoregressive models have demonstrated remarkable performance in class-conditional image generation. However, the application of next-token prediction to high-resolution text-to-image generation remains largely unexplored. In this paper, we introduce \textbf{D-JEPA$\cdot$T2I}, an autoregressive model based on continuous tokens that incorporates innovations in both architecture and training strategy to generate high-quality, photorealistic images at arbitrary resolutions, up to 4K. Architecturally, we adopt the denoising joint embedding predictive architecture (D-JEPA) while leveraging a multimodal visual transformer to effectively integrate textual and visual features. Additionally, we introduce flow matching loss alongside the proposed Visual Rotary Positional Embedding (VoPE) to enable continuous resolution learning. In terms of training strategy, we propose a data feedback mechanism that dynamically adjusts the sampling procedure based on statistical analysis and an online learning critic model. This encourages the model to move beyond its comfort zone, reducing redundant training on well-mastered scenarios and compelling it to address more challenging cases with suboptimal generation quality. For the first time, we achieve state-of-the-art high-resolution image synthesis via next-token prediction.
Figures
Figures from the paper (18 more)
Reference graph
Works this paper leans on
-
[1]
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023. 1, 3
arXiv 2023
-
[2]
A window with raindrops trickling down, overlooking a blurry city
Michael S. Albergo and Eric Vanden-Eijnden. Building 11 Figure 16. D-JEPA ·T2I can generate arbitrary aspect ratios and continuous resolutions with V oPE. Prompt: “A window with raindrops trickling down, overlooking a blurry city.” normalizing flows with stochastic interpolants, 2022. 1
2022
-
[3]
Stochastic interpolants: A unifying framework for flows and diffusions
Michael S Albergo, Nicholas M Boffi, and Eric Vanden- Eijnden. Stochastic interpolants: A unifying framework for flows and diffusions. arXiv preprint arXiv:2303.08797,
-
[4]
Rohan Anil, Andrew M Dai, Orhan Firat, Melvin John- son, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, et al. Palm 2 technical report. arXiv preprint arXiv:2305.10403, 2023. 1
arXiv 2023
-
[5]
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al. Qwen technical report. arXiv preprint arXiv:2309.16609, 2023. 1
arXiv 2023
-
[6]
Layout control by relative positional offset b
Yogesh Balaji, Seungjun Nah, Xun Huang, Arash Vahdat, Jiaming Song, Qinsheng Zhang, Karsten Kreis, Miika Ait- tala, Timo Aila, Samuli Laine, Bryan Catanzaro, Tero Kar- 12 𝑏=384 𝑏=320 𝑏=256 Figure 17. Layout control by relative positional offset b. By ad- justing b, we can generate more desirable layouts for selection. ras, and Ming-Yu Liu. ediff-i: Text-t...
2022
-
[7]
Fan Bao, Chongxuan Li, Jun Zhu, and Bo Zhang. Analytic- dpm: an analytic estimate of the optimal reverse vari- ance in diffusion probabilistic models. arXiv preprint arXiv:2201.06503, 2022. 1
arXiv 2022
-
[8]
All are worth words: A vit backbone for diffusion models
Fan Bao, Shen Nie, Kaiwen Xue, Yue Cao, Chongxuan Li, Hang Su, and Jun Zhu. All are worth words: A vit backbone for diffusion models. In Proceedings of the IEEE/CVF con- ference on computer vision and pattern recognition , pages 22669–22679, 2023. 1
2023
Show all 137 references
-
[9]
Vidu: a highly consistent, dynamic and skilled text-to-video generator with diffusion models
Fan Bao, Chendong Xiang, Gang Yue, Guande He, Hongzhou Zhu, Kaiwen Zheng, Min Zhao, Shilong Liu, Yaole Wang, and Jun Zhu. Vidu: a highly consistent, dynamic and skilled text-to-video generator with diffusion models. arXiv preprint arXiv:2405.04233, 2024. 1
2024 arXiv
-
[10]
Beit: Bert pre-training of image transformers
Hangbo Bao, Li Dong, Songhao Piao, and Furu Wei. Beit: Bert pre-training of image transformers. arXiv preprint arXiv:2106.08254, 2021. 1
2021 arXiv
-
[11]
Lumiere: A space- time diffusion model for video generation
Omer Bar-Tal, Hila Chefer, Omer Tov, Charles Her- rmann, Roni Paiss, Shiran Zada, Ariel Ephrat, Junhwa Hur, Yuanzhen Li, Tomer Michaeli, et al. Lumiere: A space- time diffusion model for video generation. arXiv preprint arXiv:2401.12945, 2024. 1
2024 arXiv
-
[12]
Improving image generation with better captions
James Betker, Gabriel Goh, Li Jing, Tim Brooks, Jianfeng Wang, Linjie Li, Long Ouyang, Juntang Zhuang, Joyce Lee, Yufei Guo, et al. Improving image generation with better captions. Computer Science. https://cdn. openai. com/papers/dall-e-3. pdf, 2(3), 2023. 1, 7, 8, 4, 14
2023
-
[13]
Stable video diffusion: Scaling latent video diffusion models to large datasets
Andreas Blattmann, Tim Dockhorn, Sumith Kulal, Daniel Mendelevitch, Maciej Kilian, Dominik Lorenz, Yam Levi, Zion English, Vikram V oleti, Adam Letts, et al. Stable video diffusion: Scaling latent video diffusion models to large datasets. arXiv preprint arXiv:2311.15127, 2023
2023 arXiv
-
[14]
Align your latents: High-resolution video synthe- sis with latent diffusion models
Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dockhorn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis. Align your latents: High-resolution video synthe- sis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,...
2023
-
[15]
Brooks, B
T. Brooks, B. Peebles, C. Holmes, W. DePue, Y . Guo, L. Jing, D. Schnurr, J. Taylor, T. Luhman, E. Luhman, C. Ng, R. Wang, and A. Ramesh. Video generation models as world simulators. OpenAI, 2024. 1
2024
-
[16]
Language models are few-shot learners
Tom B Brown. Language models are few-shot learners. arXiv preprint arXiv:2005.14165, 2020. 1
2005 arXiv
-
[17]
Maskgit: Masked generative image transformer
Huiwen Chang, Han Zhang, Lu Jiang, Ce Liu, and William T Freeman. Maskgit: Masked generative image transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11315– 11325, 2022. 4, 1
2022
-
[18]
Muse: Text- to-image generation via masked generative transformers
Huiwen Chang, Han Zhang, Jarred Barber, AJ Maschinot, Jose Lezama, Lu Jiang, Ming-Hsuan Yang, Kevin Murphy, William T Freeman, Michael Rubinstein, et al. Muse: Text- to-image generation via masked generative transformers. arXiv preprint arXiv:2301.00704, 2023. 1
2023 arXiv
-
[19]
Bfloat16: The secret to high performance on cloud tpus, 2019
Dehao Chen, Chiachen Chou, Yuanzhong Xu, and Jonathan Hseu. Bfloat16: The secret to high performance on cloud tpus, 2019. 2
2019
-
[20]
Denoising with a joint-embedding predictive architecture
Dengsheng Chen, Jie Hu, Xiaoming Wei, and Enhua Wu. Denoising with a joint-embedding predictive architecture. arXiv preprint arXiv:2410.03755, 2024. 1, 2, 3, 5, 7
2024 arXiv
-
[21]
Pixart- α: Fast training of diffu- sion transformer for photorealistic text-to-image synthesis
Junsong Chen, Jincheng Yu, Chongjian Ge, Lewei Yao, Enze Xie, Yue Wu, Zhongdao Wang, James Kwok, Ping Luo, Huchuan Lu, et al. Pixart- α: Fast training of diffu- sion transformer for photorealistic text-to-image synthesis. arXiv preprint arXiv:2310.00426, 2023. 7, 8, 1, 4, 17
-
[22]
Pixart- σ: Weak-to-strong training of diffusion transformer for 4k text-to-image generation
Junsong Chen, Chongjian Ge, Enze Xie, Yue Wu, Lewei Yao, Xiaozhe Ren, Zhongdao Wang, Ping Luo, Huchuan Lu, and Zhenguo Li. Pixart- σ: Weak-to-strong training of diffusion transformer for 4k text-to-image generation. arXiv preprint arXiv:2403.04692, 2024. 1, 7
2024 arXiv
-
[23]
Neural ordinary differential equa- tions
Ricky TQ Chen, Yulia Rubanova, Jesse Bettencourt, and David K Duvenaud. Neural ordinary differential equa- tions. Advances in neural information processing systems, 31, 2018. 1
2018
-
[24]
Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks
Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, Bin Li, Ping Luo, Tong Lu, Yu Qiao, and Jifeng Dai. Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks. arXiv prep...
2023 arXiv
-
[25]
How far are we to gpt-4v? clos- ing the gap to commercial multimodal models with open- source suites
Zhe Chen, Weiyun Wang, Hao Tian, Shenglong Ye, Zhang- wei Gao, Erfei Cui, Wenwen Tong, Kongzhi Hu, Jiapeng Luo, Zheng Ma, et al. How far are we to gpt-4v? clos- ing the gap to commercial multimodal models with open- source suites. arXiv preprint arXiv:2404.16821, 2024. 7 13 Fi...
2024 arXiv
-
[26]
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. Palm: Scaling language modeling with pathways. Journal of Machine Learning Research, 24(240): 1–113, 2023. 1
2023
-
[27]
Scaling instruction- finetuned language models
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. Scaling instruction- finetuned language models. Journal of Machine Learning Research, 25(70):1–53, 2024. 3
2024
-
[28]
Emu: Enhancing image generation models using photogenic nee- dles in a haystack, 2023
Xiaoliang Dai, Ji Hou, Chih-Yao Ma, Sam Tsai, Jialiang Wang, Rui Wang, Peizhao Zhang, Simon Vandenhende, Xi- aofang Wang, Abhimanyu Dubey, Matthew Yu, Abhishek Kadian, Filip Radenovic, Dhruv Mahajan, Kunpeng Li, Yue Zhao, Vladan Petrovic, Mitesh Kumar Singh, Simran Mot- wani, ...
2023
-
[29]
Flow matching in latent space, 2023
Quan Dao, Hao Phung, Binh Nguyen, and Anh Tran. Flow matching in latent space, 2023. 1
2023
-
[30]
Mostafa Dehghani, Josip Djolonga, Basil Mustafa, Pi- otr Padlewski, Jonathan Heek, Justin Gilmer, Andreas Steiner, Mathilde Caron, Robert Geirhos, Ibrahim Al- abdulmohsin, Rodolphe Jenatton, Lucas Beyer, Michael Tschannen, Anurag Arnab, Xiao Wang, Carlos Riquelme, Matthias Min...
2023
-
[31]
Diffusion mod- els beat gans on image synthesis
Prafulla Dhariwal and Alexander Nichol. Diffusion mod- els beat gans on image synthesis. In Advances in neural information processing systems, pages 8780–8794, 2021. 1
2021
-
[32]
Score- based generative modeling with critically-damped langevin diffusion
Tim Dockhorn, Arash Vahdat, and Karsten Kreis. Score- based generative modeling with critically-damped langevin diffusion. arXiv preprint arXiv:2112.07068, 2021. 1
2021 arXiv
-
[33]
Genie: Higher-order denoising diffusion solvers, 2022
Tim Dockhorn, Arash Vahdat, and Karsten Kreis. Genie: Higher-order denoising diffusion solvers, 2022. 1
2022
-
[34]
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020. 2
2010 arXiv
-
[35]
Structure and content-guided video synthesis with diffusion models
Patrick Esser, Johnathan Chiu, Parmida Atighehchian, Jonathan Granskog, and Anastasis Germanidis. Structure and content-guided video synthesis with diffusion models. 14 A photo of a cat playing chess.A bird made of crystalA pair of old boots covered in mud. Photo of a bear cat...
-
[36]
Scaling rectified flow transformers for high-resolution image synthesis
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas M ¨uller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, et al. Scaling rectified flow transformers for high-resolution image synthesis. In Forty-first International Conference on Machin...
-
[37]
Fluid: Scaling autoregressive text-to-image generative models with continuous tokens
Lijie Fan, Tianhong Li, Siyang Qin, Yuanzhen Li, Chen Sun, Michael Rubinstein, Deqing Sun, Kaiming He, and Yonglong Tian. Fluid: Scaling autoregressive text-to-image generative models with continuous tokens. arXiv preprint arXiv:2410.13863, 2024. 8, 4, 15
-
[38]
Ernie-vilg 2.0: Improving text-to- image diffusion model with knowledge-enhanced mixture- of-denoising-experts
Zhida Feng, Zhenyu Zhang, Xintong Yu, Yewei Fang, Lanxin Li, Xuyi Chen, Yuxiang Lu, Jiaxiang Liu, Weichong Yin, Shikun Feng, et al. Ernie-vilg 2.0: Improving text-to- image diffusion model with knowledge-enhanced mixture- of-denoising-experts. In Proceedings of the IEEE/CVF Co...
2023
-
[39]
Boost- ing latent diffusion with flow matching
Johannes S Fischer, Ming Gui, Pingchuan Ma, Nick Stracke, Stefan A Baumann, and Bj ¨orn Ommer. Boost- ing latent diffusion with flow matching. arXiv preprint arXiv:2312.07360, 2023. 1
2023 arXiv
-
[40]
If: A github repository
Deep Floyd. If: A github repository. https://github. com/deep-floyd/IF, 2023. 7
2023
-
[41]
Lumina-t2x: Transforming text into any modality, resolution, and duration via flow-based large diffusion transformers
Peng Gao, Le Zhuo, Ziyi Lin, Chris Liu, Junsong Chen, Ruoyi Du, Enze Xie, Xu Luo, Longtian Qiu, Yuhang Zhang, et al. Lumina-t2x: Transforming text into any modality, resolution, and duration via flow-based large diffusion transformers. arXiv preprint arXiv:2405.05945 ,
-
[42]
Geneval: An object-focused framework for evaluating text- to-image alignment
Dhruba Ghosh, Hannaneh Hajishirzi, and Ludwig Schmidt. Geneval: An object-focused framework for evaluating text- to-image alignment. Advances in Neural Information Pro- cessing Systems, 36, 2024. 2, 7, 8, 3, 4, 5
2024
-
[43]
Photorealistic video generation with diffusion models,
Agrim Gupta, Lijun Yu, Kihyuk Sohn, Xiuye Gu, Meera Hahn, Li Fei-Fei, Irfan Essa, Lu Jiang, and Jos ´e Lezama. Photorealistic video generation with diffusion models,
-
[44]
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceed- 16 Figure 21. Visual comparison between PixelArt-α [21] and D-JEPA·T2I. ings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016. 6
2016
-
[45]
Masked autoencoders are scal- able vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Doll´ar, and Ross Girshick. Masked autoencoders are scal- able vision learners. In Proceedings of the IEEE/CVF con- ference on computer vision and pattern recognition , pages 16000–16009, 2022. 1
2022
-
[46]
Rethinking image aesthetics assessment: Models, datasets and benchmarks
Shuai He, Yongchang Zhang, Rui Xie, Dongxiang Jiang, and Anlong Ming. Rethinking image aesthetics assessment: Models, datasets and benchmarks. In IJCAI, pages 942– 948, 2022. 1
2022
-
[47]
Classifier-free diffusion guidance
Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance. In NeurIPS 2021 Workshop on Deep Generative Models and Downstream Applications, 2021. 8, 1
2021
-
[48]
Classifier-free diffusion guidance
Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance. arXiv preprint arXiv:2207.12598, 2022. 4
2022 arXiv
-
[49]
Denoising dif- fusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising dif- fusion probabilistic models. In Advances in neural infor- mation processing systems, pages 6840–6851, 2020. 1
2020
-
[50]
Visual comparison between HunyunDit [65] and D-JEPA·T2I
Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, 17 Figure 22. Visual comparison between HunyunDit [65] and D-JEPA·T2I. Ruiqi Gao, Alexey Gritsenko, Diederik P. Kingma, Ben Poole, Mohammad Norouzi, David J. Fleet, and Tim Sali- mans. Imagen video: High definition video g...
2022
-
[51]
Cascaded diffusion models for high fidelity image generation
Jonathan Ho, Chitwan Saharia, William Chan, David J Fleet, Mohammad Norouzi, and Tim Salimans. Cascaded diffusion models for high fidelity image generation. Jour- nal of Machine Learning Research, 23(47):1–33, 2022. 1
2022
-
[52]
Training compute-optimal large language mod- els
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al. Training compute-optimal large language mod- els. arXiv preprint arXiv:2203.15556, 2022. 1
2022 arXiv
-
[53]
T2i-compbench: A comprehensive bench- mark for open-world compositional text-to-image genera- tion
Kaiyi Huang, Kaiyue Sun, Enze Xie, Zhenguo Li, and Xihui Liu. T2i-compbench: A comprehensive bench- mark for open-world compositional text-to-image genera- tion. Advances in Neural Information Processing Systems, 36:78723–78747, 2023. 2, 6, 8, 4
2023
-
[54]
Estimation of non- normalized statistical models by score matching
Aapo Hyv ¨arinen and Peter Dayan. Estimation of non- normalized statistical models by score matching. Journal of Machine Learning Research, 6(4), 2005. 1
2005
-
[55]
Tero Karras, Miika Aittala, Timo Aila, and Samuli Laine. 18 A tranquil, anime-style koi pond in a serene Japanese garden, featuring blossoming cherry trees.A raccoon wearing cowboy hat and black leather jacket is behind the backyard window. Rain droplets on the window. Transfu...
2022
-
[56]
Bert: Pre-training of deep bidirectional trans- formers for language understanding
Jacob Devlin Ming-Wei Chang Kenton and Lee Kristina Toutanova. Bert: Pre-training of deep bidirectional trans- formers for language understanding. In Proceedings of naacL-HLT, page 2. Minneapolis, Minnesota, 2019. 1
2019
-
[57]
Computational tradeoffs in image synthesis: Diffusion, masked-token, and next-token prediction
Maciej Kilian, Varun Jampani, and Luke Zettlemoyer. Computational tradeoffs in image synthesis: Diffusion, masked-token, and next-token prediction. arXiv preprint arXiv:2405.13218, 2024. 1
2024 arXiv
-
[58]
Understanding diffu- sion objectives as the elbo with simple data augmentation
Diederik Kingma and Ruiqi Gao. Understanding diffu- sion objectives as the elbo with simple data augmentation. Advances in Neural Information Processing Systems , 36,
-
[59]
Pick-a-pic: An open dataset of user preferences for text-to-image generation
Yuval Kirstain, Adam Polyak, Uriel Singer, Shahbuland Matiana, Joe Penna, and Omer Levy. Pick-a-pic: An open dataset of user preferences for text-to-image generation. Advances in Neural Information Processing Systems , 36: 36652–36663, 2023. 8, 4
2023
-
[60]
Flux.1: An open-source image genera- tion model
Black Forest Labs. Flux.1: An open-source image genera- tion model. https://www.basedlabs.ai/tools/ 19 flux1, 2024. 4
2024
-
[61]
Bloom: A 176b-parameter open-access multilingual language model
Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ili ´c, Daniel Hesslow, Roman Castagn ´e, Alexandra Sasha Luccioni, Franc ¸ois Yvon, Matthias Gall´e, et al. Bloom: A 176b-parameter open-access multilingual language model. 2023. 1
2023
-
[62]
Minimizing trajectory curvature of ode-based generative models, 2023
Sangyun Lee, Beomsu Kim, and Jong Chul Ye. Minimizing trajectory curvature of ode-based generative models, 2023. 1
2023
-
[63]
Genai-bench: Evaluating and im- proving compositional text-to-visual generation
Baiqi Li, Zhiqiu Lin, Deepak Pathak, Jiayao Li, Yixin Fei, Kewen Wu, Tiffany Ling, Xide Xia, Pengchuan Zhang, Graham Neubig, et al. Genai-bench: Evaluating and im- proving compositional text-to-visual generation. arXiv preprint arXiv:2406.13743, 2024. 2, 7, 8
2024 arXiv
-
[64]
Autoregressive image generation without vec- tor quantization
Tianhong Li, Yonglong Tian, He Li, Mingyang Deng, and Kaiming He. Autoregressive image generation without vec- tor quantization. arXiv preprint arXiv:2406.11838, 2024. 1, 3, 4
2024 arXiv
-
[65]
Hunyuan- dit: A powerful multi-resolution diffusion transformer with fine-grained chinese understanding
Zhimin Li, Jianwei Zhang, Qin Lin, Jiangfeng Xiong, Yanxin Long, Xinchi Deng, Yingfang Zhang, Xingchao Liu, Minbin Huang, Zedong Xiao, et al. Hunyuan- dit: A powerful multi-resolution diffusion transformer with fine-grained chinese understanding. arXiv preprint arXiv:2405.0874...
2024 arXiv
-
[66]
Rich human feedback for text-to-image generation
Youwei Liang, Junfeng He, Gang Li, Peizhao Li, Arseniy Klimovskiy, Nicholas Carolan, Jiao Sun, Jordi Pont-Tuset, Sarah Young, Feng Yang, et al. Rich human feedback for text-to-image generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognitio...
2024
-
[67]
Flow matching for generative modeling
Yaron Lipman, Ricky TQ Chen, Heli Ben-Hamu, Maxim- ilian Nickel, and Matt Le. Flow matching for generative modeling. arXiv preprint arXiv:2210.02747, 2022. 1, 3
2022 arXiv
-
[68]
Yaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maxi- milian Nickel, and Matthew Le. Flow matching for gener- ative modeling. In The Eleventh International Conference on Learning Representations, 2023. 1
2023
-
[69]
Lumina-mgpt: Il- luminate flexible photorealistic text-to-image generation with multimodal generative pretraining
Dongyang Liu, Shitian Zhao, Le Zhuo, Weifeng Lin, Yu Qiao, Hongsheng Li, and Peng Gao. Lumina-mgpt: Il- luminate flexible photorealistic text-to-image generation with multimodal generative pretraining. arXiv preprint arXiv:2408.02657, 2024. 1
2024 arXiv
-
[70]
Flow straight and fast: Learning to generate and transfer data with rectified flow.arXiv preprint arXiv:2209.03003, 2022
Xingchao Liu, Chengyue Gong, and Qiang Liu. Flow straight and fast: Learning to generate and transfer data with rectified flow.arXiv preprint arXiv:2209.03003, 2022. 1
2022 arXiv
-
[71]
Instaflow: One step is enough for high-quality diffusion-based text-to-image generation, 2023
Xingchao Liu, Xiwen Zhang, Jianzhu Ma, Jian Peng, and Qiang Liu. Instaflow: One step is enough for high-quality diffusion-based text-to-image generation, 2023. 1
2023
-
[72]
Fixing weight decay regularization in adam
Ilya Loshchilov, Frank Hutter, et al. Fixing weight decay regularization in adam. arXiv preprint arXiv:1711.05101, 5, 2017. 2
2017 arXiv
-
[73]
Dpm-solver++: Fast solver for guided sampling of diffusion probabilistic models, 2023
Cheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen, Chongxuan Li, and Jun Zhu. Dpm-solver++: Fast solver for guided sampling of diffusion probabilistic models, 2023. 1
2023
-
[74]
Albergo, Nicholas M
Nanye Ma, Mark Goldstein, Michael S. Albergo, Nicholas M. Boffi, Eric Vanden-Eijnden, and Saining Xie. Sit: Exploring flow and diffusion-based generative models with scalable interpolant transformers, 2024. 3, 1
2024
-
[75]
Midjourney v6 - ai art generator,
Midjourney Community. Midjourney v6 - ai art generator,
-
[76]
Freecontrol: Training-free spatial control of any text-to-image diffusion model with any condition
Sicheng Mo, Fangzhou Mu, Kuan Heng Lin, Yanli Liu, Bochen Guan, Yin Li, and Bolei Zhou. Freecontrol: Training-free spatial control of any text-to-image diffusion model with any condition. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pag...
2024
-
[77]
Glide: Towards photorealistic image genera- tion and editing with text-guided diffusion models
Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen. Glide: Towards photorealistic image genera- tion and editing with text-guided diffusion models. arXiv preprint arXiv:2112.10741, 2021. 1
2021 arXiv
-
[78]
Training lan- guage models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Car- roll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training lan- guage models to follow instructions with human feedback. Advances in neural information processing systems , 35...
2022
-
[79]
Pavlov, A
I. Pavlov, A. Ivanov, and S. Stafievskiy. Text-to- Image Benchmark: A benchmark for generative models. https://github.com/boomb0om/text2image- benchmark, 2023. Version 0.1.0. 3
2023
-
[80]
Scalable diffusion mod- els with transformers
William Peebles and Saining Xie. Scalable diffusion mod- els with transformers. In Proceedings of the IEEE/CVF In- ternational Conference on Computer Vision , pages 4195– 4205, 2023. 2, 5, 1
2023
-
[81]
Ntk-aware scaled rope allows llama models to have extended (8k+) context size without any fine-tuning and minimal perplexity degrada- tion, 2023
Bowen Peng and Jeffrey Quesnelle. Ntk-aware scaled rope allows llama models to have extended (8k+) context size without any fine-tuning and minimal perplexity degrada- tion, 2023. 4
2023
-
[82]
Film: Visual reasoning with a general conditioning layer
Ethan Perez, Florian Strub, Harm De Vries, Vincent Du- moulin, and Aaron Courville. Film: Visual reasoning with a general conditioning layer. In Proceedings of the AAAI conference on artificial intelligence, 2018. 2
2018
-
[83]
Sdxl: Improving latent diffusion models for high-resolution image synthesis, 2023
Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas M ¨uller, Joe Penna, and Robin Rombach. Sdxl: Improving latent diffusion models for high-resolution image synthesis, 2023. 1, 7, 8, 4
2023
-
[84]
Aram-Alexandre Pooladian, Heli Ben-Hamu, Carles Domingo-Enrich, Brandon Amos, Yaron Lipman, and Ricky T. Q. Chen. Multisample flow matching: Straight- ening flows with minibatch couplings, 2023. 1
2023
-
[85]
Improving language understanding by gen- erative pre-training
Alec Radford. Improving language understanding by gen- erative pre-training. 2018. 1
2018
-
[86]
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9,
-
[87]
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christo- pher D Manning, Stefano Ermon, and Chelsea Finn. Direct preference optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems, 36, 2024. 2, 6
2024
-
[88]
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, 20 and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of machine learning research, 21(140):1–67, 2020. 1
2020
-
[89]
Zero-shot text-to-image generation
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea V oss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. In Interna- tional conference on machine learning , pages 8821–8831. Pmlr, 2021. 1
2021
-
[90]
Hierarchical text-conditional image generation with clip latents, 2022
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with clip latents, 2022. 1, 4
2022
-
[91]
Deepspeed: System optimizations enable training deep learning models with over 100 billion param- eters
Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He. Deepspeed: System optimizations enable training deep learning models with over 100 billion param- eters. In Proceedings of the 26th ACM SIGKDD Interna- tional Conference on Knowledge Discovery & Data Min- ing, p...
2020
-
[92]
Gener- ating diverse high-fidelity images with vq-vae-2
Ali Razavi, Aaron Van den Oord, and Oriol Vinyals. Gener- ating diverse high-fidelity images with vq-vae-2. Advances in neural information processing systems, 32, 2019. 1
2019
-
[93]
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022. 1, 7, 4
2022
-
[94]
Im- agenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al. Im- agenet large scale visual recognition challenge. Interna- tional journal of computer vision, 115:211–252, 2015. 1
2015
-
[95]
Photorealistic text-to-image diffusion models with deep language understanding
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Sali- mans, et al. Photorealistic text-to-image diffusion models with deep language understanding. Advances in neural in- forma...
2022
-
[96]
Adversarial diffusion distillation
Axel Sauer, Dominik Lorenz, Andreas Blattmann, and Robin Rombach. Adversarial diffusion distillation. In European Conference on Computer Vision, pages 87–103. Springer, 2025. 7
2025
-
[97]
Laion-5b: An open large-scale dataset for train- ing next generation image-text models
Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Worts- man, et al. Laion-5b: An open large-scale dataset for train- ing next generation image-text models. Advances in Neural Inf...
2022
-
[98]
Make-a-video: Text-to-video generation without text-video data, 2022
Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, Devi Parikh, Sonal Gupta, and Yaniv Taigman. Make-a-video: Text-to-video generation without text-video data, 2022. 1
2022
-
[99]
Deep unsupervised learning using nonequilibrium thermodynamics
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. In International confer- ence on machine learning, pages 2256–2265. PMLR, 2015. 1
2015
-
[100]
Denoising diffusion implicit models
Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. arXiv preprint arXiv:2010.02502, 2020. 1
2010 arXiv
-
[101]
Denois- ing diffusion implicit models, 2022
Jiaming Song, Chenlin Meng, and Stefano Ermon. Denois- ing diffusion implicit models, 2022. 1
2022
-
[102]
Mugs: A multiple granularity semi-supervised method for text recognition
Qi Song, Qianyi Jiang, Lei Wang, Lingling Zhao, and Rui Zhang. Mugs: A multiple granularity semi-supervised method for text recognition. In International Conference on Document Analysis and Recognition , pages 173–188. Springer, 2023. 1
2023
-
[103]
Generative modeling by estimating gradients of the data distribution
Yang Song and Stefano Ermon. Generative modeling by estimating gradients of the data distribution. In Advances in neural information processing systems, 2019. 1
2019
-
[104]
Generative modeling by estimating gradients of the data distribution, 2020
Yang Song and Stefano Ermon. Generative modeling by estimating gradients of the data distribution, 2020. 1
2020
-
[105]
Score- based generative modeling through stochastic differential equations
Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. Score- based generative modeling through stochastic differential equations. In International Conference on Learning Repre- sentations. 1
-
[106]
Score- based generative modeling through stochastic differential equations
Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. Score- based generative modeling through stochastic differential equations. arXiv preprint arXiv:2011.13456, 2020. 1
2011 arXiv
-
[107]
Roformer: Enhanced transformer with rotary position embedding
Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu. Roformer: Enhanced transformer with rotary position embedding. Neurocomputing, 568: 127063, 2024. 1, 3, 4
2024
-
[108]
Journeydb: A benchmark for genera- tive image understanding
Keqiang Sun, Junting Pan, Yuying Ge, Hao Li, Haodong Duan, Xiaoshi Wu, Renrui Zhang, Aojun Zhou, Zipeng Qin, Yi Wang, et al. Journeydb: A benchmark for genera- tive image understanding. Advances in Neural Information Processing Systems, 36, 2024. 7
2024
-
[109]
Autoregressive model beats diffusion: Llama for scalable image generation.arXiv preprint arXiv:2406.06525, 2024
Peize Sun, Yi Jiang, Shoufa Chen, Shilong Zhang, Bingyue Peng, Ping Luo, and Zehuan Yuan. Autoregressive model beats diffusion: Llama for scalable image generation.arXiv preprint arXiv:2406.06525, 2024. 1, 8, 4
2024 arXiv
-
[110]
Ernie 3.0: Large-scale knowl- edge enhanced pre-training for language understanding and generation
Yu Sun, Shuohuan Wang, Shikun Feng, Siyu Ding, Chao Pang, Junyuan Shang, Jiaxiang Liu, Xuyi Chen, Yanbin Zhao, Yuxiang Lu, et al. Ernie 3.0: Large-scale knowl- edge enhanced pre-training for language understanding and generation. arXiv preprint arXiv:2107.02137, 2021. 1
2021 arXiv
-
[111]
Chameleon: Mixed-modal early-fusion foundation models
Chameleon Team. Chameleon: Mixed-modal early-fusion foundation models. arXiv preprint arXiv:2405.09818 ,
-
[112]
Dreamina: Free ai image generator
Dreamina Team. Dreamina: Free ai image generator. https://jimeng.jianying.com/, 2023. 7
2023
-
[113]
Gem- ini: a family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Jo- han Schalkwyk, Andrew M Dai, Anja Hauth, et al. Gem- ini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805, 2023. 1
2023 arXiv
-
[114]
Kolors: Effective training of diffusion model for photorealistic text-to-image synthesis
Kolors Team. Kolors: Effective training of diffusion model for photorealistic text-to-image synthesis. arXiv preprint,
-
[115]
Wanxiang Team. Wanx. urlhttps://tongyi.aliyun.com/wanxiang/, 2023. 7 21
2023
-
[116]
Visual autoregressive modeling: Scalable im- age generation via next-scale prediction
Keyu Tian, Yi Jiang, Zehuan Yuan, Bingyue Peng, and Li- wei Wang. Visual autoregressive modeling: Scalable im- age generation via next-scale prediction. arXiv preprint arXiv:2404.02905, 2024. 1
2024 arXiv
-
[117]
Improving and generalizing flow- based generative models with minibatch optimal transport,
Alexander Tong, Nikolay Malkin, Guillaume Huguet, Yan- lei Zhang, Jarrid Rector-Brooks, Kilian Fatras, Guy Wolf, and Yoshua Bengio. Improving and generalizing flow- based generative models with minibatch optimal transport,
-
[118]
Llama: Open and efficient foundation language mod- els
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timoth ´ee Lacroix, Bap- tiste Rozi `ere, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language mod- els. arXiv preprint arXiv:2302.13971, 2023. 1
2023 arXiv
-
[119]
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023. 1
2023 arXiv
-
[120]
A connection between score matching and denoising autoencoders
Pascal Vincent. A connection between score matching and denoising autoencoders. Neural computation, 23(7):1661– 1674, 2011. 1
2011
-
[121]
Diffusion model alignment using direct preference optimization
Bram Wallace, Meihua Dang, Rafael Rafailov, Linqi Zhou, Aaron Lou, Senthil Purushwalkam, Stefano Ermon, Caim- ing Xiong, Shafiq Joty, and Nikhil Naik. Diffusion model alignment using direct preference optimization. InProceed- ings of the IEEE/CVF Conference on Computer Vision ...
2024
-
[122]
Emu3: Next-token prediction is all you need
Xinlong Wang, Xiaosong Zhang, Zhengxiong Luo, Quan Sun, Yufeng Cui, Jinsheng Wang, Fan Zhang, Yueze Wang, Zhen Li, Qiying Yu, et al. Emu3: Next-token prediction is all you need. arXiv preprint arXiv:2409.18869, 2024. 8, 4
2024 arXiv
-
[123]
Pytorch image models
Ross Wightman. Pytorch image models. https : / / github . com / rwightman / pytorch - image - models, 2019. 6
2019
-
[124]
Liu, Lechao Xiao, Katie Ev- erett, Alex Alemi, Ben Adlam, John D
Mitchell Wortsman, Peter J. Liu, Lechao Xiao, Katie Ev- erett, Alex Alemi, Ben Adlam, John D. Co-Reyes, Izzed- din Gur, Abhishek Kumar, Roman Novak, Jeffrey Pen- nington, Jascha Sohl-dickstein, Kelvin Xu, Jaehoon Lee, Justin Gilmer, and Simon Kornblith. Small-scale proxies for...
2023
-
[125]
Show-o: One single transformer to unify multimodal under- standing and generation
Jinheng Xie, Weijia Mao, Zechen Bai, David Junhao Zhang, Weihao Wang, Kevin Qinghong Lin, Yuchao Gu, Zhijie Chen, Zhenheng Yang, and Mike Zheng Shou. Show-o: One single transformer to unify multimodal under- standing and generation. arXiv preprint arXiv:2408.12528,
-
[126]
Imagere- ward: Learning and evaluating human preferences for text- to-image generation
Jiazheng Xu, Xiao Liu, Yuchen Wu, Yuxuan Tong, Qinkai Li, Ming Ding, Jie Tang, and Yuxiao Dong. Imagere- ward: Learning and evaluating human preferences for text- to-image generation. Advances in Neural Information Pro- cessing Systems, 36, 2024. 2, 6
2024
-
[127]
Raphael: Text- to-image generation via large mixture of diffusion paths
Zeyue Xue, Guanglu Song, Qiushan Guo, Boxiao Liu, Zhuofan Zong, Yu Liu, and Ping Luo. Raphael: Text- to-image generation via large mixture of diffusion paths. Advances in Neural Information Processing Systems , 36,
-
[128]
Qwen2 technical report
An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, et al. Qwen2 technical report. arXiv preprint arXiv:2407.10671, 2024. 1
2024 arXiv
-
[129]
Xlnet: Generalized autoregressive pre- training for language understanding
Zhilin Yang. Xlnet: Generalized autoregressive pre- training for language understanding. arXiv preprint arXiv:1906.08237, 2019. 1
1906 arXiv
-
[130]
Vector-quantized image modeling with improved vqgan
Jiahui Yu, Xin Li, Jing Yu Koh, Han Zhang, Ruoming Pang, James Qin, Alexander Ku, Yuanzhong Xu, Jason Baldridge, and Yonghui Wu. Vector-quantized image modeling with improved vqgan. arXiv preprint arXiv:2110.04627, 2021. 1
2021 arXiv
-
[131]
Scaling autore- gressive models for content-rich text-to-image generation
Jiahui Yu, Yuanzhong Xu, Jing Yu Koh, Thang Luong, Gunjan Baid, Zirui Wang, Vijay Vasudevan, Alexander Ku, Yinfei Yang, Burcu Karagol Ayan, et al. Scaling autore- gressive models for content-rich text-to-image generation. arXiv preprint arXiv:2206.10789, 2(3):5, 2022. 1, 8, 4
2022 arXiv
-
[132]
Magvit: Masked generative video transformer
Lijun Yu, Yong Cheng, Kihyuk Sohn, Jos ´e Lezama, Han Zhang, Huiwen Chang, Alexander G Hauptmann, Ming- Hsuan Yang, Yuan Hao, Irfan Essa, et al. Magvit: Masked generative video transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p...
2023
-
[133]
Language model beats diffusion–tokenizer is key to visual generation
Lijun Yu, Jos ´e Lezama, Nitesh B Gundavarapu, Luca Ver- sari, Kihyuk Sohn, David Minnen, Yong Cheng, Vighnesh Birodkar, Agrim Gupta, Xiuye Gu, et al. Language model beats diffusion–tokenizer is key to visual generation. arXiv preprint arXiv:2310.05737, 2023. 1
-
[134]
Representation alignment for generation: Training diffu- sion transformers is easier than you think
Sihyun Yu, Sangkyung Kwak, Huiwon Jang, Jongheon Jeong, Jonathan Huang, Jinwoo Shin, and Saining Xie. Representation alignment for generation: Training diffu- sion transformers is easier than you think. arXiv preprint arXiv:2410.06940, 2024. 1
-
[135]
Root mean square layer normalization, 2019
Biao Zhang and Rico Sennrich. Root mean square layer normalization, 2019. 2
2019
-
[136]
Transfusion: Pre- dict the next token and diffuse images with one multi-modal model
Chunting Zhou, Lili Yu, Arun Babu, Kushal Tirumala, Michihiro Yasunaga, Leonid Shamis, Jacob Kahn, Xuezhe Ma, Luke Zettlemoyer, and Omer Levy. Transfusion: Pre- dict the next token and diffuse images with one multi-modal model. arXiv preprint arXiv:2408.11039, 2024. 4, 19
2024 arXiv
-
[137]
Lumina-next: Making lumina- t2x stronger and faster with next-dit
Le Zhuo, Ruoyi Du, Han Xiao, Yangguang Li, Dongyang Liu, Rongjie Huang, Wenze Liu, Lirui Zhao, Fu-Yun Wang, Zhanyu Ma, et al. Lumina-next: Making lumina- t2x stronger and faster with next-dit. arXiv preprint arXiv:2406.18583, 2024. 1, 4 22
2024 arXiv
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.