REVIEW 36 references
Stylized Structural Patterns for Improved Neural Network Pre-training
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Modern deep learning models in computer vision require large datasets of real images, which are difficult to curate and pose privacy and legal concerns, limiting their commercial use. Recent works suggest synthetic data as an alternative, yet models trained with it often underperform. This paper proposes a two-step approach to bridge this gap. First, we propose an improved neural fractal formulation through which we introduce a new class of synthetic data. Second, we propose reverse stylization, a technique that transfers visual features from a small, license-free set of real images onto synthetic datasets, enhancing their effectiveness. We analyze the domain gap between our synthetic datasets and real images using Kernel Inception Distance (KID) and show that our method achieves a significantly lower distributional gap compared to existing synthetic datasets. Furthermore, our experiments across different tasks demonstrate the practical impact of this reduced gap. We show that pretraining the EDM2 diffusion model on our synthetic dataset leads to an 11% reduction in FID during image generation, compared to models trained on existing synthetic datasets, and a 20% decrease in autoencoder reconstruction error, indicating improved performance in data representation. Furthermore, a ViT-S model trained for classification on this synthetic data achieves over a 10% improvement in ImageNet-100 accuracy. Our work opens up exciting possibilities for training practical models when sufficiently large real training sets are not available.
Reference graph
Works this paper leans on
-
[1]
Autoencoder configuration file (kl-f4), 2023
Stability AI. Autoencoder configuration file (kl-f4), 2023. Accessed: Nov. 15, 2024. 11
2023
-
[2]
Improving fractal pre- training
Connor Anderson and Ryan Farrell. Improving fractal pre- training. In Proceedings of the IEEE/CVF Winter Confer- ence on Applications of Computer Vision, pages 1300–1309,
-
[3]
Neuralfractal - a visual exploration of neu- ral dynamical systems
Amirabbas Asadi. Neuralfractal - a visual exploration of neu- ral dynamical systems. https://amirabbasasadi. github.io/neural-fractal/, 2021. (Last accessed on September 2024). 2, 3, 4
2021
-
[4]
Learning to see by looking at noise
Manel Baradad, Jonas Wulff, Tongzhou Wang, Phillip Isola, and Antonio Torralba. Learning to see by looking at noise. In Advances in Neural Information Processing Systems , 2021. 1, 2, 4, 5
2021
-
[5]
Procedural image programs for representation learning
Manel Baradad, Richard Chen, Jonas Wulff, Tongzhou Wang, Rogerio Feris, Antonio Torralba, and Phillip Isola. Procedural image programs for representation learning. Ad- vances in Neural Information Processing Systems, 35:6450– 6462, 2022. 1, 2, 6
work page 2022
-
[6]
Mikołaj Bi ´nkowski, Danica J Sutherland, Michael Arbel, and Arthur Gretton. Demystifying mmd gans. arXiv preprint arXiv:1801.01401, 2018. 7
arXiv 2018
-
[7]
The design and evolution of disney’s hyperion renderer
Brent Burley, Yining Karl Li, Felix Hecht, Bruce Meyer, and Matthew Hill. The design and evolution of disney’s hyperion renderer. ACM Transactions on Graphics (TOG) , 37(3):1– 22, 2018. 4
work page 2018
-
[8]
Im- age neural style transfer: A review
Qiang Cai, Mengxu Ma, Chen Wang, and Haisheng Li. Im- age neural style transfer: A review. Computers and Electri- cal Engineering, 108:108723, 2023. 3
work page 2023
Show all 36 references
-
[9]
Emerg- ing properties in self-supervised vision transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Herv ´e J´egou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerg- ing properties in self-supervised vision transformers. In Pro- ceedings of the IEEE/CVF international conference on com- puter vision, pages 9650–9660, 2021. 6, 12
2021
-
[10]
Pre-training vision models with mandelbulb variations
Benjamin Naoto Chiche, Yuto Horikawa, and Ryo Fujita. Pre-training vision models with mandelbulb variations. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition, pages 22062–22071, 2024. 4, 6
2024
-
[11]
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pages 248–255. Ieee, 2009. 1, 2, 5
2009
-
[12]
An image is worth 16x16 words: Trans- formers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Syl- vain Gelly, et al. An image is worth 16x16 words: Trans- formers for image recognition at scale. arXiv preprint a...
2010 arXiv
-
[13]
A neural algorithm of artistic style
Leon A Gatys, Alexander S Ecker, and Matthias Bethge. A neural algorithm of artistic style. arXiv preprint arXiv:1508.06576, 2015. 3, 5
2015 arXiv
-
[14]
9 amazing fractals found in nature
Shea Gunther. 9 amazing fractals found in nature. https: / / www . treehugger . com / amazing - fractals - found-in-nature-4868776 , 2024. (Last accessed on September 2024). 2
2024
-
[15]
Diffusion- enhanced patchmatch: A framework for arbitrary style trans- fer with diffusion models
Mark Hamazaspyan and Shant Navasardyan. Diffusion- enhanced patchmatch: A framework for arbitrary style trans- fer with diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 797–805, 2023. 3 9
2023
-
[16]
Neural style transfer: A review
Yongcheng Jing, Yezhou Yang, Zunlei Feng, Jingwen Ye, Yizhou Yu, and Mingli Song. Neural style transfer: A review. IEEE transactions on visualization and computer graphics , 26(11):3365–3385, 2019. 3, 9
2019
-
[17]
Analyzing and improving the training dynamics of diffusion models
Tero Karras, Miika Aittala, Jaakko Lehtinen, Janne Hellsten, Timo Aila, and Samuli Laine. Analyzing and improving the training dynamics of diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 24174–24184, 2024. 2, 6, 11, 12, 15
2024
-
[18]
Pre-training without natural images
Hirokatsu Kataoka, Kazushige Okayasu, Asato Matsumoto, Eisuke Yamagata, Ryosuke Yamada, Nakamasa Inoue, Akio Nakamura, and Yutaka Satoh. Pre-training without natural images. In Proceedings of the Asian Conference on Com- puter Vision, 2020. 1, 2, 5
2020
-
[19]
Re- placing labeled real-image datasets with auto-generated con- tours
Hirokatsu Kataoka, Ryo Hayamizu, Ryosuke Yamada, Kodai Nakashima, Sora Takashima, Xinyu Zhang, Edgar Josafat Martinez-Noriega, Nakamasa Inoue, and Rio Yokota. Re- placing labeled real-image datasets with auto-generated con- tours. In Proceedings of the IEEE/CVF Conference on C...
-
[20]
Neural neighbor style transfer
Nicholas Kolkin, Michal Kucera, Sylvain Paris, Daniel Sykora, Eli Shechtman, and Greg Shakhnarovich. Neural neighbor style transfer. arXiv e-prints, pages arXiv–2203,
-
[21]
Mi- crosoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, Lubomir Bourdev, Ross Girshick, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll ´ar, and C Lawrence Zitnick. Mi- crosoft coco: Common objects in context. arXiv preprint arXiv:1405.0312, 2014. 6
2014 arXiv
-
[22]
The fractal geometry of nature/revised and enlarged edition
Benoit B Mandelbrot. The fractal geometry of nature/revised and enlarged edition. New York, 1983. 2, 4
1983
-
[23]
Can vi- sion transformers learn without natural images? In Proceed- ings of the AAAI Conference on Artificial Intelligence, pages 1990–1998, 2022
Kodai Nakashima, Hirokatsu Kataoka, Asato Matsumoto, Kenji Iwata, Nakamasa Inoue, and Yutaka Satoh. Can vi- sion transformers learn without natural images? In Proceed- ings of the AAAI Conference on Artificial Intelligence, pages 1990–1998, 2022. 2
1990
-
[24]
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022. 2, 3, 6, 15
2022
-
[25]
Diff-nst: Diffusion interleaving for deformable neural style transfer
Dan Ruta, Gemma Canet Tarr´es, Andrew Gilbert, Eli Shecht- man, Nicholas Kolkin, and John Collomosse. Diff-nst: Diffusion interleaving for deformable neural style transfer. arXiv preprint arXiv:2307.04157, 2023. 3
2023 arXiv
-
[26]
No training, no problem: Rethinking classifier-free guidance for diffusion models
Seyedmorteza Sadat, Manuel Kansy, Otmar Hilliges, and Romann M Weber. No training, no problem: Rethinking classifier-free guidance for diffusion models. arXiv preprint arXiv:2407.02687, 2024. 6
2024 arXiv
-
[27]
Segrcdb: Seman- tic segmentation via formula-driven supervised learning
Risa Shinoda, Ryo Hayamizu, Kodai Nakashima, Nakamasa Inoue, Rio Yokota, and Hirokatsu Kataoka. Segrcdb: Seman- tic segmentation via formula-driven supervised learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 20054–20063, 2023. 2
2023
-
[28]
Visual atoms: Pre-training vision transformers with sinusoidal waves
Sora Takashima, Ryo Hayamizu, Nakamasa Inoue, Hi- rokatsu Kataoka, and Rio Yokota. Visual atoms: Pre-training vision transformers with sinusoidal waves. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18579–18588, 2023. 2, 4, 6, 16
2023
-
[29]
Common techniques for generating fractals, 2024
Wikipedia contributors. Common techniques for generating fractals, 2024. (Last accessed on September 2024). 2
2024
-
[30]
mixup: Beyond empirical risk minimization
Hongyi Zhang. mixup: Beyond empirical risk minimization. arXiv preprint arXiv:1710.09412, 2017. 8
2017 arXiv
-
[31]
The unreasonable effectiveness of deep features as a perceptual metric
Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shecht- man, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE conference on computer vision and pattern recogni- tion, pages 586–595, 2018. 6
2018
-
[32]
Inversion-based style transfer with diffusion models
Yuxin Zhang, Nisha Huang, Fan Tang, Haibin Huang, Chongyang Ma, Weiming Dong, and Changsheng Xu. Inversion-based style transfer with diffusion models. In Pro- ceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10146–10156, 2023. 3
2023
-
[33]
Places: A 10 million image database for scene recognition
Bolei Zhou, Agata Lapedriza, Aditya Khosla, Aude Oliva, and Antonio Torralba. Places: A 10 million image database for scene recognition. IEEE transactions on pattern analysis and machine intelligence, 40(6):1452–1464, 2017. 2 10 Stylized Structural Patterns for Improved Neural...
2017
-
[34]
Neural Fractal Generation In this section, we provide the pseudocode for the coloring and adaptive sampling algorithms. Algorithm 1: Dynamic Threshold Adjustment Input: z (first pass data from the rendering), ratio (desired proportion), τinit (initial threshold) Initialization...
-
[35]
The latent space of AE has 4 channels as in [1]
Hyperparameters AutoEncoder: We train the stable diffusion autoencoder (AE) following the hyperparameters in [1]. The latent space of AE has 4 channels as in [1]. We train the AE with images of resolution 128 × 128 and a batch size of 8. We use the Adam optimizer with the lear...
-
[36]
Additional Results Impact of Network Architecture on Neural Fractal: We experiment with networks of varying complexity for neu- ral fractal generation. Increasing the complexity of the net- work tends to increase the frequency of the output image (see Figure 8), making renderi...
Discussion (0). Continue with ORCID to comment.