Pith. sign in

REVIEW 36 references

Stylized Structural Patterns for Improved Neural Network Pre-training

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2506.19465 v1 pith:FWIJWBBA submitted 2025-06-24 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords syntheticdatasetsdatamodelsrealimagesimprovedtrained
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Modern deep learning models in computer vision require large datasets of real images, which are difficult to curate and pose privacy and legal concerns, limiting their commercial use. Recent works suggest synthetic data as an alternative, yet models trained with it often underperform. This paper proposes a two-step approach to bridge this gap. First, we propose an improved neural fractal formulation through which we introduce a new class of synthetic data. Second, we propose reverse stylization, a technique that transfers visual features from a small, license-free set of real images onto synthetic datasets, enhancing their effectiveness. We analyze the domain gap between our synthetic datasets and real images using Kernel Inception Distance (KID) and show that our method achieves a significantly lower distributional gap compared to existing synthetic datasets. Furthermore, our experiments across different tasks demonstrate the practical impact of this reduced gap. We show that pretraining the EDM2 diffusion model on our synthetic dataset leads to an 11% reduction in FID during image generation, compared to models trained on existing synthetic datasets, and a 20% decrease in autoencoder reconstruction error, indicating improved performance in data representation. Furthermore, a ViT-S model trained for classification on this synthetic data achieves over a 10% improvement in ImageNet-100 accuracy. Our work opens up exciting possibilities for training practical models when sufficiently large real training sets are not available.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

36 extracted references · 24 canonical work pages

  1. [1]

    Autoencoder configuration file (kl-f4), 2023

    Stability AI. Autoencoder configuration file (kl-f4), 2023. Accessed: Nov. 15, 2024. 11

  2. [2]

    Improving fractal pre- training

    Connor Anderson and Ryan Farrell. Improving fractal pre- training. In Proceedings of the IEEE/CVF Winter Confer- ence on Applications of Computer Vision, pages 1300–1309,

  3. [3]

    Neuralfractal - a visual exploration of neu- ral dynamical systems

    Amirabbas Asadi. Neuralfractal - a visual exploration of neu- ral dynamical systems. https://amirabbasasadi. github.io/neural-fractal/, 2021. (Last accessed on September 2024). 2, 3, 4

  4. [4]

    Learning to see by looking at noise

    Manel Baradad, Jonas Wulff, Tongzhou Wang, Phillip Isola, and Antonio Torralba. Learning to see by looking at noise. In Advances in Neural Information Processing Systems , 2021. 1, 2, 4, 5

  5. [5]

    Procedural image programs for representation learning

    Manel Baradad, Richard Chen, Jonas Wulff, Tongzhou Wang, Rogerio Feris, Antonio Torralba, and Phillip Isola. Procedural image programs for representation learning. Ad- vances in Neural Information Processing Systems, 35:6450– 6462, 2022. 1, 2, 6

  6. [6]

    Demystifying mmd gans

    Mikołaj Bi ´nkowski, Danica J Sutherland, Michael Arbel, and Arthur Gretton. Demystifying mmd gans. arXiv preprint arXiv:1801.01401, 2018. 7

  7. [7]

    The design and evolution of disney’s hyperion renderer

    Brent Burley, Yining Karl Li, Felix Hecht, Bruce Meyer, and Matthew Hill. The design and evolution of disney’s hyperion renderer. ACM Transactions on Graphics (TOG) , 37(3):1– 22, 2018. 4

  8. [8]

    Im- age neural style transfer: A review

    Qiang Cai, Mengxu Ma, Chen Wang, and Haisheng Li. Im- age neural style transfer: A review. Computers and Electri- cal Engineering, 108:108723, 2023. 3

Show all 36 references
  1. [9]

    Emerg- ing properties in self-supervised vision transformers

    Mathilde Caron, Hugo Touvron, Ishan Misra, Herv ´e J´egou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerg- ing properties in self-supervised vision transformers. In Pro- ceedings of the IEEE/CVF international conference on com- puter vision, pages 9650–9660, 2021. 6, 12

  2. [10]

    Pre-training vision models with mandelbulb variations

    Benjamin Naoto Chiche, Yuto Horikawa, and Ryo Fujita. Pre-training vision models with mandelbulb variations. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition, pages 22062–22071, 2024. 4, 6

  3. [11]

    Imagenet: A large-scale hierarchical image database

    Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pages 248–255. Ieee, 2009. 1, 2, 5

  4. [12]

    An image is worth 16x16 words: Trans- formers for image recognition at scale

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Syl- vain Gelly, et al. An image is worth 16x16 words: Trans- formers for image recognition at scale. arXiv preprint a...

  5. [13]

    A neural algorithm of artistic style

    Leon A Gatys, Alexander S Ecker, and Matthias Bethge. A neural algorithm of artistic style. arXiv preprint arXiv:1508.06576, 2015. 3, 5

  6. [14]

    9 amazing fractals found in nature

    Shea Gunther. 9 amazing fractals found in nature. https: / / www . treehugger . com / amazing - fractals - found-in-nature-4868776 , 2024. (Last accessed on September 2024). 2

  7. [15]

    Diffusion- enhanced patchmatch: A framework for arbitrary style trans- fer with diffusion models

    Mark Hamazaspyan and Shant Navasardyan. Diffusion- enhanced patchmatch: A framework for arbitrary style trans- fer with diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 797–805, 2023. 3 9

  8. [16]

    Neural style transfer: A review

    Yongcheng Jing, Yezhou Yang, Zunlei Feng, Jingwen Ye, Yizhou Yu, and Mingli Song. Neural style transfer: A review. IEEE transactions on visualization and computer graphics , 26(11):3365–3385, 2019. 3, 9

  9. [17]

    Analyzing and improving the training dynamics of diffusion models

    Tero Karras, Miika Aittala, Jaakko Lehtinen, Janne Hellsten, Timo Aila, and Samuli Laine. Analyzing and improving the training dynamics of diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 24174–24184, 2024. 2, 6, 11, 12, 15

  10. [18]

    Pre-training without natural images

    Hirokatsu Kataoka, Kazushige Okayasu, Asato Matsumoto, Eisuke Yamagata, Ryosuke Yamada, Nakamasa Inoue, Akio Nakamura, and Yutaka Satoh. Pre-training without natural images. In Proceedings of the Asian Conference on Com- puter Vision, 2020. 1, 2, 5

  11. [19]

    Re- placing labeled real-image datasets with auto-generated con- tours

    Hirokatsu Kataoka, Ryo Hayamizu, Ryosuke Yamada, Kodai Nakashima, Sora Takashima, Xinyu Zhang, Edgar Josafat Martinez-Noriega, Nakamasa Inoue, and Rio Yokota. Re- placing labeled real-image datasets with auto-generated con- tours. In Proceedings of the IEEE/CVF Conference on C...

  12. [20]

    Neural neighbor style transfer

    Nicholas Kolkin, Michal Kucera, Sylvain Paris, Daniel Sykora, Eli Shechtman, and Greg Shakhnarovich. Neural neighbor style transfer. arXiv e-prints, pages arXiv–2203,

  13. [21]

    Mi- crosoft coco: Common objects in context

    Tsung-Yi Lin, Michael Maire, Serge Belongie, Lubomir Bourdev, Ross Girshick, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll ´ar, and C Lawrence Zitnick. Mi- crosoft coco: Common objects in context. arXiv preprint arXiv:1405.0312, 2014. 6

  14. [22]

    The fractal geometry of nature/revised and enlarged edition

    Benoit B Mandelbrot. The fractal geometry of nature/revised and enlarged edition. New York, 1983. 2, 4

  15. [23]

    Can vi- sion transformers learn without natural images? In Proceed- ings of the AAAI Conference on Artificial Intelligence, pages 1990–1998, 2022

    Kodai Nakashima, Hirokatsu Kataoka, Asato Matsumoto, Kenji Iwata, Nakamasa Inoue, and Yutaka Satoh. Can vi- sion transformers learn without natural images? In Proceed- ings of the AAAI Conference on Artificial Intelligence, pages 1990–1998, 2022. 2

  16. [24]

    High-resolution image synthesis with latent diffusion models

    Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022. 2, 3, 6, 15

  17. [25]

    Diff-nst: Diffusion interleaving for deformable neural style transfer

    Dan Ruta, Gemma Canet Tarr´es, Andrew Gilbert, Eli Shecht- man, Nicholas Kolkin, and John Collomosse. Diff-nst: Diffusion interleaving for deformable neural style transfer. arXiv preprint arXiv:2307.04157, 2023. 3

  18. [26]

    No training, no problem: Rethinking classifier-free guidance for diffusion models

    Seyedmorteza Sadat, Manuel Kansy, Otmar Hilliges, and Romann M Weber. No training, no problem: Rethinking classifier-free guidance for diffusion models. arXiv preprint arXiv:2407.02687, 2024. 6

  19. [27]

    Segrcdb: Seman- tic segmentation via formula-driven supervised learning

    Risa Shinoda, Ryo Hayamizu, Kodai Nakashima, Nakamasa Inoue, Rio Yokota, and Hirokatsu Kataoka. Segrcdb: Seman- tic segmentation via formula-driven supervised learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 20054–20063, 2023. 2

  20. [28]

    Visual atoms: Pre-training vision transformers with sinusoidal waves

    Sora Takashima, Ryo Hayamizu, Nakamasa Inoue, Hi- rokatsu Kataoka, and Rio Yokota. Visual atoms: Pre-training vision transformers with sinusoidal waves. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18579–18588, 2023. 2, 4, 6, 16

  21. [29]

    Common techniques for generating fractals, 2024

    Wikipedia contributors. Common techniques for generating fractals, 2024. (Last accessed on September 2024). 2

  22. [30]

    mixup: Beyond empirical risk minimization

    Hongyi Zhang. mixup: Beyond empirical risk minimization. arXiv preprint arXiv:1710.09412, 2017. 8

  23. [31]

    The unreasonable effectiveness of deep features as a perceptual metric

    Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shecht- man, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE conference on computer vision and pattern recogni- tion, pages 586–595, 2018. 6

  24. [32]

    Inversion-based style transfer with diffusion models

    Yuxin Zhang, Nisha Huang, Fan Tang, Haibin Huang, Chongyang Ma, Weiming Dong, and Changsheng Xu. Inversion-based style transfer with diffusion models. In Pro- ceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10146–10156, 2023. 3

  25. [33]

    Places: A 10 million image database for scene recognition

    Bolei Zhou, Agata Lapedriza, Aditya Khosla, Aude Oliva, and Antonio Torralba. Places: A 10 million image database for scene recognition. IEEE transactions on pattern analysis and machine intelligence, 40(6):1452–1464, 2017. 2 10 Stylized Structural Patterns for Improved Neural...

  26. [34]

    Neural Fractal Generation In this section, we provide the pseudocode for the coloring and adaptive sampling algorithms. Algorithm 1: Dynamic Threshold Adjustment Input: z (first pass data from the rendering), ratio (desired proportion), τinit (initial threshold) Initialization...

  27. [35]

    The latent space of AE has 4 channels as in [1]

    Hyperparameters AutoEncoder: We train the stable diffusion autoencoder (AE) following the hyperparameters in [1]. The latent space of AE has 4 channels as in [1]. We train the AE with images of resolution 128 × 128 and a batch size of 8. We use the Adam optimizer with the lear...

  28. [36]

    Additional Results Impact of Network Architecture on Neural Fractal: We experiment with networks of varying complexity for neu- ral fractal generation. Increasing the complexity of the net- work tends to increase the frequency of the output image (see Figure 8), making renderi...

Pith tools