Pith. sign in

REVIEW 3 major objections 9 minor 77 references

Spiral-sampled GAN halves handwriting synthesis error rate

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · glm-5.2

2026-07-09 22:18 UTC pith:3CHJXHKB

load-bearing objection Solid engineering, but the headline FID gains don't survive on the paper's own preferred metric the 3 major comments →

arxiv 2607.06949 v1 pith:3CHJXHKB submitted 2026-07-08 cs.CV

SpiS-GAN: Spiral-Modulated Handwriting Synthesis with Star Operation

classification cs.CV
keywords handwriting synthesisgenerative adversarial networksspiral convolutionstar operationedge lossstyle transferdata augmentationtext recognition
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper claims that handwriting synthesis quality improves substantially when the generator samples features along elliptical spiral paths—matching the natural left-to-right, curving flow of pen strokes—rather than along rigid grid axes, and when the discriminator evaluates images through parallel spatial, spiral, and frequency pathways that preserve fine structural detail without aggressive downsampling. Combined with a Sobel edge loss that penalizes blurred stroke boundaries, this architecture produces synthetic handwriting that is more realistic, more legible, and more useful for training recognition systems than prior approaches, as demonstrated on English and Vietnamese benchmarks.

Core claim

The central claim is that handwriting generation can be improved by replacing grid-based feature sampling with a modulated elliptical spiral pattern that follows the natural trajectory of handwriting strokes, pairing it with an element-wise star multiplication that implicitly expands features into a high-dimensional space at low computational cost, and supervising stroke edges explicitly with a Sobel gradient loss. Together these changes cut the Fréchet Inception Distance on the IAM benchmark from 6.73 to 4.37 at 32-pixel resolution and from 10.17 to 4.58 at 64-pixel resolution, while also reducing downstream recognition error rates.

What carries the argument

The paper introduces three components: (1) Modulated Elliptical SpiralFC, which replaces the circular sampling path of SpiralMLP with an elliptical trajectory whose horizontal extent exceeds its vertical extent, with a triangular modulation function that expands and contracts the sampling radius within each channel group; (2) the Star-Spiral Block, which fuses spiral-sampled, spectrally gated, and spatially projected feature branches through an element-wise star product that implicitly maps features into a high-dimensional nonlinear space without widening channels; and (3) a Spiral-Modulated Discriminator that evaluates generated images through three parallel pathways—spatial, spiral, and频率—

Load-bearing premise

The paper assumes that the specific elliptical spiral sampling pattern is what captures handwriting stroke trajectories better than alternatives, but the ablation tests the entire Star-Spiral Block as a unit, so no isolated evidence shows the elliptical shape itself is the load-bearing factor rather than the star operation, spectral gating, or cross-depthwise convolution.

What would settle it

If the performance gain over the baseline persists when the elliptical spiral is replaced with a circular spiral, a random offset pattern, or a standard grid-based deformable convolution—while keeping the star operation, spectral gating, and edge loss—then the specific elliptical shape is not the mechanism driving the improvement.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • If the elliptical spiral sampling pattern genuinely captures stroke trajectories better than grid-based or circular alternatives, the same principle could transfer to other tasks involving curvilinear or flow-based structures, such as signature verification, sketch generation, or anatomical vessel tracing.
  • The Sobel edge loss provides a general mechanism for enforcing high-frequency detail in any generative model that tends to over-smooth outputs, and could be adopted in document enhancement or text-to-image generation beyond handwriting.
  • If the star operation's implicit high-dimensional expansion is effective in a GAN generator, it may offer a lightweight alternative to attention mechanisms or wider channels in other conditional generation tasks where computational budget is constrained.
  • The demonstrated cross-language transfer to Vietnamese—with its stacked diacritics and compound vowels—suggests the architecture is robust to scripts with complex spatial layouts, which could extend to Arabic, Hindi, or CJK handwriting.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The ablation study bundles MESpiralFC, the star operation, spectral gating, and cross-depthwise convolution into a single Star-Spiral Block, so the isolated contribution of the elliptical shape versus the star multiplication or spectral branch remains unclear; a finer-grained ablation would be needed to confirm which sub-component carries the performance gain.
  • The claim that elliptical sampling matches natural writing direction is intuitive but not directly validated through visualization of learned sampling patterns or comparison against alternative non-grid shapes such as parabolic or bezier-curve sampling.
  • The improvement on Vietnamese may stem partly from the edge loss preserving diacritical marks rather than from the spiral sampling itself, since diacritics are high-frequency spatial features that benefit from explicit edge supervision regardless of the generator's receptive field geometry.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 9 minor

Summary. The paper proposes SpiS-GAN, a GAN-based framework for one-shot offline handwriting synthesis. The generator replaces standard BigGAN blocks with Star-Spiral Blocks (SSBs) that combine a Modulated Elliptical SpiralFC (MESpiralFC) sampling pattern, a spectral gating branch, and the StarNet element-wise multiplication for implicit high-dimensional feature expansion. A dual-discriminator setup pairs a conventional CNN discriminator with a Spiral-Modulated discriminator using three parallel pathways (spatial, spiral, spectral). A Sobel-Regularized Edge Reconstruction Loss (SELoss) is added to enforce stroke-boundary sharpness. Experiments on IAM (English) and HANDS-VNOnDB (Vietnamese) report FID/KID improvements over FW-GAN, HiGAN+, One-DM, and others, and downstream HTR gains when synthetic data augments a 5,000-image real-data baseline.

Significance. The paper addresses a genuine problem: existing MLP-based and CNN-based handwriting generators struggle with cursive stroke trajectories and fine edge detail. The combination of spiral-offset sampling with the star operation is architecturally novel for this domain, and the Sobel edge loss is a reasonable, lightweight addition. The code is stated to be available, and the evaluation protocol follows the reproducible FW-GAN framework across two languages. These are positive factors. However, the significance of the reported gains is tempered by the metric-reliability issue detailed below.

major comments (3)
  1. §4.1 and Tables 6–7: The paper explicitly states that FID and KID have 'inherent flaws when evaluating handwriting' due to ImageNet domain mismatch, and introduces HWD to compensate. Yet the headline 'significantly outperforms' claim rests almost entirely on FID/KID. On the domain-appropriate HWD metric, the gains shrink dramatically: at 64px IAM (Table 6), HWD is 0.31 (Ours) vs 0.33 (HiGAN+) — a 0.02 gap; at 64px Vietnamese (Table 7), HWD is identical at 0.20 for HiGAN+ and Ours. At 32px, the HWD gap is 0.06 on both datasets. Without error bars, confidence intervals, or multi-seed runs, it is unclear whether these HWD differences are statistically meaningful. The authors should either (a) soften the 'significantly outperforms' language to reflect the HWD picture, or (b) provide multi-seed runs with significance tests on HWD to substantiate the claim. As it stands, the central claim is 1
  2. Tables 3, 6, 7: No confidence intervals, standard deviations, or significance tests are reported for any metric. The downstream HTR improvements (Table 3) are small in absolute terms: 0.15 CER points at 32px (10.03 vs 10.18 for FW-GAN), 0.44 at 64px (9.01 vs 9.45 for HiGAN+). Given that these are single-run results, the practical significance of these gains is uncertain. Adding at least 3-seed runs with mean ± std for the primary metrics (FID, HWD, CER) would substantially strengthen the contribution, particularly where gaps are small.
  3. Table 4 (Ablation): The ablation conflates multiple modifications within each row. Row B adds Star-Spiral Blocks, but each SSB simultaneously introduces MESpiralFC, the star operation, spectral gating, and cross-depthwise convolution (§3.2.2). The paper lists MESpiralFC and the star operation as core contributions, but the ablation provides no isolated evidence that the elliptical spiral shape (Eq. 5–7) or the star operation individually is load-bearing. A finer-grained ablation — e.g., replacing MESpiralFC with original SpiralFC, or removing the star multiplication — would clarify which components drive the FID drop from 11.37 to 5.91. Without this, the claim that the elliptical spiral 'better matches the natural writing direction' (§3.2.1) is supported only by intuition, not by isolated evidence.
minor comments (9)
  1. §2.2, final paragraph: 'a dulated Discriminator' appears to be a typo for 'Spiral-Modulated Discriminator'.
  2. §3.2.1, Eqs. 5–7: The motivation for the triangular modulation function R(i) is stated briefly. A sentence explaining why triangular (vs. Gaussian, linear, etc.) was chosen, or whether this was empirically tuned, would improve readability.
  3. §3.2.2, Eq. 11: The activation σ is applied to one branch but the specific function is not named. Clarify which activation is used (ReLU? GELU?).
  4. §3.3.7, Eq. 24: 'where where' is duplicated.
  5. §4.2: The loss weights λ_R, λ_W, λ_style are said to be dynamically calibrated via a 'specialized gradient balancing algorithm' but no reference or formula is given. A citation or brief description is needed for reproducibility.
  6. Table 2: The 64px OOV-U entry for Ours (29.40) is lower than the IV-S entry (31.27), which is counterintuitive (unseen words + unseen styles should be harder than seen words + seen styles). The authors should confirm this is correct or add discussion.
  7. Figure references: Figures 1–8 appear in sequence but Figure 9 (ablation visual) is referenced before Figures 10–12. Minor, but check ordering for narrative flow.
  8. §4.9, Table 8: The model size comparison excludes the discriminator parameters. Since the Spiral-Modulated Discriminator is a core contribution, including its parameter count would strengthen the deployment-efficiency argument.
  9. §3.2.3, Eq. 13: The gating weights a1, a2, a3 are described as independent, but Eq. 12 shows the context is computed from only two branches (excluding the spectral branch). A brief justification for why the spectral gate does not receive spectral context would help.

Circularity Check

0 steps flagged

No significant circularity found; the paper's contributions are architectural design choices validated empirically, with no derivation step that reduces to its own inputs by construction.

full rationale

The paper proposes SpiS-GAN, whose core contributions are architectural components (MESpiralFC, Star-Spiral Blocks, Spiral-Modulated Discriminator, Sobel Edge Loss) supported by empirical evaluation. There is no mathematical derivation chain where a 'prediction' or 'first-principles result' could reduce to its inputs by construction. The star operation's high-dimensional expansion property (Eqs. 8-10) is cited from StarNet [45] by external authors (Ma et al.), which is independently verifiable. The evaluation protocol, FDL, and style encoder design are adopted from FW-GAN [35], which shares two co-authors, but these are methodological choices for fair comparison—not mathematical claims whose validity depends on the cited work being true. The FDL itself traces to [62] (external). The Sobel Edge Loss (Eq. 24) is a straightforward L1 loss on fixed-filter edge maps. The ablation (Table 4) conflates multiple modifications within the Star-Spiral Block, but this is an ablation design weakness, not circularity: no component is defined in terms of the metric it then 'predicts.' The skeptic's concern about marginal HWD gains is a correctness/evaluation issue, not a circularity issue. The minor self-citation to FW-GAN for methodology is not load-bearing for the novel contributions. Score 1 reflects this minor self-citation that does not undermine the independence of the central claims.

Axiom & Free-Parameter Ledger

7 free parameters · 4 axioms · 3 invented entities

The paper introduces three new architectural/loss components, each with some empirical support via ablation, though none is isolated from confounding modifications.

free parameters (7)
  • λ_KL = 0.0001
    KL divergence loss weight, manually set (§4.2)
  • λ_FDL = 1
    Frequency distribution loss weight, manually set (§4.2)
  • λ_edge = 1
    Sobel edge loss weight, manually set (§4.2)
  • λ_R, λ_W, λ_style = dynamically calibrated
    Remaining loss weights adjusted by unspecified gradient balancing algorithm (§4.2)
  • R_x, R_y = relative to C_in, R_x > R_y
    Horizontal and vertical extents of elliptical sampling pattern; ratio chosen to match left-to-right writing flow (§3.2.1)
  • T = inherited from SpiralMLP
    Spiral rotation period, follows original SpiralFC formulation (§3.2.1)
  • k (channel groups) = unspecified
    Number of channel groups for elliptical spiral partitioning, follows SpiralMLP but exact value not stated (§3.2.1)
axioms (4)
  • domain assumption Elliptical spiral sampling patterns better match handwriting stroke trajectories than circular or grid-based patterns
    Stated in §3.2.1: 'Setting R_x > R_y stretches the pattern horizontally, thereby creating elliptical patterns to better match the natural writing direction of human handwriting.' Supported empirically but not derived from first principles.
  • domain assumption The star operation's implicit high-dimensional expansion (from StarNet [45]) is beneficial for handwriting synthesis feature interactions
    Invoked in §3.2.2 to justify using element-wise multiplication for 'rich higher-order feature modeling' in the generator. Theoretical grounding comes from [45], not from this paper.
  • domain assumption Sobel edge magnitude maps provide sufficient supervision for stroke boundary preservation
    §3.3.7 assumes L1 distance between Sobel edge maps of real and generated images is an appropriate loss. The choice of Sobel over Laplacian (used in DocDiff [65]) is justified only by experiments, not theory.
  • domain assumption FID and KID, despite known domain mismatch with handwriting, remain valid comparison metrics
    §4.1 acknowledges Inception network domain mismatch but retains FID/KID 'to maintain consistency with historical literature.' HWD is added as complement.
invented entities (3)
  • Modulated Elliptical SpiralFC (MESpiralFC) independent evidence
    purpose: Deformable convolution-style layer with elliptical spiral offsets for stroke trajectory sampling
    Ablation in Table 4 shows FID improvement when Star-Spiral Blocks (containing MESpiralFC) replace BigGAN blocks, but MESpiralFC is not isolated from other block components.
  • Spiral-Modulated Discriminator (D_SP) independent evidence
    purpose: Three-pathway discriminator with spatial, spiral, and spectral branches for multi-domain flaw detection
    Ablation step C in Table 4 shows FID improvement from 5.91 to 4.61 when this discriminator is added.
  • Sobel-Regularized Edge Reconstruction Loss (SELoss) independent evidence
    purpose: Edge-aware loss using directional Sobel gradients to enforce sharp stroke boundaries
    Ablation step D in Table 4 shows FID improvement from 4.61 to 4.37 when SELoss is added.

pith-pipeline@v1.1.0-glm · 30136 in / 4166 out tokens · 169062 ms · 2026-07-09T22:18:02.757849+00:00 · methodology

0 comments
read the original abstract

Training robust handwriting recognition (HTR) systems requires massive amounts of annotated data, which is often difficult to acquire. While synthetic handwriting generation offers a practical solution to expand training sets, existing models struggle with several core issues. First, previous approaches, even MLP-based models fail to effectively trace cursive handwriting due to fixed-grid spatial receptive field. Second, their CNN-relied discriminators usually lose structural details through aggressive downsampling, making broken connections difficult to detect. Third, existing architectures are either limited to linear feature interactions or too expensive for high-resolution synthesis. Finally, existing approaches lack explicit edge constraints, often resulting in blurred stroke boundaries. To address these challenges, this study proposes a Spiral-Modulated Handwriting Synthesis framework based on Generative Adversarial Networks (SpiS-GAN). Our generator employs Star-Spiral Blocks combining proposed Modulated Elliptical SpiralFC with the star operation to capture spatial relationships and efficiently follow complex handwriting stroke trajectories, while a Spiral-Modulated discriminator is introduced for multi-domain flaws detection. Additionally, we introduce a Sobel-Regularized Edge Reconstruction Loss that provides edge guidance, ensuring every character remains clear and legible. Evaluations on the English and Vietnamese datasets demonstrate that SpiS-GAN significantly outperforms current state-of-the-art models. The generated images are highly authentic, accurately preserve the original writer's style across languages, and successfully lower error rates when training downstream HTR systems.

Figures

Figures reproduced from arXiv: 2607.06949 by Dang Hoai Nam, Nguyen Duy Hieu, Pham Hoang Giap, Quang Huu Hieu, Vo Nguyen Le Duy.

Figure 1
Figure 1. Figure 1: Overview of the SpiS-GAN architecture. The input text is converted to one-hot [PITH_FULL_IMAGE:figures/full_fig_p009_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: (a)–(b) Grid-based; (c) Original SpiralFC; (d) Elliptical SpiralFC (ours). [PITH_FULL_IMAGE:figures/full_fig_p010_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Overview of the SpiS-GAN hierarchical generator featuring Star-Spiral Blocks [PITH_FULL_IMAGE:figures/full_fig_p011_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Detailed architecture of Star-Spiral Block [PITH_FULL_IMAGE:figures/full_fig_p012_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Architecture of the original StarBlock from StarNet [45]. [PITH_FULL_IMAGE:figures/full_fig_p013_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Illustration of the Dual Discriminator Architecture. [PITH_FULL_IMAGE:figures/full_fig_p014_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Detailed architecture of SP-MLP Block. As the network goes deeper, the spatial resolution gradually decreases while the number of feature channels increases, enabling the model to extract fea￾tures at multiple levels, from fine stroke textures and individual character shapes to broader structural patterns. The final feature map is then con￾densed through global average pooling and passed through a linear l… view at source ↗
Figure 8
Figure 8. Figure 8: Visualization of Sobel edge magnitude maps. [PITH_FULL_IMAGE:figures/full_fig_p022_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Visual progression of the ablation study. Red boxes indicate faint/blurred edges; [PITH_FULL_IMAGE:figures/full_fig_p035_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Qualitative comparison of handwriting generated from the IAM dataset across [PITH_FULL_IMAGE:figures/full_fig_p040_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: Qualitative comparison of handwriting generated from the HANDS-VNOnDB [PITH_FULL_IMAGE:figures/full_fig_p040_11.png] view at source ↗
Figure 12
Figure 12. Figure 12: Qualitative reconstruction results. Each row corresponds to a different model, [PITH_FULL_IMAGE:figures/full_fig_p042_12.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

77 extracted references · 77 canonical work pages · 19 internal anchors

  1. [1]

    H. T. Nguyen, C. T. Nguyen, P. T. Bao, M. Nakagawa, A database of unconstrained vietnamese online handwriting and recognition exper- iments by recurrent neural networks, Pattern Recognition 78 (2018) 291–306.doi:https://doi.org/10.1016/j.patcog.2018.01.013. URLhttps://www.sciencedirect.com/science/article/pii/ S0031320318300141

  2. [2]

    Kleber, S

    F. Kleber, S. Fiel, M. Diem, R. Sablatnig, CVL-DataBase: An Off-Line Database for Writer Retrieval, Writer Identification and Word Spotting, in: ICDAR, 2013

  3. [3]

    Pratikakis, K

    I. Pratikakis, K. Zagori, P. Kaddas, B. Gatos, Icfhr 2018 competition on handwritten document image binarization (h-dibco 2018), in: 2018 16th International Conference on Frontiers in Handwriting Recognition (ICFHR), 2018, pp. 489–493.doi:10.1109/ICFHR-2018.2018.00091

  4. [4]

    R. D. Lins, Nabuco - two decades of document processing in latin amer- ica, J. Univers. Comput. Sci. 17 (2011) 151–161. URLhttps://api.semanticscholar.org/CorpusID:2896293

  5. [5]

    Generating Sequences With Recurrent Neural Networks

    A. Graves, Generating Sequences with Recurrent Neural Networks, arXiv preprint arXiv:1308.0850 (2013)

  6. [6]

    A. K. Bhunia, S. Khan, H. Cholakkal, R. M. Anwer, F. S. Khan, M. Shah, Handwriting Transformers, in: ICCV, 2021

  7. [7]

    Aksan, F

    E. Aksan, F. Pece, O. Hilliges, DeepWriting: Making digital ink editable via deep generative modeling, in: CHI, ACM, 2018

  8. [8]

    Aksan, O

    E. Aksan, O. Hilliges, STCN: Stochastic Temporal Convolutional Net- works, in: ICLR, 2018. 43

  9. [9]

    Kotani, S

    A. Kotani, S. Tellex, J. Tompkin, Generating Handwriting via Decou- pled Style Descriptors, in: ECCV, 2020

  10. [10]

    B. Ji, T. Chen, Generative Adversarial Network for Handwritten Text, arXiv preprint arXiv:1907.11845 (2019)

  11. [11]

    J. Wang, C. Wu, Y.-Q. Xu, H.-Y. Shum, Combining Shape and Physical Models for On-line Cursive Handwriting Synthesis, IJDAR 7 (4) (2005) 219–227

  12. [12]

    Z. Lin, L. Wan, Style-preserving english handwriting synthesis, Pattern Recognit. 40 (7) (2007) 2097–2109

  13. [13]

    A. O. Thomas, A. Rusu, V. Govindaraju, Synthetic Handwritten CAPTCHAs, Pattern Recognit. 42 (12) (2009) 3365–3373

  14. [14]

    Haines, O

    T. Haines, O. Mac Aodha, G. Brostow, My Text in Your Handwriting, ACM Trans. Graphics 35 (3) (2016)

  15. [15]

    Alonso, B

    E. Alonso, B. Moysset, R. Messina, Adversarial Generation of Hand- written Text Images Conditioned on Sequences, in: ICDAR, IEEE Computer Society, 2019

  16. [16]

    Fogel, H

    S. Fogel, H. A verbuch-Elor, S. Cohen, S. Mazor, R. Litman, Scrabble- GAN: Semi-Supervised Varying Length Handwritten Text Generation, in: CVPR, 2020

  17. [17]

    L. Kang, P. Riba, Y. Wang, M. Rusi˜ nol, A. Fornés, M. Villegas, GAN- writing: Content-Conditioned Generation of Styled Handwritten Word Images, in: ECCV, 2020

  18. [18]

    J. Gan, W. Wang, HiGAN: Handwriting Imitation Conditioned on Arbitrary-Length Texts and Disentangled Styles, in: AAAI, 2021

  19. [19]

    Mattick, M

    A. Mattick, M. Mayr, M. Seuret, A. Maier, V. Christlein, SmartPatch: Improving Handwritten Word Imitation with Patch Discriminators, in: ICDAR, 2021

  20. [20]

    C. Luo, Y. Zhu, L. Jin, Z. Li, D. Peng, SLOGAN: Handwriting Style Synthesis for Arbitrary-Length and Out-of-Vocabulary Text, IEEE Trans. Neural Netw. Learn. Syst. (2022)

  21. [21]

    Krishnan, R

    P. Krishnan, R. Kovvuri, G. Pang, B. Vassilev, T. Hassner, TextStyle- Brush: Transfer of Text Aesthetics from a Single Example, arXiv e- prints (2021) arXiv–2106. 44

  22. [22]

    Davis, C

    B. Davis, C. Tensmeyer, B. Price, C. Wigington, B. Morse, R. Jain, Text and Style Conditioned GAN for Generation of Offline Handwriting Lines, in: BMVC, 2020

  23. [23]

    Brock, J

    A. Brock, J. Donahue, K. Simonyan, Large Scale GAN Training for High Fidelity Natural Image Synthesis, in: ICLR, 2019

  24. [24]

    Dhariwal, A

    P. Dhariwal, A. Nichol, Diffusion models beat gans on image synthesis, in: Proceedings of the 35th International Conference on Neural In- formation Processing Systems, NIPS ’21, Curran Associates Inc., Red Hook, NY, USA, 2021

  25. [25]

    Z. Wang, H. Zheng, P. He, W. Chen, M. Zhou, Diffusion-gan: Training gans with diffusion (2022).arXiv:2206.02262

  26. [26]

    Y. Xu, Y. Zhao, Z. Xiao, T. Hou, Ufogen: You forward once large scale text-to-image generation via diffusion gans (2023).arXiv:2311.09257

  27. [27]

    Nikolaidou, G

    K. Nikolaidou, G. Retsinas, V. Christlein, M. Seuret, G. Sfikas, E. B. Smith, H. Mokayed, M. Liwicki, WordStylist: Styled Verbatim Hand- written Text Generation with Latent Diffusion Models, in: ICDAR, 2023

  28. [28]

    Y. Zhu, Z. Li, T. Wang, M. He, C. Yao, Conditional text image gener- ation with diffusion models, in: Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 14235–14245

  29. [29]

    G. Dai, Y. Zhang, Q. Ke, Q. Guo, S. Huang, One-shot diffusion mim- icker for handwritten text generation, in: European Conference on Computer Vision, 2024

  30. [30]

    Pippi, S

    V. Pippi, S. Cascianelli, R. Cucchiara, Handwritten Text Generation from Visual Archetypes, in: CVPR, 2023

  31. [31]

    D. Kass, E. Vats, AttentionHTR: Handwritten Text Recognition Based on Attention Encoder-Decoder Networks, arXiv preprint arXiv:2201.09390 (2022)

  32. [32]

    Vaswani, N

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, I. Polosukhin, Attention is all you need, in: NeurIPS, 2017. 45

  33. [33]

    Do You Even Need Attention? A Stack of Feed-Forward Layers Does Surprisingly Well on ImageNet

    L. Melas-Kyriazi, Do you even need attention? a stack of feed-forward layers does surprisingly well on imagenet (2021).arXiv:2105.02723

  34. [34]

    Y. Tang, K. Han, J. Guo, C. Xu, Y. Li, C. Xu, Y. Wang, An image patch is a wave: Phase-aware vision mlp, in: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, 2022, p. 10925–10934.doi:10.1109/cvpr52688.2022.01066. URLhttp://dx.doi.org/10.1109/CVPR52688.2022.01066

  35. [35]

    Tong Dang Khoa, D

    H. Tong Dang Khoa, D. Hoai Nam, V. Nguyen Le Duy, Fw-gan: Frequency-driven handwriting synthesis with wave-modulated mlp generator, Expert Systems with Applications 299 (2026) 130175. doi:https://doi.org/10.1016/j.eswa.2025.130175. URLhttps://www.sciencedirect.com/science/article/pii/ S095741742503790X

  36. [36]

    S. Chen, E. Xie, C. Ge, R. Chen, D. Liang, P. Luo, Cyclemlp: A mlp- like architecture for dense visual predictions, IEEE Transactions on Pattern Analysis and Machine Intelligence 45 (12) (2023) 14284–14300

  37. [37]

    D. Lian, Z. Yu, X. Sun, S. Gao, As-mlp: An axial shifted mlp architec- ture for vision, arXiv preprint arXiv:2107.08391 (2021)

  38. [38]

    H. Mu, B. U. Tayyab, N. Chua, Spiralmlp: A lightweight vision mlp architecture, in: 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (W ACV), IEEE, 2025, pp. 8627–8637

  39. [39]

    C. Yeh, Y. Chen, A. Wu, C. Chen, F. Viégas, M. Wattenberg, At- tentionviz: A global view of transformer attention (2023).arXiv: 2305.03210. URLhttps://arxiv.org/abs/2305.03210

  40. [40]

    Hinton, A

    G. Hinton, A. Krizhevsky, I. Sutskever, Imagenet classification with deep convolutional neural networks, Advances in Neural Information Processing Systems (2012) 1097–1105doi:10.1145/3065386

  41. [41]

    K. He, X. Zhang, S. Ren, J. Sun, Deep Residual Learning for Image Recognition, in: CVPR, 2016

  42. [42]

    Simonyan, A

    K. Simonyan, A. Zisserman, Very deep convolutional networks for large- scale image recognition, in: ICLR, 2015. 46

  43. [43]

    Z. Liu, H. Mao, C.-Y. Wu, C. Feichtenhofer, T. Darrell, S. Xie, A convnet for the 2020s, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 11976–11986

  44. [44]

    Dosovitskiy, L

    A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al., An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale, in: ICLR, 2021

  45. [45]

    X. Ma, X. Dai, Y. Bai, Y. Wang, Y. Fu, Rewrite the stars, in: Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024

  46. [46]

    J. Yang, C. Li, X. Dai, J. Gao, Focal modulation networks, in: Advances in Neural Information Processing Systems (NeurIPS), 2022

  47. [47]

    Y. Rao, W. Zhao, Y. Tang, J. Zhou, S.-N. Lim, J. Lu, Hornet: Efficient high-order spatial interactions with recursive gated convolutions, in: Advances in Neural Information Processing Systems (NeurIPS), 2022

  48. [48]

    Guo, C.-Z

    M.-H. Guo, C.-Z. Lu, Z.-N. Liu, M.-M. Cheng, S.-M. Hu, Visual atten- tion network, Computational visual media 9 (4) (2023) 733–752

  49. [49]

    Y. Long, Z. Deng, B. Hao, G. Liu, K. Sun, S. Zhao, Rspnet: Towards universal lightweight road surface perception network, Expert Systems with Applications (2026) 132831

  50. [50]

    J. Xu, Y. Hu, Z. Gou, Y. Lu, D. Cui, Automated non-invasive method for measuring chicken comb and wattle using multi-camera system and cw-measure-pose, Expert Systems with Applications (2025) 129420

  51. [51]

    M. Yu, W. Chen, J. Hou, A lightweight network for concrete crack detection in complex environments based on cloud-edge collaboration, Expert Systems with Applications (2025) 130204

  52. [52]

    Z.-Q. J. X. Zhi-Qin John Xu, Y. Z. Yaoyu Zhang, T. L. Tao Luo, Y. X. Yanyang Xiao, Z. M. Zheng Ma, Frequency principle: Fourier analysis sheds light on deep neural networks, Communications in Computational Physics 28 (5) (2020) 1746–1767.doi:10.4208/cicp.oa-2020-0085. URLhttp://dx.doi.org/10.4208/cicp.OA-2020-0085

  53. [53]

    T. Luo, Z. Ma, Z.-Q. J. Xu, Y. Zhang, Theory of the frequency princi- ple for general deep neural networks, arXiv preprint arXiv:1906.09235 (2019). 47

  54. [54]

    Jacot, F

    A. Jacot, F. Gabriel, C. Hongler, Neural tangent kernel: Convergence and generalization in neural networks, in: Advances in neural informa- tion processing systems, 2018, pp. 8571–8580

  55. [55]

    T. Luo, Z. Ma, Z.-Q. J. Xu, Y. Zhang, On the exact computation of linear frequency principle dynamics and its generalization, SIAM Journal on Mathematics of Data Science 4 (4) (2022) 1272–1292.doi: 10.1137/21m1444400. URLhttp://dx.doi.org/10.1137/21m1444400

  56. [56]

    Explicitizing an Implicit Bias of the Frequency Principle in Two-layer Neural Networks

    Y. Zhang, Z.-Q. J. Xu, T. Luo, Z. Ma, Explicitizing an implicit bias of the frequency principle in two-layer neural networks (2019).arXiv: 1905.10264

  57. [57]

    Spectrum Dependent Learning Curves in Kernel Regression and Wide Neural Networks

    B. Bordelon, A. Canatar, C. Pehlevan, Spectrum dependent learning curves in kernel regression and wide neural networks (2020).arXiv: 2002.02561

  58. [58]

    Y. Cao, Z. Fang, Y. Wu, D.-X. Zhou, Q. Gu, Towards understanding the spectral bias of deep learning, arXiv preprint arXiv:1912.01198 (2019)

  59. [59]

    The Convergence Rate of Neural Networks for Learned Functions of Different Frequencies

    R. Basri, D. Jacobs, Y. Kasten, S. Kritchman, The convergence rate of neural networks for learned functions of different frequencies (2019). arXiv:1906.00425

  60. [60]

    G. Yang, H. Salman, A fine-grained spectral perspective on neural net- works, arXiv preprint arXiv:1907.10599 (2019)

  61. [61]

    W. E, C. Ma, L. Wu, Machine learning from a continuous viewpoint, arXiv preprint arXiv:1912.12777 (2019)

  62. [62]

    Z. Ni, J. Wu, Z. Wang, W. Yang, H. Wang, L. Ma, Misalignment- robust frequency distribution loss for image transformation, in: 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, 2024, p. 2910–2919.doi:10.1109/cvpr52733.2024. 00281. URLhttp://dx.doi.org/10.1109/CVPR52733.2024.00281

  63. [63]

    M. A. Souibgui, S. Biswas, S. K. Jemni, Y. Kessentini, A. Fornés, J. Lladós, U. Pal, Docentr: An end-to-end document image enhance- ment transformer, in: 2022 26th International Conference on Pattern Recognition (ICPR), IEEE, 2022, pp. 1699–1705. 48

  64. [64]

    M. A. Souibgui, Y. Kessentini, De-gan: A conditional generative ad- versarial network for document enhancement, IEEE Transactions on Pattern Analysis and Machine Intelligence 44 (3) (2020) 1180–1191

  65. [65]

    Z. Yang, B. Liu, Y. Xxiong, L. Yi, G. Wu, X. Tang, Z. Liu, J. Zhou, X. Zhang, Docdiff: Document enhancement via residual diffusion mod- els, in: Proceedings of the 31st ACM International Conference on Mul- timedia, 2023, pp. 2795–2806

  66. [66]

    Y. Rao, W. Zhao, Z. Zhu, J. Lu, J. Zhou, Global filter networks for image classification, Advances in neural information processing systems 34 (2021) 980–993

  67. [67]

    J. H. Lim, J. C. Ye, Geometric GAN, arXiv preprint arXiv:1705.02894 (2017)

  68. [68]

    Graves, S

    A. Graves, S. Fernández, F. Gomez, J. Schmidhuber, Connectionist temporal classification: labelling unsegmented sequence data with re- current neural networks, in: Proceedings of the 23rd International Conference on Machine Learning, ICML ’06, Association for Com- puting Machinery, New York, NY, USA, 2006, p. 369–376.doi: 10.1145/1143844.1143891. URLhttps...

  69. [69]

    X. Chen, Y. Duan, R. Houthooft, J. Schulman, I. Sutskever, P. Abbeel, Infogan: Interpretable representation learning by information maximiz- ing generative adversarial nets (2016).arXiv:1606.03657

  70. [70]

    J.-Y. Zhu, R. Zhang, D. Pathak, T. Darrell, A. A. Efros, O. Wang, E. Shechtman, Toward multimodal image-to-image translation (2017). arXiv:1711.11586

  71. [71]

    Lee, H.-Y

    H.-Y. Lee, H.-Y. Tseng, J.-B. Huang, M. Singh, M.-H. Yang, Di- verse Image-to-Image Translation via Disentangled Representations, Springer International Publishing, 2018, p. 36–52.doi:10.1007/ 978-3-030-01246-5_3. URLhttp://dx.doi.org/10.1007/978-3-030-01246-5_3

  72. [72]

    J. Gan, W. Wang, J. Leng, X. Gao, HiGAN+: Handwriting Imitation GAN with Disentangled Representations, ACM Trans. Graphics 42 (1) (2022) 1–17. 49

  73. [73]

    A. K. Bhunia, A. Das, A. K. Bhunia, P. S. R. Kishore, P. P. Roy, Handwriting Recognition in Low-Resource Scripts Using Adversarial Learning, in: CVPR, IEEE, 2019

  74. [74]

    DiffusionPen: Towards Controlling the Style of Handwritten Text Generation

    K. Nikolaidou, G. Retsinas, G. Sfikas, M. Liwicki, Diffusionpen: To- wards controlling the style of handwritten text generation, arXiv preprint arXiv:2409.06065 (2024)

  75. [75]

    Pippi, F

    V. Pippi, F. Quattrini, S. Cascianelli, R. Cucchiara, HWD: A Novel Evaluation Score for Styled Handwritten Text Generation, in: BMVC, 2023

  76. [76]

    D. P. Kingma, J. Ba, Adam: A method for stochastic optimization, in: ICLR, 2015

  77. [77]

    M. Li, T. Lv, J. Chen, L. Cui, Y. Lu, D. Florencio, C. Zhang, Z. Li, F. Wei, Trocr: Transformer-based optical character recognition with pre-trained models, Proceedings of the AAAI Conference on Artificial Intelligence 37 (11) (2023) 13094–13102.doi:10.1609/aaai.v37i11. 26538. URLhttp://dx.doi.org/10.1609/aaai.v37i11.26538 50