Pith. sign in

REVIEW 2 major objections 5 minor 52 references

Neural networks plus pathfinders find Seiberg-duality chains for modest quivers faster than blind or pure physics search.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-31 01:43 UTC pith:T4P5Z3V3

load-bearing objection Solid empirical ML tool for tracing Seiberg dualities on modest quivers; hybrids beat BFS and LCA with public code and honest failure analysis. the 2 major comments →

arxiv 2607.28628 v1 pith:T4P5Z3V3 submitted 2026-07-30 hep-th cs.AIcs.LGhep-ph

Learning to Trace Seiberg Dualities

classification hep-th cs.AIcs.LGhep-ph
keywords Seiberg dualityquiver mutationsgraph neural networkspathfindingcomputational complexitysupersymmetric gauge theoriesA* searchtoric Calabi-Yau
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Given two supersymmetric quiver gauge theories, deciding whether they are related by a chain of Seiberg dualities (quiver mutations) and finding a short chain is combinatorially hard once the number of nodes grows. This paper trains graph networks with transformers on mutation trees grown from D-brane seed theories, then uses the networks as distance estimators and next-move advisers inside A* and beam-search pathfinders. For quivers of roughly ten nodes the hybrid searchers outperform both unguided breadth-first search and a pure rank-minimizing physics heuristic while keeping near-perfect success rates. The same machinery yields a practical complexity measure and a benchmark task for AI applied to theoretical physics.

Core claim

For quivers with a modest number of nodes (order 10), transformer-plus-MLP graph networks used as policies inside bidirectional A* and beam pathfinders, especially when hybridized with a lowest-common-ancestor rank heuristic, find Seiberg-duality paths more efficiently than unguided BFS or pure deterministic LCA search while maintaining near-100% success.

What carries the argument

Hybrid pathfinders that combine a Distance GNN (heuristic estimating mutation distance), an Adviser GNN (policy over which node to dualize next), and a physics-informed Lowest-Common-Ancestor cost that prefers rank-reducing moves; the networks are trained on BFS-generated mutation trees from toric Calabi–Yau seed quivers.

Load-bearing premise

Distances used for training and scoring are taken from BFS trees via longest common path prefixes and are only upper bounds once quiver symmetries are ignored, biasing the data toward pairs that share a low-rank common ancestor—the same criterion the physics baseline exploits.

What would settle it

Retrain the same architectures on a dataset that fully quotients by quiver automorphisms and re-measures efficiency ratios against LCA; if the hybrid advantage disappears or success rates fall below the pure LCA baseline at comparable complexity, the claimed outperformance is an artifact of the biased distance definition.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Computational complexity of duality chains can be read off as C = D log10 K and used to estimate how far a holographic RG flow proceeds down a warped throat.
  • The same hybrid pathfinders give a practical tool for deciding whether two given quivers are Seiberg dual and for enumerating short connecting sequences.
  • The task supplies a concrete benchmark for frontier AI models on a well-defined theoretical-physics search problem.
  • Hybrid NN-plus-physics search remains superior to pure LCA up to roughly 1.5–2 times the training complexity before degrading.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same mutation-pathfinding setup could be reused for other cluster-algebra or BPS-quiver problems where the mutation graph is infinite and non-monotonic heuristics appear.
  • Once automorphism-aware distances are available, the residual gap between Hybrid LCA and pure LCA would quantify how much genuine pattern learning the networks have acquired beyond the rank-minimizing bias built into the training trees.
  • Scaling the training set to larger node counts would test whether the observed complexity threshold (roughly 1.5–2× training C) continues to hold or collapses, giving a concrete scaling law for this class of physics search tasks.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper studies the computational problem of deciding whether two 4d N=1 quiver gauge theories are related by a sequence of Seiberg dualities (quiver mutations) and of finding short duality paths. Starting from toric Calabi–Yau seed quivers, the authors generate mutation trees by BFS, train a Distance GNN (DGNN) to regress mutation distance and an Adviser GNN (AGNN) to propose the next node to dualize, and embed both networks as heuristics/costs inside bidirectional A*, beam search, a physics-inspired Lowest-Common-Ancestor (LCA) rank heuristic, and hybrid combinations. On in-distribution and out-of-distribution quivers with O(10) nodes they report that transformer+MLP architectures, especially Hybrid and Hybrid-LCA pathfinders, substantially outperform unguided BFS and improve on pure LCA in efficiency while retaining near-100% success; a controlled complexity study (C = D log10 K) estimates the hybrids remain superior up to roughly 1.5–2× the training complexity before degrading.

Significance. The work supplies a concrete, reproducible benchmark that sits at the intersection of Seiberg duality, cluster-algebra mutations, and modern graph ML. Strengths that should be credited explicitly include: (i) public checkpoints and pathfinder code, (ii) systematic ID/OOD tables, success heatmaps and efficiency ratios (Tables 2a–2b, §§6–7), (iii) an honest failure-mode section that retrains at lower complexity to locate the breaking point (Eq. 7.3), and (iv) transparent disclosure that training distances are upper bounds from BFS path prefixes and that the dataset is biased toward low-rank common ancestors. If the empirical claims hold under the stated regime, the paper offers both a practical tool for duality searches and a useful stress-test for frontier AI models applied to theoretical physics.

major comments (2)
  1. [§3.1, Eq. (3.5); §6.2.2] §3.1, Eq. (3.5): duality distance is defined from BFS path-prefix LCAs without full quotienting by quiver automorphisms, so the reported d is an upper bound and the training set is biased toward pairs that share a low-rank common ancestor—the same criterion the LCA baseline exploits. The authors flag this (§6.1.5–6.2.2, Fig. 24) and still show gains on infinite-mutation OOD families never seen in training, but a quantitative bound on the residual bias (e.g., symmetry-reduced distances on a subsample, or ER recomputed after random node relabeling) is needed before the ~1.1–1.2× Hybrid-vs-LCA claim can be treated as fully independent of the generation procedure.
  2. [§5.2, Table 1] §5.2 and Table 1: the Hybrid-LCA cost weights (c_det dec/eq/inc, λ_det cost, λ_AR, λ_DGNN, λ_LCA) are stated to have been optimized for 100% success and maximal ER. It is unclear whether this optimization used a held-out validation split distinct from the 500-pair evaluation samples of §6. Without that separation the modest efficiency gains over pure LCA risk mild contamination by the evaluation distribution; a short statement of the hyperparameter protocol (or a re-evaluation with frozen weights chosen only on a validation slice) would remove the ambiguity.
minor comments (5)
  1. [§2.2] §2.2: typographical slip “the far more generic cas is” → “case”.
  2. [§4.1.1, Fig. 8] Fig. 8c and related MAE plots: the median error is negative (ˆd < d). A one-sentence remark that the network systematically under-estimates large distances (consistent with the non-monotonicity discussion) would help the reader interpret the heuristic quality.
  3. [§6.2, Appendix A.2] Appendix A.2 / Fig. 37: several quivers are anomalous as 4d N=1 theories. The text already notes they remain valid as BPS quivers; a brief cross-reference in the main OOD discussion (§6.2) would prevent confusion.
  4. [§7] The complexity measure C = D log10 K is introduced only in §7. A forward pointer in the introduction or §5.3 would make the later breaking-point analysis easier to anticipate.
  5. [Note Added / §1] References: the concurrent DualityCert work [25] is noted; a one-sentence clarification of the complementary goals (path-finding/complexity vs. verifier-gated claim repair) already present in the Note Added could be moved into the introduction for readers who skip the note.

Circularity Check

1 steps flagged

No load-bearing circularity: empirical pathfinder benchmarks on BFS-generated labels, with disclosed LCA-affinity bias that does not force the claimed efficiency gains.

specific steps
  1. other [§3.1 Eq. (3.5); §6.2.2]
    "d(QA, QB)=|PA|+|PB|−2|PLCA|.... beating the LCA pathfinder on the datasets we generated, as mentioned already in the previous sections, is the most difficult test for our pathfinders because the LCA pathfinder follows the same criterion (but opposite) that we used to generate the dataset in the first place"

    Training distances and pair structure are defined via longest common mutation-path prefixes (i.e., LCAs of the BFS tree). The deterministic LCA baseline optimizes the same low-rank-ancestor geometry. This creates a mild affinity between labels and one baseline, so part of the Hybrid-vs-LCA contest is run on data shaped by LCA logic. It is not a by-construction identity: Hybrid still explores fewer nodes ~84% of the time at equal path length, sometimes finds longer paths, and loses to LCA on several OOD slices; BFS and infinite-mutation OOD checks remain independent.

full rationale

The paper is an applied ML/pathfinding study, not a first-principles derivation of a physical law. Seiberg mutation rules (Eqs. 2.2–2.4) are taken as given; training pairs and distance labels are produced by BFS over duality trees (Alg. 1, Eq. 3.5) and used to supervise DGNN/AGNN; pathfinders are then scored by success rate and nodes explored against BFS and LCA baselines on held-out ID and OOD quivers. That pipeline is standard supervised learning plus search, not a self-definitional loop: nothing equates a fitted parameter to the reported efficiency ratio by construction, and Hybrid/Hybrid-LCA can and do underperform LCA on small-K, finite-mutation, and anomalous OOD families (Figs. 28–30). The only mild affinity is that Eq. 3.5 distances are LCA-prefix lengths and the dataset therefore favors pairs that share a low-rank ancestor—the same structural cue the deterministic LCA baseline exploits—which the authors explicitly flag (§3.1, §6.1.5–6.2.2). That is an evaluation-bias caveat, already bounded by bidirectional-BFS comparisons, infinite-mutation OOD families never seen in training (Fig. 38), and the controlled complexity stress test (§7). No self-citation uniqueness theorem, smuggled ansatz, or renaming of a known result carries the central claim. Score 1 reflects only that disclosed generation affinity, not forced circularity.

Axiom & Free-Parameter Ledger

4 free parameters · 4 axioms · 2 invented entities

The central empirical claim rests on standard Seiberg/Fomin–Zelevinsky mutation rules, a deliberate truncation of QFT data to (adjacency, ranks), a BFS-derived distance that ignores full automorphism reduction, and several fitted pathfinder weights. No new physical entities are postulated; the invented objects are algorithmic.

free parameters (4)
  • Hybrid-LCA cost weights (c_det dec, c_det eq, c_det inc, λ_det cost, λ_AR, λ_DGNN, λ_LCA) = c_det dec=0.3, c_det eq=2.7, c_det inc=3.1, λ_det cost=1.3 (others tuned analogously)
    Optimized to maximize efficiency ratio at 100% success rate rather than derived from first principles (§5.2).
  • AGNN beam width B = B=3
    Fixed from top-k accuracy curves; controls completeness–efficiency tradeoff.
  • DGNN/AGNN architecture widths and depths (H=64, 3 GNN layers, 2 transformer layers, dropout 0.2, etc.) = H=64; DGNN ~250 epochs; AGNN stopped at epoch 39–49
    Standard ML hyperparameters chosen for training stability; affect absolute accuracy.
  • LCA rank-change costs c_dec:c_eq:c_inc = ~1:10:100
    Hand-set hierarchy that defines the physics baseline.
axioms (4)
  • domain assumption Seiberg duality on a quiver node is exactly the Fomin–Zelevinsky mutation of (A,N) with vector-like pairs deleted and ranks required non-negative (Eqs. 2.2–2.4).
    Standard in the quiver-gauge-theory literature; superpotential is ignored except for automatic mass deletion of newly generated vector-like pairs (§2.1).
  • ad hoc to paper Duality distance equals |P_A|+|P_B|−2|P_LCA| computed from BFS path prefixes in the mutation tree (Eq. 3.5), without full symmetry reduction.
    Defines both training labels and the notion of ‘shortest path’ used in evaluation; authors note it is an upper bound when automorphisms exist (§3.1, §4.1.1).
  • domain assumption Anomaly cancellation (weighted arrow balance) is required for training seeds; OOD tests may drop it.
    Physical for 4D N=1 QFTs; relaxed for BPS-quiver-style OOD stress tests (§2.1, §6.2).
  • standard math Two quivers are identified up to node permutation via Weisfeiler–Lehman hashing during search.
    Standard graph-isomorphism heuristic used to terminate pathfinders (§5.2).
invented entities (2)
  • Distance GNN (DGNN) and Adviser GNN (AGNN) no independent evidence
    purpose: Regress mutation distance and propose next dualization node for pathfinders.
    Custom graph-transformer architectures adapted from MPNN/GraphGPS; no claim of new physics.
  • Hybrid LCA pathfinder and complexity C=D log10 K no independent evidence
    purpose: Combine NN policies with rank-reducing physics heuristic; diagnose scaling limits.
    Algorithmic constructs introduced to organize experiments (§5.2, §7).

pith-pipeline@v1.2.0-daily-grok45 · 49788 in / 3497 out tokens · 76739 ms · 2026-07-31T01:43:46.750012+00:00 · methodology

0 comments
read the original abstract

Dualities play an important role in establishing both microscopic and emergent phenomena in a wide range of physical systems. In practice, though, it can often be computationally challenging to establish when two systems are dual, even when all of the "rules of the game" are well-known. Said differently, when confronted with two systems, how can one efficiently establish that they are in fact dual? In this paper we use machine learning methods to address this question for Seiberg dualities of supersymmetric quiver gauge theories. Mathematically, this involves establishing mutations of quivers, which is in turn a variation on the theme of "learning to unknot". On the one hand, this leads us to a practical tool for establishing the computational complexity of different dualities. On the other hand, it also allows us to study how different network architectures learn how to trace Seiberg dualities. We find that for quivers with a modest number of quiver nodes (of order $10$), different network architectures consisting of transformers and multi-layer perceptrons tend to outperform deterministic algorithms. Supplementing the network by well-established pathfinder algorithms (essentially "Google Maps for quivers") leads to an additional improvement in the efficiency and accuracy of the search strategy. We anticipate that this class of questions can serve as a useful benchmark for frontier AI models applied to theoretical physics.

Figures

Figures reproduced from arXiv: 2607.28628 by Alessandro Mininno, Gary Shiu, Jonathan J. Heckman, Shani Meynet.

Figure 1
Figure 1. Figure 1: Schematic representation of Seiberg duality at the level of quiver theories. The [PITH_FULL_IMAGE:figures/full_fig_p010_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Seiberg Duality Tree for C 3/Z3 inspired by [32, [PITH_FULL_IMAGE:figures/full_fig_p011_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Example of dualization of C 3/Z (1,1,1) 3 with respect to the red node (top quiver node). In the context of string based constructions, RG flows/duality moves correspond to motion along a warped throat. Moving further into the bulk of a string compactification thus has a direct bearing on how energy scales impact the computational complexity of a given collection of quivers. 2.3 Seed and Validation Theorie… view at source ↗
Figure 4
Figure 4. Figure 4: Schematic representation of the BFS generated by the mutation [PITH_FULL_IMAGE:figures/full_fig_p014_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Schematic representation of the duality distance [PITH_FULL_IMAGE:figures/full_fig_p015_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Schematic representations of DGNN and AGNN architectures. [PITH_FULL_IMAGE:figures/full_fig_p016_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Training logs for the DGNN model The network parameters are optimized by minimizing the Mean Squared Error (MSE) between the prediction ˆd and the path distance d computed at the generation of the database. For a given batch of size B, the loss function is defined as: LMSE = 1 B X B k=1 ( ˆdk − dk) 2 . (4.9) While MSE is utilized for gradient descent to penalize outlier predictions, the accuracy of the mod… view at source ↗
Figure 8
Figure 8. Figure 8: Benchmark results for DGNN. Because we apply the DGNN as the heuristic function of an A∗ search, as we will ex￾plain in Section 5.2, we have decided to perform further benchmark tests on the kind of heuristic function that the DGNN can provide. One test is to check if the heuristic function is monotonic, so that we reduce the possibility that the A∗ search requires processing the same node multiple times. … view at source ↗
Figure 9
Figure 9. Figure 9: Evaluation of DGNN monotonicity. The non-monotonicity of the heuristic results [PITH_FULL_IMAGE:figures/full_fig_p021_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Training logs for the Autoregressive model. [PITH_FULL_IMAGE:figures/full_fig_p023_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: Benchmark results for the AGNN. While next-step accuracy degrades at larger [PITH_FULL_IMAGE:figures/full_fig_p025_11.png] view at source ↗
Figure 12
Figure 12. Figure 12: Policy Margin analysis. While Top-1 inversions cause some paths to fail, the high [PITH_FULL_IMAGE:figures/full_fig_p026_12.png] view at source ↗
Figure 13
Figure 13. Figure 13: Schematic representations of the two kinds of pathfinders we consider in this work. [PITH_FULL_IMAGE:figures/full_fig_p028_13.png] view at source ↗
Figure 14
Figure 14. Figure 14: Success Heatmap 2 4 6 8 10 12 100 101 102 103 104 Distance EER K=3 K=4 K=5 K=6 K=7 K=8 K=9 K=10 K=11 K=12 K=13 BFS (a) Efficiency Ratio 2 4 6 8 10 12 0 50 100 150 Distance GE K=3 K=4 K=5 K=6 K=7 K=8 K=9 K=10 K=11 K=12 K=13 Opt (b) Guidance Efficiency [PITH_FULL_IMAGE:figures/full_fig_p038_14.png] view at source ↗
Figure 15
Figure 15. Figure 15: Global performance of the DGNN pathfinder tested on theories in the testing set [PITH_FULL_IMAGE:figures/full_fig_p038_15.png] view at source ↗
Figure 16
Figure 16. Figure 16: Success Heatmap The AGNN pathfinder is the only pathfinder in this work that is unidirectional. However, for consistency in the analysis with the other pathfinders, we still compare it with the bidi￾rectional BFS baseline. Additionally, the beam search we performed set the beam width to B = 3. This is the reason why, in Figure 17a, for quivers with K = 3 and small distances, the performance of the AGNN pa… view at source ↗
Figure 17
Figure 17. Figure 17: Global performance of the AGNN pathfinder evaluated on theories in the testing [PITH_FULL_IMAGE:figures/full_fig_p040_17.png] view at source ↗
Figure 18
Figure 18. Figure 18: Success Heatmap has a cost function given by the probability distribution predicted by the AGNN model.29 For ID theories, the Hybrid pathfinder resolves the guidance inefficiency of DGNN and the success-rate drop of the AGNN pathfinder, giving the so-far most efficient NN-guided pathfinder with a 100% SR in finding a path (if the path exists). 6.1.4 LCA Pathfinder Performance In this section, we discuss t… view at source ↗
Figure 19
Figure 19. Figure 19: Global performance of the Hybrid pathfinder tested on theories in the testing set [PITH_FULL_IMAGE:figures/full_fig_p042_19.png] view at source ↗
Figure 20
Figure 20. Figure 20: Success Heatmap 41 [PITH_FULL_IMAGE:figures/full_fig_p042_20.png] view at source ↗
Figure 21
Figure 21. Figure 21: The global performance of the LCA pathfinder tested on the ID dataset. [PITH_FULL_IMAGE:figures/full_fig_p043_21.png] view at source ↗
Figure 22
Figure 22. Figure 22: EER for the three pathfinders using LCA baseline in the ID dataset. [PITH_FULL_IMAGE:figures/full_fig_p044_22.png] view at source ↗
Figure 23
Figure 23. Figure 23: Nodes explored and average EER per nodes and distances of the Hybrid pathfinder [PITH_FULL_IMAGE:figures/full_fig_p045_23.png] view at source ↗
Figure 24
Figure 24. Figure 24: Total percentage of path deviations (Shorter and Longer) found in ID dataset by [PITH_FULL_IMAGE:figures/full_fig_p046_24.png] view at source ↗
Figure 25
Figure 25. Figure 25: EER for the Hybrid LCA pathfinder on the ID dataset [PITH_FULL_IMAGE:figures/full_fig_p047_25.png] view at source ↗
Figure 26
Figure 26. Figure 26: Nodes explored and average EER per nodes and distances of the Hybrid LCA [PITH_FULL_IMAGE:figures/full_fig_p048_26.png] view at source ↗
Figure 27
Figure 27. Figure 27: Comparison of paths found in ID dataset by Hybrid LCA pathfinder when com [PITH_FULL_IMAGE:figures/full_fig_p049_27.png] view at source ↗
Figure 28
Figure 28. Figure 28: Success Heatmap for Theories in Figure [PITH_FULL_IMAGE:figures/full_fig_p050_28.png] view at source ↗
Figure 29
Figure 29. Figure 29: Success Heatmap for Theories in Figure [PITH_FULL_IMAGE:figures/full_fig_p051_29.png] view at source ↗
Figure 30
Figure 30. Figure 30: EER for the pathfinders using LCA baseline in the OOD dataset of Figure [PITH_FULL_IMAGE:figures/full_fig_p053_30.png] view at source ↗
Figure 31
Figure 31. Figure 31: EER for the pathfinders using LCA baseline in the OOD dataset of Figure [PITH_FULL_IMAGE:figures/full_fig_p054_31.png] view at source ↗
Figure 32
Figure 32. Figure 32: Success Rate, Effective Efficiency Ratio (with respect to LCA baseline) and Effec [PITH_FULL_IMAGE:figures/full_fig_p056_32.png] view at source ↗
Figure 33
Figure 33. Figure 33: Success Rate, Effective Efficiency Ratio (with respect to LCA baseline) and Effec [PITH_FULL_IMAGE:figures/full_fig_p057_33.png] view at source ↗
Figure 34
Figure 34. Figure 34: Examples of C 3/Zn quivers theories. The corresponding superpotential can be derived from (A.3). The quiver for this theory is shown in Figure 35a. 3. The 4d N = 1 SCFTs on D3-branes probing C(dP2). This theory is characterized by five gauge groups and bifundamental fields X15 , Y15 , X21 , X25 , X31 , X32 , X42 , X43 , X53 , X54 , Y54 , (A.6) subjected to the superpotential WC(dP2) = X15X53X31 − X25X53X3… view at source ↗
Figure 35
Figure 35. Figure 35: Quivers of C(dPn) theories. The corresponding superpotentials are in Eqs. (A.5), (A.7) and (A.9). Y2k+1|2k−1 , Y2k+2|2k (k = 1 . . . q), Y2k+2|2k−1 (k = q + 1 . . . p), (A.11) where α = 1, 2 is the doublet index. These fields enter to a superpotential as WC(Y p,q) = X q k=1 ϵαβ  U α 2k−1|2kV β 2k|2k+1Y2k+1|2k−1 − V α 2k|2k+1U β 2k+1|2k+2Y2k+2|2k  + X p k=q+1 (−1)k−q−1 ϵαβU α 2k−1|2kZ2k|2k+1U β 2k+1|2k+2… view at source ↗
Figure 36
Figure 36. Figure 36: Examples of C(Y p,q) quivers theories. The corresponding superpotential can be derived from (A.12). 61 [PITH_FULL_IMAGE:figures/full_fig_p062_36.png] view at source ↗
Figure 37
Figure 37. Figure 37: Quiver families considered in [13]. We comment that some of these quivers are intrinsically anomalous when inerpreted as 4D N = 1 QFTs (i.e., inconsistent), but are perfectly consistent when viewed as the BPS quivers of 4D N = 2 QFTs. As such, they provide an interesting testing ground for studying mutations of quivers. 62 [PITH_FULL_IMAGE:figures/full_fig_p063_37.png] view at source ↗
Figure 38
Figure 38. Figure 38: Further quiver theories used to test the NN. Nodes with a rank other than one [PITH_FULL_IMAGE:figures/full_fig_p064_38.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

52 extracted references · 43 linked inside Pith

  1. [1]

    Electric - Magnetic Duality in Supersymmetric Non-Abelian Gauge Theories,

    N. Seiberg, “Electric - Magnetic Duality in Supersymmetric Non-Abelian Gauge Theories,”Nucl. Phys. B435(1995) 129–146,arXiv:hep-th/9411149

  2. [2]

    Cluster Algebras I: Foundations,

    S. Fomin and A. Zelevinsky, “Cluster Algebras I: Foundations,”arXiv:math/0104151

  3. [3]

    Cluster Algebras II: Finite Type Classification,

    S. Fomin and A. Zelevinsky, “Cluster Algebras II: Finite Type Classification,”Invent. Math.154no. 1, (2003) 63–121,arXiv:math/0208229

  4. [4]

    Supergravity and a Confining Gauge Theory: Duality Cascades andχSB Resolution of Naked Singularities,

    I. R. Klebanov and M. J. Strassler, “Supergravity and a Confining Gauge Theory: Duality Cascades andχSB Resolution of Naked Singularities,”JHEP08(2000) 052, arXiv:hep-th/0007191

  5. [5]

    Superconformal Field Theory on Three-Branes at a Calabi-Yau Singularity,

    I. R. Klebanov and E. Witten, “Superconformal Field Theory on Three-Branes at a Calabi-Yau Singularity,”Nucl. Phys. B536(1998) 199–218,arXiv:hep-th/9807080

  6. [6]

    Gravity Duals of Supersymmetric SU(N)×SU(N+M) Gauge Theories,

    I. R. Klebanov and A. A. Tseytlin, “Gravity Duals of Supersymmetric SU(N)×SU(N+M) Gauge Theories,”Nucl. Phys. B578(2000) 123–138, arXiv:hep-th/0002159

  7. [7]

    Chaotic Duality in String Theory,

    S. Franco, Y.-H. He, C. Herzog, and J. Walcher, “Chaotic Duality in String Theory,” Phys. Rev. D70(2004) 046006,arXiv:hep-th/0402120

  8. [8]

    Statistical Inference and String Theory,

    J. J. Heckman, “Statistical Inference and String Theory,”Int. J. Mod. Phys. A30 no. 26, (2015) 1550160,arXiv:1305.3621 [hep-th]

  9. [9]

    Relative Entropy and Proximity of Quantum Field Theories,

    V. Balasubramanian, J. J. Heckman, and A. Maloney, “Relative Entropy and Proximity of Quantum Field Theories,”JHEP05(2015) 104,arXiv:1410.6809 [hep-th]

  10. [10]

    Misanthropic Entropy and Renormalization as a Communication Channel,

    R. Fowler and J. J. Heckman, “Misanthropic Entropy and Renormalization as a Communication Channel,”Int. J. Mod. Phys. A37no. 16, (2022) 2250109, arXiv:2108.02772 [hep-th]

  11. [11]

    Higher Symmetries of 5D Orbifold SCFTs,

    M. Del Zotto, J. J. Heckman, S. N. Meynet, R. Moscrop, and H. Y. Zhang, “Higher Symmetries of 5D Orbifold SCFTs,”Phys. Rev. D106no. 4, (2022) 046010, arXiv:2201.08372 [hep-th]

  12. [12]

    Quiver Approach to Symmetry Theories,

    V. Chakrabhavi, M. Cvetiˇ c, J. J. Heckman, and S. Meynet, “Quiver Approach to Symmetry Theories,”arXiv:2605.30354 [hep-th]

  13. [13]

    Quiver Mutations, Seiberg Duality and Machine Learning,

    J. Bao, S. Franco, Y.-H. He, E. Hirst, G. Musiker, and Y. Xiao, “Quiver Mutations, Seiberg Duality and Machine Learning,”Phys. Rev. D102no. 8, (2020) 086013, arXiv:2006.10783 [hep-th]. 69

  14. [14]

    BPS spectroscopy with reinforcement learning,

    F. Carta, A. Gauntlett, F. Griffin, and Y.-H. He, “BPS spectroscopy with reinforcement learning,”Phys. Lett. B868(2025) 139646,arXiv:2501.14863 [hep-th]

  15. [15]

    Learning to Unknot,

    S. Gukov, J. Halverson, F. Ruehle, and P. Su lkowski, “Learning to Unknot,”Mach. Learn. Sci. Tech.2no. 2, (2021) 025035,arXiv:2010.16263 [math.GT]

  16. [16]

    Grokking modular arithmetic,

    A. Gromov, “Grokking modular arithmetic,”arXiv:2301.02679 [cs.LG]

  17. [17]

    Arkani-Hamed, J

    N. Arkani-Hamed, J. L. Bourjaily, F. Cachazo, A. B. Goncharov, A. Postnikov, and J. Trnka,Grassmannian Geometry of Scattering Amplitudes. Cambridge University Press, 4, 2016.arXiv:1212.5605 [hep-th]

  18. [18]

    Lectures on D-branes, gauge theories and Calabi-Yau singularities,

    Y.-H. He, “Lectures on D-branes, gauge theories and Calabi-Yau singularities,” in1st Hangzhou-Beijing International Summer School. 8, 2004.arXiv:hep-th/0408142

  19. [19]

    D-branes on Calabi-Yau manifolds,

    P. S. Aspinwall, “D-branes on Calabi-Yau manifolds,” inTheoretical Advanced Study Institute in Elementary Particle Physics (TASI 2003): Recent Trends in String Theory, pp. 1–152. 3, 2004.arXiv:hep-th/0403166

  20. [20]

    A Comprehensive Survey of Brane Tilings,

    S. Franco, Y.-H. He, C. Sun, and Y. Xiao, “A Comprehensive Survey of Brane Tilings,” Int. J. Mod. Phys. A32no. 23n24, (2017) 1750142,arXiv:1702.03958 [hep-th]

  21. [21]

    Attention Is All You Need,

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention Is All You Need,”arXiv e-prints(June, 2017) arXiv:1706.03762,arXiv:1706.03762 [cs.CL]

  22. [22]

    A Formal Basis for the Heuristic Determination of Minimum Cost Paths,

    P. E. Hart, N. J. Nilsson, and B. Raphael, “A Formal Basis for the Heuristic Determination of Minimum Cost Paths,”IEEE Transactions on Systems Science and Cybernetics4no. 2, (1968) 100–107

  23. [23]

    B. T. Lowerre,The Harpy speech recognition system. PhD thesis, Carnegie Mellon University, Pennsylvania, Apr., 1976

  24. [24]

    Proficient Graph Neural Network Design by Accumulating Knowledge on Large Language Models,

    J. Wang, H. Liu, S. Di, Z. Wang, J. Wang, L. Chen, and X. Zhou, “Proficient Graph Neural Network Design by Accumulating Knowledge on Large Language Models,” arXiv e-prints(Aug., 2024) arXiv:2408.06717,arXiv:2408.06717 [stat.ML]

  25. [25]

    DualityCert: Verifier-Gated Language-Model Repair of Broken Duality Claims in Quantum Field Theory,

    X. Yu, “DualityCert: Verifier-Gated Language-Model Repair of Broken Duality Claims in Quantum Field Theory,”arXiv:2607.23614 [cs.CR]

  26. [26]

    Dynamical Supersymmetry Breaking in Supersymmetric QCD,

    I. Affleck, M. Dine, and N. Seiberg, “Dynamical Supersymmetry Breaking in Supersymmetric QCD,”Nucl. Phys. B241(1984) 493–534

  27. [27]

    The Runaway quiver,

    K. A. Intriligator and N. Seiberg, “The Runaway quiver,”JHEP02(2006) 031, arXiv:hep-th/0512347. 70

  28. [28]

    Seiberg Duality for Quiver Gauge Theories,

    D. Berenstein and M. R. Douglas, “Seiberg Duality for Quiver Gauge Theories,” arXiv:hep-th/0207027

  29. [29]

    Exceptional Collections and del Pezzo Gauge Theories,

    C. P. Herzog, “Exceptional Collections and del Pezzo Gauge Theories,”JHEP04 (2004) 069,arXiv:hep-th/0310262

  30. [30]

    Seiberg Duality is an Exceptional Mutation,

    C. P. Herzog, “Seiberg Duality is an Exceptional Mutation,”JHEP08(2004) 064, arXiv:hep-th/0405118

  31. [31]

    D-Branes on Vanishing del Pezzo Surfaces,

    P. S. Aspinwall and I. V. Melnikov, “D-Branes on Vanishing del Pezzo Surfaces,” JHEP12(2004) 042,arXiv:hep-th/0405134

  32. [32]

    Duality walls, duality trees and fractional branes,

    S. Franco, A. Hanany, Y.-H. He, and P. Kazakopoulos, “Duality walls, duality trees and fractional branes,”arXiv:hep-th/0306092

  33. [33]

    Neural Message Passing for Quantum Chemistry,

    J. Gilmer, S. S. Schoenholz, P. F. Riley, O. Vinyals, and G. E. Dahl, “Neural Message Passing for Quantum Chemistry,”arXiv e-prints(Apr., 2017) arXiv:1704.01212, arXiv:1704.01212 [cs.LG]

  34. [34]

    Recipe for a General, Powerful, Scalable Graph Transformer,

    L. Ramp´ aˇ sek, M. Galkin, V. P. Dwivedi, A. T. Luu, G. Wolf, and D. Beaini, “Recipe for a General, Powerful, Scalable Graph Transformer,”arXiv e-prints(May, 2022) arXiv:2205.12454,arXiv:2205.12454 [cs.LG]

  35. [35]

    Attention Is All You Need,

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention Is All You Need,” inAdvances in Neural Information Processing Systems, vol. 30, pp. 5998–6008. Curran Associates, Inc., 2017

  36. [36]

    Crystal Melting and Black Holes,

    J. J. Heckman and C. Vafa, “Crystal Melting and Black Holes,”JHEP09(2007) 011, arXiv:hep-th/0610005

  37. [37]

    BPS Quivers and Spectra of Complete N=2 Quantum Field Theories,

    M. Alim, S. Cecotti, C. Cordova, S. Espahbodi, A. Rastogi, and C. Vafa, “BPS Quivers and Spectra of Complete N=2 Quantum Field Theories,”Commun. Math. Phys.323(2013) 1185–1227,arXiv:1109.4941 [hep-th]

  38. [38]

    Transforming Calabi-Yau Constructions: Generating New Calabi-Yau Manifolds with Transformers,

    J. H. T. Yip, C. Arnal, F. Charton, and G. Shiu, “Transforming Calabi-Yau Constructions: Generating New Calabi-Yau Manifolds with Transformers,” arXiv:2507.03732 [hep-th]

  39. [39]

    Sampling string vacua using generative models,

    M. Walden and M. Larfors, “Sampling string vacua using generative models,”Mach. Learn. Sci. Tech.7no. 1, (2026) 015018,arXiv:2509.16029 [hep-th]

  40. [40]

    Generating Special Triangulations with Transformers,

    C. Arnal, J. H. T. Yip, F. Charton, and G. Shiu, “Generating Special Triangulations with Transformers,” in . 6, 2026.arXiv:2606.26660 [hep-th]

  41. [41]

    Exploring Line Bundle Standard Models with Transformers,

    J. H. T. Yip, A. Mininno, and G. Shiu, “Exploring Line Bundle Standard Models with Transformers,”arXiv:2607.00078 [hep-th]. 71

  42. [42]

    Reconstructing conformal field theoretical compositions with Transformers,

    H. Cao, G. Merz, K. Cranmer, and G. Shiu, “Reconstructing conformal field theoretical compositions with Transformers,”arXiv:2605.01072 [hep-th]

  43. [43]

    Holographic Complexity Equals Bulk Action?,

    A. R. Brown, D. A. Roberts, L. Susskind, B. Swingle, and Y. Zhao, “Holographic Complexity Equals Bulk Action?,”Phys. Rev. Lett.116no. 19, (2016) 191301, arXiv:1509.07876 [hep-th]

  44. [44]

    Comments on Holographic Complexity,

    D. Carmi, R. C. Myers, and P. Rath, “Comments on Holographic Complexity,”JHEP 03(2017) 118,arXiv:1612.00433 [hep-th]

  45. [45]

    Quantum Complexity of Time Evolution with Chaotic Hamiltonians,

    V. Balasubramanian, M. Decross, A. Kar, and O. Parrikar, “Quantum Complexity of Time Evolution with Chaotic Hamiltonians,”JHEP01(2020) 134,arXiv:1905.05765 [hep-th]

  46. [46]

    Generalized Complexity Distances and Non-Invertible Symmetries,

    J. J. Heckman, R. J. Hicks, and C. Murdia, “Generalized Complexity Distances and Non-Invertible Symmetries,”arXiv:2604.14275 [hep-th]

  47. [47]

    Finite Heisenberg groups from nonAbelian orbifold quiver gauge theories,

    B. A. Burrington, J. T. Liu, and L. A. Pando Zayas, “Finite Heisenberg groups from nonAbelian orbifold quiver gauge theories,”Nucl. Phys. B794(2008) 324–347, arXiv:hep-th/0701028

  48. [48]

    Vafa–Witten Invariants from Exceptional Collections,

    G. Beaujard, J. Manschot, and B. Pioline, “Vafa–Witten Invariants from Exceptional Collections,”Commun. Math. Phys.385no. 1, (2021) 101–226,arXiv:2004.14466 [hep-th]

  49. [49]

    Layer Normalization,

    J. Lei Ba, J. R. Kiros, and G. E. Hinton, “Layer Normalization,”arXiv e-prints(July,

  50. [50]

    Rectifier Nonlinearities Improve Neural Network Acoustic Models,

    A. L. Maas, “Rectifier Nonlinearities Improve Neural Network Acoustic Models,” in . 2013

  51. [51]

    Demystifying Oversmoothing in Attention-Based Graph Neural Networks,

    X. Wu, A. Ajorlou, Z. Wu, and A. Jadbabaie, “Demystifying Oversmoothing in Attention-Based Graph Neural Networks,”arXiv e-prints(May, 2023) arXiv:2305.16102,arXiv:2305.16102 [cs.LG]. 72

  52. [2016]

    arXiv:1607.06450,arXiv:1607.06450 [stat.ML]