Pith. sign in

REVIEW 2 major objections 3 minor 51 references

Scattering amplitudes can be treated as executable programs, and an external generate-evaluate-select loop with coding agents can search this space to discover hybrid representations that are far faster and less cancellation-prone than stan

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-02 06:11 UTC pith:A5QBVV43

load-bearing objection Serious, well-audited demonstration that program search can cross amplitude representations, though the causal role of the evolutionary loop is untested by a missing no-evolution control. the 2 major comments →

arxiv 2607.21629 v1 pith:A5QBVV43 submitted 2026-07-14 hep-ph hep-th

Scattering Amplitudes as Programs: Self-Evolving Search for Theory and Event Generation

classification hep-ph hep-th
keywords scattering amplitudesprogram searchevolutionary computationcoding agentsBCFW recursionsplit-helicity amplitudesR-invariantsmatrix-element generation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper argues that a scattering amplitude is not a fixed formula but a program, and that equivalent programs for the same physics can differ enormously in speed, scaling, and numerical stability. It embeds repository-scale coding agents inside an external loop where a frozen evaluator checks correctness and scores objectives, letting a population of amplitude programs evolve. Three searches support the claim: a scaling search moves from BCFW recursion to a split-helicity transfer dynamic program with an 805x geometric-mean speed-up; a structural search reorganizes NMHV terms into R-invariant-style cells and glued supercells, cutting spurious cancellation; and an exact-operation search across twenty QCD and electroweak processes reduces counted arithmetic 47.7x and post-hoc Python runtime 5.9x. In matched pure-gluon tests from n=4 to n=6, the evolved exact-sum kernel is 277x faster than the tested Sherpa-Comix call and within 13.7x of process-specialised compiled MadGraph Fortran at -O3. The paper presents this as initial evidence that parts of amplitude optimisation can be made systematic, while deferring native generator integration and end-to-end event throughput as future tests.

Core claim

The central claim is that an externally evaluated evolutionary search over complete executable amplitude programs can traverse qualitatively different analytic representations, not just tune implementation constants. Starting from transparent seeds (BCFW recursion, CSW/MHV-vertex terms, and a Berends-Giele matrix-element engine), the search discovers: (1) a fixed-negative-helicity split-helicity zigzag sum reorganised as a finite-state transfer dynamic program, with source-level O(k^2 n) work and O(n) memory, giving an 805x geometric-mean speed-up on the 34-cell scored grid; (2) an NMHV term decomposition in which neighbouring R-invariant-style cells that share a spurious denominator are poi

What carries the argument

The core object is the executable amplitude program P(x) mapping physics input (theory, process, multiplicity, helicity, phase-space point, couplings) to an amplitude or squared matrix element; the search treats any analytic representation, recursion, colour/helicity organisation, and reuse scheme as a program to be mutated. The search machinery is an external generate-evaluate-select loop: a repository-scale coding agent proposes edited programs inside a sandbox, a frozen evaluator applies hard numerical-correctness gates and a declared objective (scaling fit, summation-conditioning proxy, or exact operation count), and a population with islands and archives retains successful lineages. The

Load-bearing premise

The correctness of every evolved program is certified only by agreement with author-supplied reference implementations at a finite set of phase-space points and helicity configurations, so the entire edifice rests on the assumption that no unobserved kinematic region hides a numerical error.

What would settle it

Take any evolved champion and evaluate it at a fresh phase-space point outside its validation set using an independent high-precision implementation (for example, 40-digit arithmetic for the scaling run); a relative disagreement above the stated tolerance, or a measured scaling exponent on the specialised split-helicity route that exceeds the claimed O(n) at fixed k, would refute the central claims. For the structural gluing, a symbolic polynomial-divisibility check at the shared boundary that fails would refute the claimed cancellation beyond the numerically tested trajectory.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • If the central claim holds, amplitude optimisation becomes a search problem: a fixed objective and a frozen correctness gate can induce changes of mathematical representation without any supplied target implementation.
  • For fixed negative-helicity count, tree-level gluon amplitudes can be evaluated in linear time in multiplicity using the split-helicity transfer path, so high-multiplicity tails in exact-sum kernels become substantially cheaper than BCFW-type recursion.
  • The structural result implies that inter-term cancellation near spurious boundaries can be reduced by gluing cells with common denominator factors, cutting the 95th-percentile summation condition number roughly in half and reducing term counts from 8-13 to 2-3.
  • The 47.7x exact-operation reduction is phase-space invariant by construction (the counted programs cannot branch on kinematics), so the aggregate arithmetic saving applies across all kinematical points, not only benchmark configurations.
  • The matched pure-gluon timing comparison implies that exact full-colour, full-helicity matrix-element kernels in current production generators contain at least an order-of-magnitude headroom, with the gap to process-specialised compiled Fortran narrowing from 77.7x at n=4 to 13.7x at n=6.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The search's rediscovery of split-helicity transfer and R-invariant-style organisations without being seeded with those formulae suggests that many historical 'analytic breakthroughs' in amplitude theory are discoverable by evaluator-guided exploration; a direct test would be to run the same machinery on N^2MHV or higher sectors with no additional hints.
  • The structural gluing is verified only by high-precision numerical trajectories near named spurious boundaries, not by a symbolic polynomial-divisibility certificate; a symbolic proof of the cell-gluing identity would strengthen the result to a theorem, while any counterexample along the boundary hypersurface would bound the method's reach.
  • The generator comparison isolates one exact-sum kernel and explicitly excludes phase-space integration, colour/helicity sampling, matching, and showering; the practical conclusion is conditional on whether the same arithmetic reductions survive native integration and end-to-end event throughput, which the paper identifies as the next decisive experiment.
  • A testable extension is to port the evolved engine's arithmetic graph into a compiled code-generation backend (similar to how MadGraph's generated source is compiled at -O0/-O2/-O3) and measure whether the 277x-vs-Sherpa and 13.7x-vs-MadGraph gaps translate to real event-generation wall-clock speedups.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 3 minor

Summary. The paper develops a framework in which scattering amplitudes are treated as executable programs and searched over by LLM-based 'repository-scale coding agents' embedded in an external generate–evaluate–select loop with frozen evaluators and objectives. Three searches are reported: (i) a scaling search that evolves a BCFW seed into a fixed-k split-helicity transfer dynamic program, reporting an 805x geometric-mean speed-up on the scored grid; (ii) a structural search that reorganises NMHV rational expressions into R-invariant-style cells and glued supercells, reducing a cancellation-conditioning score from 0.3758 to 0.6048 and term counts from 8–13 to 2–3; and (iii) a tree-level matrix-element search over twenty QCD/EW processes that reduces the exactly counted real-operation total by 47.7x and yields a 5.9x post-hoc Python runtime improvement, with matched pure-gluon comparisons against Sherpa-Comix and MadGraph at n=4,5,6. The claims are carefully scoped: the generator comparison is an isolated exact-sum kernel call, not an end-to-end event-generation benchmark.

Significance. If the empirical results hold, the paper provides initial evidence that program search can move between mathematical representations of scattering amplitudes and assemble hybrid algorithms that would be unlikely to emerge from low-level tuning alone. The reported audits are unusually thorough for this type of work: fresh-phase-space checks, exact operation-count invariance under rejection of numerical branching, high-precision mpmath anchors, tolerance-ladder justification, cost-reweighting robustness, independent holdout samples, and controlled boundary trajectories. The frozen archives and versioned reproductions strengthen the reliability of the computational claims. The central weakness is causal attribution: the paper argues that the external population dynamics provide the selection pressure, but no no-evolution control is reported, so the respective roles of the LLM's pretrained knowledge and the evolutionary loop are not established.

major comments (2)
  1. [§3, Figure 1, Table 1] The claim 'The coding agent alone is not the experiment; the frozen ruler and population dynamics provide the selection pressure' is asserted but not demonstrated. No control is reported in which the same coding agent, seed, frozen evaluator, prompts, and budget are run without the external population/island/inspiration structure, e.g. a single-lineage iterative edit–test–repair session. Given the high submission pass rate (114/115) and the visible local tests, the agent may be competent enough on its own that much of the observed gain comes from LLM priors rather than from the generate–evaluate–select loop. This directly affects the paper's 'self-evolving' attribution. I recommend either adding such a control or substantially softening the causal language to describe an agentic program-search system whose components are not individually ablated.
  2. [§4.2, Appendix C] The structural claim that shared-boundary gluing 'removes' spurious cancellation is supported only by finite-point numerical agreement and controlled one-parameter trajectories; the released implementation constructs the quotient by pointwise division and does not provide a symbolic polynomial-divisibility certificate, as the paper itself states. The controlled six-point check (Figure 8, Appendix C) is strong numerical evidence, but the central assertion that the representation has less inter-term cancellation would be materially strengthened by a symbolic proof for the identified six-point boundary or, failing that, an explicit statement that the result is a numerical observation rather than an algebraic theorem. Since the paper is otherwise unusually careful about evidence, this limitation should be foregrounded in the main text, not only in an appendix.
minor comments (3)
  1. [§4.2] The notation '(0,2,4)' for negative-helicity positions is used several times before being explained; a one-sentence definition (zero-based positions of the negative-helicity legs) would improve readability.
  2. [Table 2] The table footnote refers to Q_0.95 but the definition appears only later in §4.2/Appendix B; please define it at first use in the table caption.
  3. [§4.1] The statement that the generation-47 score equals 0.75(2.8916)+0.25 log(1+804.7) is helpful, but the geometric-mean speed-up is elsewhere quoted as 804.7 and 805x; use a consistent number of significant digits throughout.

Circularity Check

0 steps flagged

No circularity: the reported gains are the frozen declared objectives or post-hoc external checks, not fitted quantities relabelled as predictions.

full rationale

The derivation chain is self-contained. The physical amplitudes are fixed by the theory, and every evolved program must pass numerical-correctness gates against independent references: Berends–Giele (itself checked against Parke–Taylor), MadGraph5_aMC@NLO, and frozen project engines previously validated point-by-point against MadGraph. The headline numbers are the declared selection objectives, not quantities fitted afterward and then presented as predictions: the scaling run selects on fitted exponents and geometric-mean speed, the structural run on the summation-conditioning proxy of Eq. (20), and the tree-complexity run on exactly counted arithmetic operations. The paper explicitly separates selection metrics from post-hoc evidence, e.g. the tree-complexity runtime 'was neither selected nor shown to the agent' (Section 4.3), and it reports fresh-point audits in Appendix B and a controlled boundary analysis in Appendix C. The structural gluing is candidly described as numerically verified along controlled trajectories, not a symbolic polynomial-divisibility certificate, and the 10^-81 agreement is correctly labelled as a same-arithmetic algebraic comparison rather than an independent 80-digit validation. The only self-citation, [31], is the companion repository storing frozen artefacts and is a data-availability disclosure, not a load-bearing premise. The skeptic's missing no-evolution control concerns causal attribution—whether gains come from the external loop or from LLM priors—but it is not a definitional reduction of a claimed result to its inputs. No circular step can be exhibited from the paper's equations or citation chain.

Axiom & Free-Parameter Ledger

4 free parameters · 6 axioms · 0 invented entities

No new physical entities are postulated. The 'free parameters' are experimental design choices (score weights, tolerances, cost conventions); the paper's robustness checks show the headline factors survive reasonable variations. The axioms are standard amplitude-theory background plus the practical assumptions of finite-point certification and a frozen evaluator.

free parameters (4)
  • Scaling score weights = 0.75 S_scaling + 0.25 S_speed
    Chosen by hand in Eq. (19) to balance scaling and speed; determines which champion is selected. The 805x result is under this weighting, but fresh-point audits confirm the champion remains fastest.
  • Structural score parameters = 0.95 quantile, 10^-3 drop threshold, Spole formula
    Define the cancellation objective; the champion score 0.6048 depends on these choices. Holdout audits show the improvement persists on unseen points.
  • Tree-complexity operation cost weights = Equal unit weights
    Default convention in Eq. (21); reweighting study (Table 6) shows aggregate reductions 47.7-57.2x and geometric means 11.7-12.4x, so the central claim is robust.
  • Scaling gate tolerance ladder = 10^-6 to 3x10^-3
    Loosened at high n; justified by a 40-digit mpmath anchor bounding reference error below 1.0x10^-8, leaving headroom at every rung.
axioms (6)
  • standard math Spinor-helicity formalism and colour-ordered decomposition correctly represent QCD tree amplitudes
    Background for all benchmarks; used in Eq. (4) of Section 2.1.
  • standard math Berends–Giele recursion and BCFW recursion are correct references for tree gluon amplitudes
    Used as correctness references and seed programs in Sections 2.2 and 4.1.
  • domain assumption Finite-point numerical matching at tolerance 10^-6 to references implies program correctness
    The evolved programs are certified only on finitely many phase-space points (Section 2.3); the paper assumes no hidden bugs outside the test set.
  • standard math The split-helicity zigzag sum of Ref. [46] and its reorganisations are correct
    The scaling champion implements and reorganises this known formula (Section 4.1).
  • domain assumption The external evaluator and objective are frozen and authoritative
    The self-evolving claim relies on the separation between proposal and evaluation (Section 3).
  • domain assumption Timing measurements on a single laptop are representative
    All wall-clock times used one MacBook Pro; background load was not logged (Section 5).

pith-pipeline@v1.3.0-alltime-deepseek · 25748 in / 12634 out tokens · 221685 ms · 2026-08-02T06:11:52.060880+00:00 · methodology

0 comments
read the original abstract

By viewing scattering amplitudes as computer programs, we connect two goals: exposing useful analytic structure and constructing efficient numerical evaluators for collider phenomenology. Equivalent programs can differ sharply in multiplicity scaling, arithmetic complexity, cancellation, and runtime. Amplitude calculation therefore defines a structured search problem over analytic representations, recursive algorithms, colour and helicity organisation, and reuse of intermediate objects. We embed repository-scale coding agents inside an external generate-evaluate-select loop with frozen evaluators and objectives, and study three optimisation targets. A scaling search moves from BCFW recursion to a specialised fixed-k split-helicity transfer algorithm, reaching an 805x geometric-mean speed-up on the scored grid. A structural search reorganises NMHV terms into R-invariant-style cells and glued supercells, reducing inter-term cancellation. Across twenty QCD and electroweak processes, an exact-operation search reduces counted arithmetic by 47.7x and gives a 5.9x post-hoc Python runtime improvement. In matched pure-gluon component tests from n=4 to n=6, the evolved engine's speed advantage over the tested Sherpa-Comix exact-sum call grows from 10x to 277x, while its gap to process-specialised MadGraph5_aMC@NLO Fortran compiled at -O3 narrows from 77.7x to 13.7x. The searches move between mathematical representations and combine recursion, symmetry, basis reduction, dynamic programming, and shared computation into hybrid amplitude programs. They provide initial evidence that parts of amplitude optimisation can be made systematic through self-evolving program search, while leaving native generator integration and end-to-end event throughput as future tests. The generator comparison is an isolated exact matrix-element call, not a modification of MadGraph or Sherpa.

Figures

Figures reproduced from arXiv: 2607.21629 by Sven Krippendorf, Yi Gu.

Figure 1
Figure 1. Figure 1: Architecture of the externally evaluated self-evolving program-search system. A repository [PITH_FULL_IMAGE:figures/full_fig_p008_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Multiplicity scaling along the completed fifty-generation run. Each point is the median wall [PITH_FULL_IMAGE:figures/full_fig_p013_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Evolution of the cancellation-conditioning score [PITH_FULL_IMAGE:figures/full_fig_p014_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Exactly counted real arithmetic operations [PITH_FULL_IMAGE:figures/full_fig_p016_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Compiler-controlled pure-gluon matrix-element timing comparison for [PITH_FULL_IMAGE:figures/full_fig_p018_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Kinematic robustness of the scaling champion. Left: per-cell speed-up over the seed at ten [PITH_FULL_IMAGE:figures/full_fig_p023_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Controlled approaches to named denominators in the structural representations. Left: the CSW [PITH_FULL_IMAGE:figures/full_fig_p025_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: High-precision pointwise test of the six-point gluing operation. The magnitudes [PITH_FULL_IMAGE:figures/full_fig_p026_8.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

51 extracted references · 38 linked inside Pith

  1. [1]

    The automated computation of tree-level and next-to-leading order 26 differential cross sections, and their matching to parton shower simulations,

    J. Alwall, R. Frederix, S. Frixione, V. Hirschi, F. Maltoni, O. Mattelaer, H. S. Shao, T. Stelzer, P. Torrielli, and M. Zaro, “The automated computation of tree-level and next-to-leading order 26 differential cross sections, and their matching to parton shower simulations,”JHEP07(2014) 079, arXiv:1405.0301 [hep-ph]

  2. [2]

    Event generation with SHERPA 1.1

    T. Gleisberg, S. H¨ oche, F. Krauss, M. Sch¨ onherr, S. Schumann, F. Siegert, and J. Winter, “Event generation with SHERPA 1.1”JHEP02(2009) 007,arXiv:0811.4622 [hep-ph]

  3. [3]

    Event Generators for High-Energy Physics Experiments,

    J. M. Campbellet al., “Event Generators for High-Energy Physics Experiments,”SciPost Phys. 16no. 5, (2024) 130,arXiv:2203.11110 [hep-ph]

  4. [4]

    ATLAS Software and Computing HL-LHC Roadmap,

    ATLAS Collaboration, “ATLAS Software and Computing HL-LHC Roadmap,” Tech. Rep. CERN-LHCC-2022-005, CERN, 2022. https://cds.cern.ch/record/2802918/files/LHCC-G-182.pdf

  5. [5]

    CMS Phase-2 Computing Model: Update Document,

    CMS Offline Software and Computing, “CMS Phase-2 Computing Model: Update Document,” Tech. Rep. CMS-NOTE-2022-008, CERN, 2022. https://cds.cern.ch/record/2815292/files/NOTE2022_008.pdf

  6. [6]

    Challenges in Monte Carlo Event Generator Software for High-Luminosity LHC,

    A. Valassiet al., “Challenges in Monte Carlo Event Generator Software for High-Luminosity LHC,”Computing and Software for Big Science5no. 1, (2021) 12,arXiv:2004.13687 [hep-ph]

  7. [7]

    Modelling and computational improvements to the simulation of single vector-boson plus jet processes for the ATLAS experiment,

    ATLAS Collaboration, “Modelling and computational improvements to the simulation of single vector-boson plus jet processes for the ATLAS experiment,”JHEP08(2022) 089, arXiv:2112.09588 [hep-ex]

  8. [8]

    Accelerating LHC event generation with simplified pilot runs and fast PDFs,

    E. Bothmann, A. Buckley, I. A. Christidi, C. G¨ utschow, S. H¨ oche, M. Knobbe, T. Martin, and M. Sch¨ onherr, “Accelerating LHC event generation with simplified pilot runs and fast PDFs,”Eur. Phys. J. C82(2022) 1128,arXiv:2209.00843 [hep-ph]

  9. [9]

    A Portable Parton-Level Event Generator for the High-Luminosity LHC,

    E. Bothmann, T. Childers, W. Giele, S. H¨ oche, J. Isaacson, and M. Knobbe, “A Portable Parton-Level Event Generator for the High-Luminosity LHC,”SciPost Phys.17(2024) 081, arXiv:2311.06198 [hep-ph]

  10. [10]

    Data-parallel leading-order event generation in MadGraph5 aMC@NLO,

    S. Hageb¨ ock, D. Massaro, O. Mattelaer, S. Roiser, A. Valassi, and Z. Wettersten, “Data-parallel leading-order event generation in MadGraph5 aMC@NLO,”arXiv:2507.21039 [hep-ph]

  11. [11]

    Spinor Techniques for Calculatingp¯p→W ±/Z0 + Jets,

    R. Kleiss and W. J. Stirling, “Spinor Techniques for Calculatingp¯p→W ±/Z0 + Jets,”Nucl. Phys. B262(1985) 235–262

  12. [12]

    Calculating scattering amplitudes efficiently,

    L. J. Dixon, “Calculating scattering amplitudes efficiently,” inTheoretical Advanced Study Institute in Elementary Particle Physics (TASI 95): QCD and Beyond, pp. 539–584. 1996. arXiv:hep-ph/9601359

  13. [13]

    Duality and Multi - Gluon Scattering,

    M. L. Mangano, S. J. Parke, and Z. Xu, “Duality and Multi - Gluon Scattering,”Nucl. Phys. B 298(1988) 653–672

  14. [14]

    Recursive Calculations for Processes with n Gluons,

    F. A. Berends and W. T. Giele, “Recursive Calculations for Processes with n Gluons,”Nucl. Phys. B306(1988) 759–808

  15. [15]

    New recursion relations for tree amplitudes of gluons,

    R. Britto, F. Cachazo, and B. Feng, “New recursion relations for tree amplitudes of gluons,”Nucl. Phys. B715(2005) 499–522,arXiv:hep-th/0412308

  16. [16]

    Direct proof of tree-level recursion relation in Yang-Mills theory,

    R. Britto, F. Cachazo, B. Feng, and E. Witten, “Direct proof of tree-level recursion relation in Yang-Mills theory,”Phys. Rev. Lett.94(2005) 181602,arXiv:hep-th/0501052

  17. [17]

    One loop n point gauge theory amplitudes, unitarity and collinear limits,

    Z. Bern, L. J. Dixon, D. C. Dunbar, and D. A. Kosower, “One loop n point gauge theory amplitudes, unitarity and collinear limits,”Nucl. Phys. B425(1994) 217–260, arXiv:hep-ph/9403226

  18. [18]

    The Amplituhedron,

    N. Arkani-Hamed and J. Trnka, “The Amplituhedron,”JHEP10(2014) 030,arXiv:1312.2007 [hep-th]

  19. [19]

    Transforming the Bootstrap: Using Transformers to Compute Scattering Amplitudes in Planar N= 4 Super Yang-Mills Theory,

    T. Cai, G. W. Merz, F. Charton, N. Nolte, M. Wilhelm, K. Cranmer, and L. J. Dixon, “Transforming the Bootstrap: Using Transformers to Compute Scattering Amplitudes in Planar N= 4 Super Yang-Mills Theory,”Mach. Learn. Sci. Tech.5no. 3, (2024) 035073, arXiv:2405.06107 [cs.LG]. 27

  20. [20]

    Recurrent Features of Amplitudes in PlanarN= 4 Super Yang-Mills Theory,

    T. Cai, F. Charton, K. Cranmer, L. J. Dixon, G. W. Merz, and M. Wilhelm, “Recurrent Features of Amplitudes in PlanarN= 4 Super Yang-Mills Theory,”JHEP04(2025) 143, arXiv:2501.05743 [hep-th]

  21. [21]

    Simplifying Polylogarithms with Machine Learning,

    A. Dersy, M. D. Schwartz, and X. Zhang, “Simplifying Polylogarithms with Machine Learning,” Int. J. Data Sci. Math. Sci.01no. 02, (2023) 135–179,arXiv:2206.04115 [cs.LG]

  22. [22]

    Learning the Simplicity of Scattering Amplitudes,

    C. Cheung, A. Dersy, and M. D. Schwartz, “Learning the Simplicity of Scattering Amplitudes,” SciPost Phys.18(2025) 040,arXiv:2408.04720 [hep-th]

  23. [23]

    An LLM-driven framework for cosmological model-building and exploration,

    N. Mudur, C. Cuesta-Lazaro, M. W. Toomey, and D. P. Finkbeiner, “An LLM-driven framework for cosmological model-building and exploration,” inLLM for Scientific Discovery: Reasoning, Assistance, and Collaboration (LM4Sci at COLM). 2025. https://openreview.net/forum?id=xnPJOtPmK3

  24. [24]

    DiscoverPhysics: Benchmarking LLMs for Out-of-the-Box Scientific Thinking,

    M. L. Wiemann, L. M. Smith, P. Melchior, S. Mishra-Sharma, A. G. Wilson, P. Izmailov, and C. Cuesta-L´ azaro, “DiscoverPhysics: Benchmarking LLMs for Out-of-the-Box Scientific Thinking,”arXiv:2605.26087 [stat.ML]

  25. [25]

    MadEvolve: Evolutionary Optimization of Cosmological Algorithms with Large Language Models,

    T. Li, S. Zang, and M. M¨ unchmeyer, “MadEvolve: Evolutionary Optimization of Cosmological Algorithms with Large Language Models,”arXiv:2602.15951 [astro-ph.CO]

  26. [26]

    Single-minus gluon tree amplitudes are nonzero,

    A. Guevara, A. Lupsasca, D. Skinner, A. Strominger, and K. Weil, “Single-minus gluon tree amplitudes are nonzero,”arXiv:2602.12176 [hep-th]

  27. [27]

    Surface Water Wave Scattering and the Hydrotope,

    N. Arkani-Hamed, F. Calisto, N. Ussembayev, W. W. Zhao, and Z. Zhou, “Surface Water Wave Scattering and the Hydrotope,”arXiv:2606.28280 [hep-th]

  28. [28]

    Explainable AI-assisted Optimization for Feynman Integral Reduction,

    Z.-Y. Song, T.-Z. Yang, Q.-H. Cao, M.-x. Luo, and H. X. Zhu, “Explainable AI-assisted Optimization for Feynman Integral Reduction,”JHEP06(2026) 225,arXiv:2502.09544 [hep-ph]

  29. [29]

    Refining Integration-by-Parts Reduction of Feynman Integrals with Machine Learning,

    M. von Hippel and M. Wilhelm, “Refining Integration-by-Parts Reduction of Feynman Integrals with Machine Learning,”JHEP05(2025) 185,arXiv:2502.05121 [hep-th]

  30. [30]

    Efficient AI-Inspired Reduction of Feynman Integrals via Tube Seeding,

    J. Berman, F. Charton, A. Luna, M. Wilhelm, and M. Zeng, “Efficient AI-Inspired Reduction of Feynman Integrals via Tube Seeding,”arXiv:2606.10698 [hep-ph]

  31. [31]

    Scattering Amplitudes as Programs: Frozen Release Artefacts

    Y. Gu and S. Krippendorf, “Scattering Amplitudes as Programs: Frozen Release Artefacts.” Github repository, commit ab8c8682713ff8592bb55e456229619bff7cfd12, 2026. https://github.com/YiGu310/scattering-amplitudes-program-search/tree/ ab8c8682713ff8592bb55e456229619bff7cfd12

  32. [32]

    Multiparton amplitudes in gauge theories,

    M. L. Mangano and S. J. Parke, “Multiparton amplitudes in gauge theories,”Phys. Rept.200 (1991) 301–367,arXiv:hep-th/0509223

  33. [33]

    A brief introduction to modern amplitude methods,

    L. J. Dixon, “A brief introduction to modern amplitude methods,” inTheoretical Advanced Study Institute in Elementary Particle Physics: Particle Physics: The Higgs Boson and Beyond, pp. 31–67. 2014.arXiv:1310.5353 [hep-ph]

  34. [34]

    Elvang and Y.-t

    H. Elvang and Y.-t. Huang,Scattering Amplitudes in Gauge Theory and Gravity. Cambridge University Press, 2015.arXiv:1308.1697 [hep-th]

  35. [35]

    High-precision calculation of multiloop Feynman integrals by difference equations,

    S. Laporta, “High-precision calculation of multiloop Feynman integrals by difference equations,” Int. J. Mod. Phys. A15(2000) 5087–5159,arXiv:hep-ph/0102033

  36. [36]

    A General algorithm for calculating jet cross-sections in NLO QCD,

    S. Catani and M. H. Seymour, “A General algorithm for calculating jet cross-sections in NLO QCD,”Nucl. Phys. B485(1997) 291–419,arXiv:hep-ph/9605323. [Erratum: Nucl.Phys.B 510, 503–504 (1998)]

  37. [37]

    An Amplitude fornGluon Scattering,

    S. J. Parke and T. R. Taylor, “An Amplitude fornGluon Scattering,”Phys. Rev. Lett.56(1986) 2459

  38. [38]

    Perturbative gauge theory as a string theory in twistor space,

    E. Witten, “Perturbative gauge theory as a string theory in twistor space,”Commun. Math. Phys. 252(2004) 189–258,arXiv:hep-th/0312171. 28

  39. [39]

    MHV vertices and tree amplitudes in gauge theory,

    F. Cachazo, P. Svrcek, and E. Witten, “MHV vertices and tree amplitudes in gauge theory,” JHEP09(2004) 006,arXiv:hep-th/0403047

  40. [40]

    New Relations for Gauge-Theory Amplitudes,

    Z. Bern, J. J. M. Carrasco, and H. Johansson, “New Relations for Gauge-Theory Amplitudes,” Phys. Rev. D78(2008) 085011,arXiv:0805.3993 [hep-ph]

  41. [41]

    A New Monte Carlo Treatment of Multiparticle Phase Space at High-Energies,

    R. Kleiss, W. J. Stirling, and S. D. Ellis, “A New Monte Carlo Treatment of Multiparticle Phase Space at High-Energies,”Comput. Phys. Commun.40(1986) 359–373

  42. [42]

    Mathematical discoveries from program search with large language models,

    B. Romera-Paredes, M. Barekatain, A. Novikov, M. Balog, M. P. Kumar,et al., “Mathematical discoveries from program search with large language models,”Nature625(2024) 468–475

  43. [43]

    AlphaEvolve: A coding agent for scientific and algorithmic discovery,

    A. Novikov, N. Vu, M. Eisenberger, E. Dupont, P.-S. Huang, A. Z. Wagner, S. Shirobokov, B. Kozlovskii, F. J. R. Ruiz, A. Mehrabian, M. P. Kumar, A. See, S. Chaudhuri, G. Holland, A. Davies, S. Nowozin, P. Kohli, and M. Balog, “AlphaEvolve: A coding agent for scientific and algorithmic discovery,”arXiv:2506.13131 [cs.AI]

  44. [44]

    AdaEvolve: Adaptive LLM Driven Zeroth-Order Optimization,

    M. Cemri, S. Agrawal, A. Gupta, S. Liu, A. Cheng, Q. Mang, A. Naren, L. E. Erdogan, K. Sen, M. Zaharia, A. Dimakis, and I. Stoica, “AdaEvolve: Adaptive LLM Driven Zeroth-Order Optimization,”arXiv:2602.20133 [cs.NE]

  45. [45]

    ShinkaEvolve: Towards Open-Ended And Sample-Efficient Program Evolution,

    R. T. Lange, Y. Imajuku, and E. Cetin, “ShinkaEvolve: Towards Open-Ended And Sample-Efficient Program Evolution,”arXiv:2509.19349 [cs.CL]

  46. [46]

    All Split Helicity Tree-Level Gluon Amplitudes,

    R. Britto, B. Feng, R. Roiban, M. Spradlin, and A. Volovich, “All Split Helicity Tree-Level Gluon Amplitudes,”Phys. Rev. D71(2005) 105017,arXiv:hep-th/0503198

  47. [47]

    Eliminating spurious poles from gauge-theoretic amplitudes,

    A. Hodges, “Eliminating spurious poles from gauge-theoretic amplitudes,”JHEP05(2013) 135, arXiv:0905.1473 [hep-th]

  48. [48]

    A Note on Polytopes for Scattering Amplitudes,

    N. Arkani-Hamed, J. L. Bourjaily, F. Cachazo, A. Hodges, and J. Trnka, “A Note on Polytopes for Scattering Amplitudes,”JHEP04(2012) 081,arXiv:1012.6030 [hep-th]

  49. [49]

    Comix, a new matrix element generator,

    T. Gleisberg and S. H¨ oche, “Comix, a new matrix element generator,”JHEP12(2008) 039, arXiv:0808.3674 [hep-ph]

  50. [50]

    Speeding up MadGraph5 aMC@NLO,

    O. Mattelaer and K. Ostrolenk, “Speeding up MadGraph5 aMC@NLO,”Eur. Phys. J. C81no. 5, (2021) 435,arXiv:2102.00773 [hep-ph]

  51. [51]

    HL-LHC Computing Review Stage-2, Common Software Projects: Event Generators,

    HSF Physics Event Generator Working Group, “HL-LHC Computing Review Stage-2, Common Software Projects: Event Generators,”arXiv:2109.14938 [hep-ph]. 29