REVIEW 3 major objections 7 minor 1 cited by
A cheap consistency certificate for Seiberg duality lets language models repair broken claims, but which search policy wins depends on the model.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-30 17:35 UTC pith:R6JSBELH
load-bearing objection Clean preregistered evidence that a cheap QFT consistency certificate actually moves LM repair, with a real model-dependent flip in how you should use it. the 3 major comments →
DualityCert: Verifier-Gated Language-Model Repair of Broken Duality Claims in Quantum Field Theory
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
On a preregistered benchmark of 145 broken Seiberg-duality claims, verifier-gated retry improves final repair success over a single attempt by +8.3 percentage points on one confirmatory model and +7.1 on the other. Under an equal eleven-attempt budget, a stop-first strategy portfolio loses to independent verifier-filtered resampling on one model and beats it on the other, while every winning policy relies on the same consistency certificate. Feedback content helps only on one of the two models.
What carries the argument
DualityCert consistency certificate: a symbolic verdict that a candidate electric–magnetic quiver pair passes a fixed registry of obligations (’t Hooft anomalies, superpotential R-consistency, trial central-charge match, bounded classical chiral-ring proxy) with none failing—used both as interaction-time feedback and as a final judge at least as strict as what the agent saw.
Load-bearing premise
Success under this certificate is taken to measure real verifier-exploitation skill, even though the checks are a finite, incomplete proxy that can miss physics the encoding does not cover.
What would settle it
Re-run the same locked 145-fixture protocol on the two confirmatory models and check whether the iteration gain and the E4 sign reversal between portfolio and best-of-n disappear, or whether depth-two fixtures remain near floor for every policy.
If this is right
- Language-model agents can be scored and compared on physics claims by native consistency certificates rather than by answer matching or full formal proofs.
- Verifier-gated iteration is a reliable gain on this benchmark; the better budget-matched search policy is not universal across models.
- The same schema–obligation–certificate–feedback–strict-judge pattern can be reused for other dualities and consistency-driven areas of QFT and string theory.
- Releasing the verifier, fixtures, protocol, and per-attempt records makes the policy comparison independently regenerable.
Where Pith is reading between the lines
- Model-specific sensitivity to obligation names versus resampling suggests pairing each model with its preferred exploitation policy rather than shipping one default loop.
- Because depth-two repairs sit near the floor, useful next benchmarks likely need richer obligations or multi-edit operators, not only more attempts.
- Areas with sharp protected quantities—bootstrap bounds, amplitude factorization, tadpole cancellation—are natural next certificate targets under the same design pattern.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces DualityCert, a symbolic verifier that evaluates candidate Seiberg-duality claims (ordered quiver pairs) against a registry of consistency obligations: 't Hooft anomaly matching, superpotential R-charge consistency, trial central-charge matching, and a bounded classical chiral-ring proxy. Passing claims receive a consistency certificate, explicitly scoped as "no tested inconsistency found," not a proof. The verifier is used as a repair environment for LLM agents on a preregistered benchmark of 145 depth-one-perturbed dual pairs from six toric seed families. With the analysis frozen (commit 813fdb0) before confirmatory calls, the authors report: content-free verifier-gated retry beats single-shot by +8.3 pp (DeepSeek-Chat) and +7.1 pp (Qwen-Plus), Holm-adjusted p<0.002; under an equal eleven-attempt budget, the stop-first strategy portfolio vs independent verifier-filtered resampling reverses sign between the two models (−10.3 pp vs +14.7 pp); feedback-content gains (E1, E5) appear only on Qwen-Plus; a preregistered MiniMax-M2.5 extension reproduces the iteration gain and the negative E4 ordering. Statistics use fixture-clustered binomial GEE with Holm families. Code, fixtures, manifests, quarantine ledgers, and per-attempt records are released.
Significance. If the results hold, this is a methodologically strong demonstration of verifier-gated LLM repair in a domain without proof assistants, and an honestly reported non-universality result (the E4 sign reversal) that cautions against exporting verifier-exploitation policies across models. The empirical hygiene is well above the norm for this literature: protocol freeze before confirmatory calls, preregistered endpoints and multiplicity control, an equal-budget E4 comparison constructed by deterministic replay over complete ss/gr/vf components, prespecified contamination rules with byte-exact quarantine and SHA-256 ledgers, per-replication sign consistency reported alongside the pooled estimates, and full artifact release enabling regeneration of every table. The certificate semantics are stated with appropriate modesty, and the masked-feedback control (E5) is a genuinely informative design element for separating semantic feedback content from retry structure. The physics scope is narrow (one duality class, bounded classical proxies, no a-maximization, no Witten-anomaly check), but the authors disclose each gap explicitly and do not overclaim.
major comments (3)
- [§3, §B.2 (benchmark selection vs. evaluation criterion)] The benchmark is closed under the verifier that judges success: fixtures are kept only if the perturbed pair fails in scope while the seed certifies (52 candidates dropped for label mismatch, App. B.2), and repair success is certification by the same registry. The comparative policy claims (E1/E2/E4/E5) are internally valid under this construction, but the absolute success levels, and the word 'repair' in the title and abstract, inherit it: the study measures repair of verifier-detectable depth-one breakage. This should be stated explicitly in §3 and the abstract. A cheap, concrete out-of-selection check is available: the registry already contains opt-in obligations not used in fixture selection (e.g., index matching, Table 2); running them on the certified repairs would test whether 'repair' transfers beyond the selection surface.
- [§3 and App. C (multiplicity family structure)] The text describes 'one paper-wide Holm adjustment,' but the primary family covers only the six {E1,E2,E4}×{models} tests; E5 is a separate two-hypothesis family and the MiniMax extension a separate three-hypothesis family (Table 1 note, §C). The paper therefore contains 11 confirmatory tests with FWER controlled only within families, so 'paper-wide' overstates the control. This is likely cosmetic — the Qwen E5 result (unadjusted p=0.0065) plausibly survives a joint Holm over the eight primary+E5 tests — but the family structure should be described accurately wherever the control guarantee is stated, or the families merged.
- [§2 (singlet fixtures: coincident verifier configs)] For the 34 singlet-containing fixtures the R-charge grading makes the interaction-time and final-judge configurations coincide (§2, App. B.1), so the 'final judge at least as strict' anti-gaming element is vacuous on ~23% of the benchmark; protection there rests on withheld residuals and the copy guard. The pooled endpoints mix both regimes. A preregistered-style sensitivity re-estimating E1, E2, and E4 on the 111 singlet-free fixtures would show that the central comparisons are not driven by the subset where the agent effectively sees the final judge. This is a robustness ask, not an objection to the disclosed design.
minor comments (7)
- [§2, Eq. (1)] As printed, the gaugino sum Σ_a (N_a²−1) lacks the (R−1)³ factor, which vanishes for gauginos (R=1); the formula as written is correct for Tr R but not for Tr R³, where the gaugino contribution is zero. Presumably a transcription/typography issue, but it should be corrected since Eq. (1) is the paper's one displayed computation.
- [App. F (model versioning)] Exact reproduction is impossible because the API endpoints drift even with pinned code and fixtures. Please record the provider model snapshot identifiers and campaign dates in App. F alongside the commit hashes, and state this limitation explicitly in the reproduction section.
- [§3 (E4 replay specification)] The E4 replay would benefit from two clarifications: (i) how the deterministic portfolio replay treats a stage whose attempt passed the interaction-time verifier (L=3) but fails the final judge (L=5) — is the stage credited, or does escalation continue?; (ii) whether best-of-n draws were judged online at final strictness ('stopping at the first final certificate') or rescored post hoc. The answers affect the interpretation of budget equality.
- [Table 4 (replication variance)] The single-shot counts for DeepSeek-Chat vary from 8/145 to 17/145 across replications — replication-level variance comparable to the E2 effect itself. One sentence noting that replication fixed effects absorb this, and that per-replication signs are nonetheless consistent (as claimed in §4), would pre-empt misreading of the GEE marginal estimates.
- [§B.1, Table 2 (certificate content)] Given that 12 of 23 registry obligations are NOT APPLICABLE / NOT IMPLEMENTED / UNKNOWN on positive fixtures, consider reporting the number of obligations actually judged per certificate (or noting it is constant at 11 across the benchmark) so readers can calibrate what certification meant per repair.
- [§2 (notation)] The schema node labels 'SU(2) 0, SU(2) 1, ...' are confusing on first reading — they suggest fixed SU(2) gauge factors while ranks in fact vary (seed ranks 2–4, Table 3). Clarify that 'SU(2)' here is a schema label, or rename.
- [Refs/Acknowledgments] Reference [1] is the author's own arXiv preprint; the Acknowledgments also disclose Claude Max access during early development. Both are fine, but the disclosure of AI assistance in development could be made slightly more specific (which parts of the pipeline it touched), given the paper's subject.
Circularity Check
No significant circularity: empirical policy contrasts against a fixed symbolic checker, not a self-defined prediction.
full rationale
The paper’s load-bearing claims are preregistered experimental risk differences (E1–E5) on LM repair success under DualityCert, not first-principles derivations of dualities. The verifier encodes standard, independently known consistency checks (’t Hooft anomalies, R(W)=2, trial central-charge matching, bounded classical chiral-ring proxy); the certificate is explicitly scoped as “no tested inconsistency was found, not that the duality is proven.” Benchmark fixtures are built by dualizing seed quivers then applying depth-one perturbations that the same registry labels failed while seeds certify—this makes the task well-posed (a repair exists and is needed) but does not force the measured success rates or the sign-reversing E4 contrast between models. No free parameters are fitted to outcomes and re-presented as predictions; GEE endpoints, Holm families, and contamination rules were frozen before confirmatory API calls. Self-citations (e.g. the author’s related LLM-physics evaluation work) are background only and not load-bearing for the policy results. The derivation chain is therefore self-contained empirical measurement against an external symbolic oracle, not circular reduction of outputs to inputs.
Axiom & Free-Parameter Ledger
free parameters (4)
- Interaction/final chiral-ring cutoffs L=3 / L=5 (singlet-free) =
L=3 interaction, L=5 final
- Repair round cap K=5 and best-of-n budget 2K+1=11 =
K=5, n=11
- Replication count R=3 =
R=3
- Depth-one perturbation magnitudes (e.g. R-shift −1/3, rank ±1) =
four operator classes; ranks 2–4 seeds
axioms (6)
- domain assumption If two 4d N=1 theories are IR-dual, protected 't Hooft anomalies, consistent R(W)=2 superpotential assignments, and (trial) central charges computed from the encoded R-symmetry must match within the encoded global symmetry content.
- ad hoc to paper Bounded single-trace cyclic words modulo truncated F-terms are an acceptable classical proxy for chiral-ring comparison in this repair setting, even though word length is not duality-invariant and quantum relations are omitted.
- ad hoc to paper For SU(2) nodes, the implemented cubic condition may be a chirality-balance convention; the Z2 Witten anomaly need not be checked for the certificate used in this study.
- ad hoc to paper No a-maximization is required: comparing trial a on the encoded R-charges is enough for the central-charge obligation in this benchmark.
- standard math Binomial GEE with logit link, fixture clustering, exchangeable working correlation, robust sandwich covariance, and Holm control of the primary six-hypothesis family yields valid inference for the preregistered risk differences.
- domain assumption Provider-fault API failures matching a prespecified predicate are exogenous and may be quarantined/re-run without biasing model-quality endpoints; malformed model output always counts.
invented entities (3)
-
DualityCert consistency certificate
independent evidence
-
145-fixture depth-one repair benchmark (six toric seed families)
independent evidence
-
Feedback projections (verifier / masked / generic) and stop-first strategy portfolio
independent evidence
read the original abstract
We present DualityCert, a symbolic verifier for candidate Seiberg-duality claims in four-dimensional N=1 quiver gauge theories. The verifier evaluates 't Hooft anomaly matching, superpotential R-charge consistency, central-charge matching, and a bounded chiral-ring proxy. A claim that passes receives a consistency certificate, which states that no tested inconsistency was found, not that the duality is proven. We use the verifier as a repair environment for language-model agents, which receive a deliberately broken claim and must edit it until it certifies. On a preregistered benchmark of 145 broken claims, with the analysis fixed before the first confirmatory model call, verifier-gated retry improves final repair success over a single attempt by +8.3 percentage points (pp) on deepseek-chat and +7.1 pp on qwen-plus (Holm-adjusted p<0.002). Under an equal budget of eleven attempts, the stop-first strategy portfolio underperforms independent verifier-filtered resampling by 10.3 percentage points on deepseek-chat but outperforms it by 14.7 points on qwen-plus, reversing the ordering of the two tested verifier-exploitation policies across the two confirmatory models. On qwen-plus, category-level verifier feedback is worth +8.7 pp over content-free retry, and interpretable obligation identities alone are worth +6.4 pp over structurally identical masked feedback. Neither effect is detected on deepseek-chat. Separately, a preregistered MiniMax-M2.5 extension again finds an iteration gain and independent verifier-filtered resampling outperforming the strategy portfolio. Which policy is better thus differs between the two models, while every winning policy uses the same cheap certificate. The verifier, benchmark, protocol, and all per-attempt records are released.
Figures
Forward citations
Cited by 1 Pith paper
-
Learning to Trace Seiberg Dualities
Hybrid graph-transformer networks guiding A* and beam search find Seiberg-duality paths between ~10-node quivers more efficiently than BFS or pure physics heuristics, with a measured complexity breaking point.
Reference graph
Works this paper leans on
-
[1]
Grading the Unspoken: Evaluating Tacit Reasoning in Quantum Field Theory and String Theory with LLMs
Xingyang Yu, Yinghuan Zhang, Yufei Zhang, and Zijun Cui. Grading the Unspoken: Evaluating Tacit Reasoning in Quantum Field Theory and String Theory with LLMs. 2026. arXiv:2604.14188
Pith/arXiv arXiv 2026
-
[2]
Daniel J. H. Chung, Zhiqi Gao, Yurii Kvasiuk, Tianyi Li, Moritz M ¨unchmeyer, Maja Rudolph, Frederic Sala, and Sai Chaitanya Tadepalli. Theoretical physics benchmark (TPBench)—a dataset and study of AI reasoning capabilities in theoretical physics.Mach. Learn. Sci. Tech., 6(3):030505, 2025. doi: 10.1088/ 2632-2153/adfcb0. arXiv:2502.15815
Pith/arXiv arXiv 2025
-
[3]
Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research Benchmark
Minhui Zhu et al. Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research Benchmark. 2025. arXiv:2509.26574
Pith/arXiv arXiv 2025
-
[4]
Kaiyu Yang, Aidan M. Swope, Alex Gu, et al. LeanDojo: Theorem Proving with Retrieval-Augmented Language Models. 2023. arXiv:2306.15626
Pith/arXiv arXiv 2023
-
[5]
Trieu H. Trinh, Yuhuai Wu, Quoc V . Le, He He, and Thang Luong. Solving olympiad geometry without human demonstrations.Nature, 625:476–482, 2024. doi: 10.1038/s41586-023-06747-5
-
[6]
Olympiad-level formal mathematical reasoning with reinforcement learning.Nature, 2025
Thomas Hubert, Rishi Mehta, Laurent Sartran, et al. Olympiad-level formal mathematical reasoning with reinforcement learning.Nature, 2025. doi: 10.1038/s41586-025-09833-y
-
[7]
Resolution of Erd ˝os Problem #728: a writeup of Aristotle’s Lean proof
Nat Sothanaphan. Resolution of Erd ˝os Problem #728: a writeup of Aristotle’s Lean proof. 2026. arXiv:2601.07421
arXiv 2026
-
[8]
Learning to Disprove: Formal Counterexample Generation with Large Language Models
Zenan Li, Zhaoyu Li, Kaiyu Yang, Xiaoxing Ma, and Zhendong Su. Learning to Disprove: Formal Counterexample Generation with Large Language Models. 2026. arXiv:2603.19514
arXiv 2026
-
[9]
N. Seiberg. Electric - magnetic duality in supersymmetric nonAbelian gauge theories.Nucl. Phys. B, 435: 129–146, 1995. doi: 10.1016/0550-3213(94)00023-8. arXiv:hep-th/9411149
Pith/arXiv arXiv 1995
-
[10]
Kenneth A. Intriligator and N. Seiberg. Lectures on supersymmetric gauge theories and electric-magnetic duality.Nucl. Phys. B Proc. Suppl., 45BC:1–28, 1996. doi: 10.1016/0920-5632(95)00626-5. arXiv:hep- th/9509066
arXiv 1996
-
[11]
HepLean: Digitalising high energy physics.Comput
Joseph Tooby-Smith. HepLean: Digitalising high energy physics.Comput. Phys. Commun., 308:109457,
-
[12]
Large Language Models Cannot Self-Correct Reasoning Yet
Jie Huang, Xinyun Chen, Swaroop Mishra, et al. Large Language Models Cannot Self-Correct Reasoning Yet. 2023. arXiv:2310.01798
Pith/arXiv arXiv 2023
-
[13]
Training Verifiers to Solve Math Word Prob- lems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, et al. Training Verifiers to Solve Math Word Prob- lems. 2021. arXiv:2110.14168. 15
Pith/arXiv arXiv 2021
-
[14]
Teaching Large Language Models to Self-Debug
Xinyun Chen, Maxwell Lin, Nathanael Sch ¨arli, et al. Teaching Large Language Models to Self-Debug
-
[15]
Self-Refine: Iterative Refinement with Self-Feedback
Aman Madaan, Niket Tandon, Prakhar Gupta, et al. Self-Refine: Iterative Refinement with Self-Feedback
-
[16]
Reflexion: Language Agents with Verbal Rein- forcement Learning
Noah Shinn, Federico Cassano, Edward Berman, et al. Reflexion: Language Agents with Verbal Rein- forcement Learning. 2023. arXiv:2303.11366
Pith/arXiv arXiv 2023
-
[17]
Hunter Lightman, Vineet Kosaraju, Yura Burda, et al. Let’s Verify Step by Step. 2023. arXiv:2305.20050
Pith/arXiv arXiv 2023
-
[18]
Mathematical discoveries from program search with large language models.Nature, 625:468–475, 2024
Bernardino Romera-Paredes, Mohammadamin Barekatain, Alexander Novikov, et al. Mathematical discoveries from program search with large language models.Nature, 625:468–475, 2024. doi: 10.1038/s41586-023-06924-6
-
[19]
Quiver Mutations, Seiberg Duality and Machine Learning.Phys
Jiakang Bao, Sebasti ´an Franco, Yang-Hui He, Edward Hirst, Gregg Musiker, and Yan Xiao. Quiver Mutations, Seiberg Duality and Machine Learning.Phys. Rev. D, 102(8):086013, 2020. doi: 10.1103/ PhysRevD.102.086013. arXiv:2006.10783
Pith/arXiv arXiv 2020
-
[20]
Douglas, Sarah Hoback, Anna Mei, and Ron Nissim
Michael R. Douglas, Sarah Hoback, Anna Mei, and Ron Nissim. Formalization of QFT. 2026. arXiv:2603.15770
arXiv 2026
-
[21]
Michael R. Douglas. Axioms for physical reasoning: codifying the Seiberg–Witten solution in Lean
-
[22]
PhysProver: Advancing Automatic Theorem Proving for Physics
Hanning Zhang et al. PhysProver: Advancing Automatic Theorem Proving for Physics. 2026. arXiv:2601.15737
arXiv 2026
-
[23]
Towards Verifiable and Self-Correcting AI Physicists for Quantum Many- Body Simulations
Ken Deng, Xiangfei Wang, Guijing Duan, Chen Mo, Junkun Huang, Runqing Zhang, Ling Qian, Zhiguo Huang, Jize Han, and Di Luo. Towards Verifiable and Self-Correcting AI Physicists for Quantum Many- Body Simulations. 2026. arXiv:2604.00149
Pith/arXiv arXiv 2026
-
[24]
Edward Witten. An SU(2) Anomaly.Phys. Lett. B, 117:324–328, 1982. doi: 10.1016/0370-2693(82) 90728-6
-
[25]
Naturalness, chiral symmetry, and spontaneous chiral symmetry breaking.NATO Sci
Gerard ’t Hooft. Naturalness, chiral symmetry, and spontaneous chiral symmetry breaking.NATO Sci. Ser. B, 59:135–157, 1980. doi: 10.1007/978-1-4684-7571-5 9
-
[26]
D. Anselmi, D. Z. Freedman, Marcus T. Grisaru, and A. A. Johansen. Nonperturbative formulas for central functions of supersymmetric gauge theories.Nucl. Phys. B, 526:543–571, 1998. doi: 10.1016/ S0550-3213(98)00278-8. arXiv:hep-th/9708042
Pith/arXiv arXiv 1998
-
[27]
Kenneth A. Intriligator and Brian Wecht. The Exact superconformal R symmetry maximizes a.Nucl. Phys. B, 667:183–200, 2003. doi: 10.1016/S0550-3213(03)00459-0. arXiv:hep-th/0304128
Pith/arXiv arXiv 2003
-
[28]
Douglas, Nathan Seiberg, and Edward Witten
Freddy Cachazo, Michael R. Douglas, Nathan Seiberg, and Edward Witten. Chiral rings and anomalies in supersymmetric gauge theory.JHEP, 12:071, 2002. doi: 10.1088/1126-6708/2002/12/071. arXiv:hep- th/0211170
arXiv 2002
-
[29]
Michael R. Douglas and Gregory W. Moore. D-branes, quivers, and ALE instantons. 1996. arXiv:hep- th/9603167
arXiv 1996
-
[30]
D-brane gauge theories from toric singularities and toric du- ality.Nucl
Bo Feng, Amihay Hanany, and Yang-Hui He. D-brane gauge theories from toric singularities and toric du- ality.Nucl. Phys. B, 595:165–200, 2001. doi: 10.1016/S0550-3213(00)00699-4. arXiv:hep-th/0003085
Pith/arXiv arXiv 2001
- [31]
-
[32]
Kennaway, David Vegh, and Brian Wecht
Sebastian Franco, Amihay Hanany, Kristian D. Kennaway, David Vegh, and Brian Wecht. Brane dimers and quiver gauge theories.JHEP, 01:096, 2006. doi: 10.1088/1126-6708/2006/01/096. arXiv:hep- th/0504110
arXiv 2006
-
[33]
Chris E. Beasley and M. Ronen Plesser. Toric duality is Seiberg duality.JHEP, 12:001, 2001. doi: 10.1088/1126-6708/2001/12/001. arXiv:hep-th/0109053
Pith/arXiv arXiv 2001
-
[34]
Kung-Yee Liang and Scott L. Zeger. Longitudinal data analysis using generalized linear models. Biometrika, 73:13–22, 1986. doi: 10.1093/biomet/73.1.13
-
[35]
A simple sequentially rejective multiple test procedure.Scand
Sture Holm. A simple sequentially rejective multiple test procedure.Scand. J. Statist., 6(2):65–70, 1979. 16
1979
- [2025]
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.