{"id":"1020692b-acf4-416c-81e5-cb38295eeed3","arxiv_id":"2511.13163","paper_version":5,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A physics-inspired neural network trained only on experimental Au+Au data interpolates dN/dη, pT spectra, and v2 to RHIC energies without published measurements.","lead":"A neural network designed to mirror the stages of a heavy-ion collision was trained on real RHIC data to predict particle multiplicity, transverse-momentum spectra, and elliptic flow across collision energies. It reproduces held-out measurements and is then used to estimate these observables at RHIC energies that have not yet been measured.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Test-set leakage across observables at the same collision energy invalidates the claimed generalization to unseen energies","rationale":"I agree with the reader that the architecture-ablation claim is fragile and that the external validation is partially circular, but I find a more direct and load-bearing problem: the held-out test protocol itself is contaminated. The paper's headline capability—predicting observables at collision energies not yet explored—requires demonstrating that the network generalizes to energies completely absent from training. Table I shows that the claimed test-set energies are not globally held out: every energy appears in the training set through at least one other observable. Given the shared-input design (energy is a scalar input and all observables share hidden layers), this leaks information about the held-out energy into the model. Consequently, the quantitative support for 'generalization to new energies' in Figs. 4, 9, and 12 is weaker than presented. This is not a matter of disagreement with physics consensus; it is an internal evaluation flaw. The concern is addressable with a leave-one-energy-out retraining, so the paper remains a plausible proof-of-concept, but the evidence for its central claim is materially weaker. The reader's CONDITIONAL verdict already reflects caution, so I do not move the verdict; I would simply add this specific condition to the acceptance criteria. This is why I select UNCHANGED and partial agreement: I share the reader's concern about validation fragility but identify a distinct, more mechanistic mechanism.","tokens_in":11872,"tokens_out":8079,"duration_ms":80833,"concrete_test":"Run leave-one-energy-out cross-validation: for each energy E in {7.7, 9.2, 11.5, 14.5, 19.6, 27, 39, 62.4, 130, 200}, remove all three observables at E from the training set, retrain the model with identical architecture and hyperparameters, and evaluate on all available experimental data at E. Compare these held-out-energy test losses to the losses reported in Table II and the visual accuracy in Figs. 4, 9, and 12. If the test losses increase by more than a factor of 2 for any observable, the original test protocol was inflated by cross-observable leakage and the claimed generalization to unexplored energies is not established.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that a network trained on three observables over a range of RHIC energies can predict those observables at energies not yet explored. The held-out tests are meant to support this generalization, but Table I reveals that every test energy already appears in the training set through a different observable. Specifically: dN/dη is tested at 130 GeV while v2 at 130 GeV is in the training set; v2 is tested at 19.6 and 39 GeV while dN/dη at 19.6 and pT spectra at 39 GeV are in training; pT spectra are tested at 19.6 and 130 GeV while dN/dη at 19.6 and v2 at 130 GeV are in training. Because the network takes the collision energy as an input and shares hidden representations across all three output tasks, training on any observable at energy E gives the network information about E that can be exploited when predicting a different observable at E. Thus the reported test losses do not measure generalization to a truly unseen energy; they measure cross-observable transfer at an energy already represented in training. The actual target energies—17.3, 54.4, 13.7, 9.2—are absent from the entire training set, so the quantitative evaluation never probes the regime the paper claims to address. The external validation (CLVisc, multiplicity scaling) is indirect and partly calibrated to the same observable family, so it cannot independently rescue the generalization claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a physics-inspired neural network trained exclusively on experimental Au+Au data at RHIC to reproduce and predict three bulk observables: charged-particle pseudorapidity density dN/dη, elliptic flow v2(pT), and transverse-momentum spectra dN/(2πpTdpTdη). The architecture uses a stage-inspired layering (QGP, chemical freeze-out, kinetic freeze-out, detector) with locally connected Fibonacci-sphere hidden layers and a dual-input design. After training on multiple energies and centralities, the authors report low held-out test losses, show that the structured architecture outperforms fully connected and Dropout baselines, and make predictions for energies they describe as unexplored (17.3, 54.4, 13.7, and 9.2 GeV). These predictions are checked against CLVisc hydrodynamic calculations and against the global energy dependence of total charged-particle multiplicity per participant pair.","tokens_in":12276,"tokens_out":5109,"duration_ms":53293,"significance":"If the generalization claim were fully established, the paper would provide a useful fast empirical surrogate for bulk heavy-ion observables and an interesting demonstration that physically motivated inductive biases can improve multi-task ML performance in nuclear physics. The architecture-ablation results (Table II) are suggestive, and the decision to train only on experimental data is a strength. However, the quantitative evaluation does not actually test the claimed extrapolation to unseen collision energies, and the external validation is partly calibrated to the same observables used for training. The paper's own concluding caveat that it is 'an empirical interpolator' is appropriate, but the central predictive claim is currently stronger than the evidence supports.","major_comments":[{"comment":"The held-out test set does not test generalization to unseen collision energies. Every test energy appears in the training set through at least one other observable: dN/dη is tested at 130 GeV while v2 at 130 GeV is trained; v2 is tested at 19.6 and 39 GeV while dN/dη at 19.6 and pT spectra at 39 GeV are trained; pT spectra are tested at 19.6 and 130 GeV while dN/dη at 19.6 and v2 at 130 GeV are trained. Since collision energy is an input and all outputs share hidden representations, training at energy E for one task gives the network usable information about E for the other tasks. Thus Table II and Figs. 4, 9, and 12 demonstrate cross-observable transfer at already represented energies, not energy generalization. The true target energies (17.3, 54.4, 13.7, 9.2) never appear as held-out tests. An energy-exclusive split (removing all observables at test energies from training) is needed t","section":"Table I; §II (shared layers); §III"},{"comment":"The external validation is not independent. The CLVisc parameters in Table III are tuned to reproduce the most-central charged-particle dN/dη, which is one of the three trained observables, so agreement between the NN and CLVisc partly reflects consistency with the same training family. Likewise, the multiplicity-per-participant fit in Fig. 6 uses exactly the dN/dη/multiplicity data family on which the network was trained. These checks therefore confirm smoothness and consistency rather than genuine predictive power at unmeasured energies. The authors should either validate on observables absent from the training set (e.g., identified-particle spectra, higher-order harmonics, or independent model constraints) or explicitly label Figs. 5–7 as consistency checks rather than validation.","section":"§III, Figs. 5–7; Appendix A"},{"comment":"The architecture-improvement claim is not quantified robustly. Table II reports a single 'best' training and test loss with no standard deviation across random seeds and no explicit model-selection rule. If the test loss is used to select the epoch or architecture, the comparison is partially fitted to the test set. In addition, the causal interpretation of the local-connectivity improvement is underdetermined: a sparse structured network may outperform fully connected and Dropout baselines because of the regularization induced by sparsity, not because the Fibonacci-sphere geometry encodes the collision dynamics. Please provide repeated-run statistics, a validation-based selection protocol, and an ablation against an equally sparse but non-physical connectivity pattern.","section":"Table II; §III"},{"comment":"The 10% error band used for all predictions is ad hoc and not derived from the model or from data coverage. Adding a fixed 10% band does not quantify extrapolation uncertainty, especially at energies and kinematic regions far from the training distribution. For a paper whose central claim is filling data gaps at RHIC, some form of uncertainty quantification — for example, ensemble/Bayesian methods or propagation of training-density-dependent variance — is required before 'accuracy' of the 13.7, 17.3, 54.4, and 9.2 GeV predictions can be assessed.","section":"§III, Figs. 5, 10, 13"}],"minor_comments":[{"comment":"No code, trained weights, or processed dataset are provided, and the preprocessing, centrality binning, and random seed are not fully specified. Given the ML-focused contribution, releasing the code and data pipeline would be essential for reproducibility.","section":"General; reproducibility"},{"comment":"The captions say 'the bands are the error bars,' but the bands are the fixed 10% uncertainty added to the NN predictions, not experimental error bars. Please rephrase to avoid implying they are measured uncertainties.","section":"Captions for Figs. 5, 10, 13"},{"comment":"The phrase 'collision energies not yet explored experimentally at RHIC' is overstated. For example, 9.2 GeV appears in the v2 training set, and 39 GeV appears in the pT-spectra training set. The predictions are for observables at energies not yet measured for that observable; the text should say so explicitly.","section":"Abstract and Section IV"},{"comment":"The Fibonacci-grid formula is given with a somewhat loose notation for x_n and y_n (the square-root factor should be parenthesized consistently), and the reader is not told how the seven-nearest-neighbour connections are computed once the sphere points are projected. Please specify the adjacency criterion precisely.","section":"Eq. (1); §II"},{"comment":"The concluding claim that the network's success suggests 'an empirical indication of underlying universality' is speculative. The training and test sets are all from one collision system and one observable family; this statement should be removed or heavily qualified.","section":"Section IV"}],"recommendation":"major_revision","confidential_remarks":"The paper appears to be a revised version (v5), yet the central validation issue remains: the held-out tests share collision energies with the training set across observables, so the headline claim of predicting at unexplored energies is not quantitatively supported. I am not recommending rejection, because the empirical interpolator idea and the architecture ablations are salvageable, and an energy-exclusive split or a reframed cross-observable claim would address the core problem. I would encourage the editor to require that the authors either provide energy-exclusive held-out results with uncertainty estimates and repeated-run statistics, or substantially soften the generalization claim."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my read. The genuinely new thing is a single network, trained purely on experimental data, that simultaneously reproduces dN/dη, v2 and pT spectra across RHIC energies and centralities, with a collision-stage-inspired locally connected architecture. The ablations in Table II are real evidence that the sparse structured connectivity and split input help; that is a legitimate, useful design lesson. The paper is also honest about being an empirical interpolator and flags the beam-remnant region.\n\nThe problem is the stress-test one, and it's real. Table I: every test energy already appears in the training set through another observable. 130 GeV appears as v2 training while dN/dη is tested at 130; 19.6 appears as dN/dη and pT training while v2 is tested; 39 appears as pT training while v2 is tested. Since the network gets energy as input and shares hidden features across outputs, the test losses measure cross-observable transfer at an energy the network has already seen, not generalization to a new energy. So the central quantitative claim—that the held-out tests support predictions at unexplored energies—does not follow. The actual predicted energies (17.3, 54.4, 13.7, 9.2) are never tested against data; the CLVisc comparison is partly circular because CLVisc is tuned to the most central dN/dη, one of the trained observables; the multiplicity scaling fit is indirect.\n\nOther soft spots: no code or data, no error bars on the loss numbers, and the 10% band is a convention, not a calibrated uncertainty. These are all addressable. I don't see a fatal flaw in the surrogate idea itself; a network trained on all available data could still be a useful interpolator, and the authors' predictions are plausible. But the evidence as presented is weaker than the abstract implies.\n\nThis is exactly a paper that should go to peer review—serious referees should see it—but not in current form. The authors should redo the held-out evaluation leaving out all observables at the test energies, report seed-to-seed spread, release code and processed data, and either calibrate uncertainties or label the bands as schematic. With that, the surrogate claim would be credible. I'd send it back for major revision.","headline":"A promising physics-guided multi-observable surrogate for RHIC, but the held-out tests leak energy information across observables, so the headline generalization to unseen energies is not actually demonstrated.","tokens_in":12719,"tokens_out":3039,"would_cite":true,"duration_ms":27769,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["25.75.Dw","25.75.-q","24.10.Nz"],"model":"deepseek-v4-flash","headline":"A single neural network, designed to mirror the stages of a heavy-ion collision and trained exclusively on existing experimental data, can reproduce and predict three bulk observables in Au+Au collisions across RHIC energies.","keywords":["deep learning","heavy-ion collisions","Au+Au","RHIC","pseudorapidity density","transverse momentum spectra","elliptic flow","physics-guided neural network"],"falsifier":"Compare the network's predicted dN/dη, v2(pT), and pT spectra at 54.4 GeV Au+Au (or another energy the paper treats as unmeasured) against newly published high-statistics data from the same accelerator: systematic disagreement beyond the claimed ~10% error band would falsify the surrogate claim. A second, more targeted falsifier: if a random sparse graph with the same degree distribution as the Fibonacci-sphere network achieves the same test loss on the test set, then the claim that the physics prior causes the improvement is falsified.","tokens_in":11792,"feed_emoji":"⚛️","tokens_out":13357,"duration_ms":99980,"temperature":0.7,"pith_summary":"This paper tries to establish a practical claim: a single neural network, designed to mirror the stages of a heavy-ion collision and trained only on existing experimental measurements, can reproduce three bulk observables—charged-particle pseudorapidity density, transverse-momentum spectra, and elliptic flow—across a wide range of RHIC collision energies and centralities, and can then predict those observables at energies not yet measured. The network's architecture encodes the collision timeline (quark-gluon plasma, chemical freeze-out, kinetic freeze-out, detector) and uses locally connected layers whose neurons are spread on expanding spheres via a Fibonacci grid, a geometric prior the authors call a uniformly expanding fireball. They show that this structured design, together with a dual-input split (collision parameters in one place, final-particle kinematics in another), markedly outperforms fully connected baselines and Dropout alternatives. If these claims hold, the trained network is an efficient empirical surrogate for filling data gaps at RHIC and supports the idea that bulk observables are governed by a small number of effective macroscopic parameters.","feed_headline":"One network predicts unmeasured RHIC collision observables","feed_subtitle":"Trained only on real data, it reproduces spectra and flow, and forecasts unmeasured beam energies.","key_machinery":"The load-bearing mechanism is the locally connected spherical-layer stack: neurons in the quark-gluon plasma, chemical freeze-out, and kinetic freeze-out layers are distributed on spheres of increasing radius by the Fibonacci grid, and each neuron connects to its seven nearest neighbours on the neighbouring sphere, creating a sparse, physically motivated connectivity that suppresses long-range correlations and regularizes the network. The complementary mechanism is the split dual input: one input carries the collision system's parameters (masses, energy, centrality range), the other carries the final-particle kinematic coordinate (pseudorapidity or transverse momentum, with a categorical mar","core_discovery":"The central discovery is that a single physics-structured network can simultaneously learn the mapping from collision parameters (nucleon masses, collision energy, centrality) and final-particle kinematics (pseudorapidity or transverse momentum, with a categorical marker for the kinematic bin) to three distinct observables, using only real experimental data, and can interpolate to collision energies that were not part of the training set. The architecture mirrors the collision timeline: layers corresponding to the quark-gluon plasma, chemical freeze-out, kinetic freeze-out, and detector, with each hidden layer's neurons placed on a sphere of increasing radius using a Fibonacci grid and conne","pith_inferences":["A decisive control experiment is missing from the paper's ablations: replacing the Fibonacci-sphere seven-nearest-neighbour adjacency with a random sparse graph of the same degree. If the random graph matches the test loss, the specific fireball geometry is not the cause of the improvement; if it degrades, the physics prior is doing genuine work.","Surrogate validation would be stronger if it used an observable outside the training family, such as identified-particle ratios or femtoscopic radii, since the current checks (multiplicity systematics and a hydrodynamic model) are anchored to multiplicities and spectra that the network already sees.","A natural stress test for the interpolation claim is to query the network across a domain boundary, e.g., train on RHIC Au+Au energies and ask for predictions at LHC energies or at very low energies; a graceful extrapolation would support the low-dimensional universality reading, while a sharp failure would delineate the surrogate's valid range.","Beyond prediction, the same architecture could be inverted: replace the observed kinematics in the input with the latent layer activations and use the trained network as a learned summary statistic to constrain QGP properties (shear viscosity, initial conditions) from experimental data, which the paper only gestures at as 'fast emulators'."],"forward_implications":["The trained network can act as a fast, experiment-only surrogate for generating pseudorapidity density, transverse-momentum spectra, and elliptic flow at RHIC energies where data are sparse or absent, without invoking a specific microscopic model.","Because the network was trained without model-generated synthetic data, its success shows that a well-structured network can learn directly from experimental measurements, bypassing the need for hydrodynamic simulations in interpolation tasks.","The physics-motivated architecture (local connections plus split input) is the reason for the performance gain; replacing these with Dropout or removing them degrades results, so the design template matters.","The predictions are consistent with a viscous hydrodynamic calculation and with the global energy dependence of total charged-particle multiplicity per participant pair, supporting the reliability of the surrogate.","The network's ability to reproduce multiple bulk observables suggests an effective universality—bulk particle production appears governed by a few macroscopic parameters even though the network never receives those parameters explicitly."],"fun_headline_variants":["Single network forecasts unseen RHIC energies","Physics-inspired net predicts unmeasured collisions","One model maps collision parameters to observables","DL surrogate predicts RHIC observables at new energies","Neural net trained on real data interpolates to new beam energies"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"That the geometric prior of a uniformly expanding fireball—encoded as seven-nearest-neighbour connections on Fibonacci-grid spheres of increasing radius—matches the actual correlation structure of the data; if it does not, the ablation results only show that a sparse structured network outperforms the baselines, not that physics guidance is responsible.","fun_headline_variants_meta":{"raw":{"variants":["Single network forecasts unseen RHIC energies","Physics-inspired net predicts unmeasured collisions","One model maps collision parameters to observables","DL surrogate predicts RHIC observables at new energies","Neural net trained on real data interpolates to new beam energies"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000181,"raw_usage":{"total_tokens":1147,"prompt_tokens":747,"completion_tokens":400,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":491,"completion_tokens_details":{"reasoning_tokens":339}},"tokens_in":491,"tokens_out":400,"duration_ms":3541,"temperature":1.0,"reasoning_tokens":339,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T21:53:33.499389+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare the network's predicted dN/dη, v2(pT), and pT spectra at 54.4 GeV Au+Au (or another energy the paper treats as unmeasured) against newly published high-statistics data from the same accelerator: systematic disagreement beyond the claimed ~10% error band would falsify the surrogate claim. A second, more targeted falsifier: if a random sparse graph with the same degree distribution as the Fibonacci-sphere network achieves the same test loss on the test set, then the claim that the physics prior causes the improvement is falsified.","supporting_citations":[],"review_version":1}