{"id":"d2550d54-99b3-4064-bcd6-cadc07b32bd0","arxiv_id":"2501.12439","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A new suite of cosmological simulations combines the IllustrisTNG galaxy formation model with warm and self-interacting dark matter, showing that galaxies look similar across models while halo structure and power spectra change.","lead":"Astronomers ran the same galaxy formation model in six different dark matter universes, including warm and self-interacting dark matter, to see which cosmic structures and galaxies survive. The AIDA-TNG simulation suite gives researchers a shared platform to test whether small-scale tensions with cold dark matter point to new physics.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claim that TNG produces realistic galaxies in all DM models rests on the untested transferability of a CDM-calibrated subgrid model; similarity across AIDA runs alone cannot distinguish physics from rigidity.","rationale":"The reader's weakest assumption is exactly the transferability of the unchanged TNG subgrid model to WDM and SIDM halo potentials. My stress-test pass finds that this is the most load-bearing premise behind the abstract's central claim. The paper is honest about having kept the model unchanged and does test several galaxy properties, but the tests are all internal comparisons within a single galaxy-formation code. Such comparisons can establish that the TNG model does not produce large galaxy differences across these dark-matter models, but they cannot by themselves establish that the resulting galaxy population is realistic in an absolute sense or that the similarity would survive a different, equally plausible subgrid implementation. The comparison in Section 4.1 to EAGLE-based SIDM simulations already hints that baryonic responses can depend on the galaxy-formation model, which strengthens the concern. I therefore agree with the CONDITIONAL verdict: the concern is real but not fatal, and the concrete test above would either retire or confirm it. No verdict change is needed from the reader's conditional assessment.","tokens_in":29472,"tokens_out":5719,"duration_ms":70270,"concrete_test":"Compare AIDA-TNG WDM3 and SIDM1 (and WDM1 where available) against published EAGLE-based ADM simulations over the same dark-matter models and overlapping halo-mass/resolution ranges: Oman et al. 2024 for WDM, and Robertson et al. 2018/2019 or TANGO-SIDM for SIDM. Specifically, overlay the stellar mass function, Mstar-Mhalo relation, gas fraction, and galaxy size distributions from AIDA and EAGLE for CDM versus WDM/SIDM. If both independent subgrid models show the same small DM-model-induced differences, the AIDA conclusion is robust; if EAGLE shows, for example, a clear SIDM or WDM shift in stellar mass or gas fraction where TNG does not, the AIDA similarity is a TNG-calibration artifact rather than a physical insensitivity.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract and Section 5 claim that the TNG galaxy formation model 'can produce a realistic galaxy population in all scenarios,' despite being calibrated on CDM. The load-bearing premise is stated in Section 2.3: the fiducial TNG subgrid model is kept entirely unchanged in all runs. For the central claim to hold, the subgrid prescriptions for star formation, feedback, and black hole growth must respond correctly to altered dark-matter potentials, not merely reproduce similar results because they are insensitive to those potentials. The evidence in Figures 9 and 13 shows near-identical stellar mass functions, stellar and gas mass fractions, SMBH masses, and SFRDs across CDM, WDM, and SIDM. But all of these curves are generated by the same TNG code. Agreement between runs of one code cannot separate a physically robust insensitivity from a subgrid model whose calibrated feedback is too rigid to notice changes in halo concentration, core structure, or potential depth. The paper's own comparison to EAGLE-based SIDM simulations in Section 4.1 reveals that baryonic response at 10^12-10^13 Msun differs between galaxy formation models, so the conclusion is not obviously model-independent. Without an independent galaxy-formation model or a direct observational test in the mass range where ADM alters halo structure, the 'realistic in all scenarios' claim remains conditional.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces the AIDA-TNG project, a suite of cosmological magnetohydrodynamic simulations run with the Arepo code and the IllustrisTNG galaxy formation model, in cold dark matter (CDM), three warm dark matter (WDM) models with particle masses 1, 3, and 5 keV, and two self-interacting dark matter (SIDM) models with constant (1 cm^2/g) and velocity-dependent (Correa 2021) cross-sections. Each model is run as dark-matter-only and full-physics versions of two cosmological boxes, 110.7 and 51.7 Mpc, at two resolution levels, using the same initial conditions as the corresponding TNG100 and TNG50 runs. The paper presents first results on the halo mass function and its redshift evolution, the stellar mass function, halo density profiles, the concentration-mass relation, galaxy scaling relations (stellar mass-halo mass, gas fraction, SMBH mass, star formation rate density, galaxy sizes), and the matter power spectrum. The central claim is that, despite the TNG subgrid model being calibrated on CDM, galaxy properties such as stellar and gas mass fractions, stellar mass function, SMBH masses, and SFRD are very similar across all dark matter scenarios, while differences appear in halo structure (cores, concentrations) and galaxy sizes (SIDM galaxies about 20 percent larger).","tokens_in":29734,"tokens_out":7808,"duration_ms":81489,"significance":"If the results hold, AIDA-TNG will be a valuable community resource: it combines cosmological volumes, a well-tested baryonic model, and multiple dark matter alternatives in matched initial conditions, with DMO/FP pairs that allow baryonic and dark-matter effects to be separated. Strengths of the paper include the transparent description of initial conditions, resolution limits, and artificial-fragmentation masking for WDM; the use of existing TNG initial conditions for controlled comparisons; and the public availability of the data. The conclusion that global galaxy properties are insensitive to the dark matter model is interesting and, if correct, has practical importance for interpreting observations. However, the paper's strongest claim, that the TNG model 'can produce a realistic galaxy population in all scenarios,' rests on the untested transferability of a CDM-calibrated subgrid model to altered dark-matter potentials, and this assumption is acknowledged but not independently validated.","major_comments":[{"comment":"The central claim that the TNG galaxy formation model 'can produce a realistic galaxy population in all scenarios' is stronger than the presented evidence. The evidence in Figs. 9 and 13 shows that one subgrid model (TNG) yields similar galaxy properties across the AIDA runs, but this similarity cannot by itself distinguish a physically robust insensitivity from a subgrid model that is too rigid to respond to changes in halo potential. The paper states in Sec. 2.3 that the TNG model is kept entirely unchanged; this is a reasonable design choice for a first study, but the 'realistic' conclusion requires either a quantitative observational test in the mass range where ADM actually changes halo structure (e.g., M_vir < 1e11 Msun for WDM3/WDM1 or the dwarf regime for SIDM), or an explicit statement that the conclusion is conditional on the TNG subgrid model remaining valid. The paper's own comparison in Sec. 4.1 to EAGLE-based SIDM simulations shows that baryonic response differs between galaxy formation models at 1e12-1e13 Msun, so the result is not known to be galaxy-formation-model independent. Please either soften the claim to 'consistent with the TNG subgrid model remaining approximately valid' or add the missing quantitative test.","section":"Abstract; Section 5; Section 2.3"},{"comment":"The abstract claims that the TNG model produces a realistic galaxy population in all scenarios, but the full-physics runs do not include WDM5. Table 1 shows no FP WDM5 run in any of the presented boxes, and Sec. 2.1 states that a FP version of the 50/A WDM5 box was deliberately not created. Thus the galaxy-population similarity is not simulated for the 5 keV WDM model; it is inferred from the DMO run being close to CDM. This is a load-bearing gap for the 'all scenarios' phrasing. Please either add a full-physics WDM5 run (even at lower resolution) or restrict the conclusion to the scenarios for which full-physics runs exist.","section":"Abstract; Table 1; Section 2.1"},{"comment":"The WDM1 FP/DMO halo mass function ratio at z >= 2 shows a low-mass excess that the authors themselves attribute to a possible effect of artificial fragmentation ('this could be a non-trivial consequence of artificial fragmentation'). Since the paper only masks haloes below M_lim rather than removing spurious haloes, the low-mass behaviour in Fig. 8 for WDM1 is not quantitatively robust. This matters because the paper uses Fig. 8 to argue that baryonic effects on the halo mass function are similar across dark matter models. Applying the Lovell et al. (2014) sphericity-based spurious-halo removal, or at least showing the ratio with and without haloes below M_lim, would strengthen this specific conclusion.","section":"Section 3; Figure 8"}],"minor_comments":[{"comment":"The abstract states that the simulations resolve haloes down to 10^8 Msun, but the mass functions in Fig. 6 use a 100-particle limit and the lowest 50/A dark matter particle mass gives 100*m_DM ~ 4e8 Msun; please reconcile the quoted mass range with the actual resolution limit.","section":"Abstract; Section 3"},{"comment":"The text uses 'WIMPS'; the correct acronym is 'WIMPs'.","section":"Section 1"},{"comment":"The phrase 'a economic use' should be 'an economic use'.","section":"Section 2"},{"comment":"The paper quotes half-mode masses for the WDM models but does not provide the computed M_lim values for each run; a small table or appendix listing M_lim per box and resolution would help readers interpret the dashed-line regions in Fig. 6.","section":"Section 2.1"},{"comment":"In the bottom-left panel, the caption says 'The dotted lines mark the 1σ region' but it is unclear whether this is the scatter of the simulations or an observational reference; please specify.","section":"Figure 13"},{"comment":"The caption notes that observed sizes are projected half-light radii while simulated sizes are 3D half-mass radii, but the text should reiterate this caveat because it directly affects the interpretation of the ~20% size difference.","section":"Figure 14"},{"comment":"The description of the vSIDM model as having a cross-section 'inversely proportional to the relative velocity' is a simplification; the Correa (2021) model has a more specific velocity dependence, so please rephrase to avoid implying a pure 1/v scaling.","section":"Section 6.2"}],"recommendation":"major_revision","confidential_remarks":"This is a solid and useful resource paper, and the simulation suite is well designed. My main concern is the gap between the abstract's 'realistic galaxy population in all scenarios' and the actual evidence, which is limited to one subgrid model and excludes WDM5 from the full-physics runs. The transferability assumption in Sec. 2.3 is acknowledged, but the paper should either test it more directly or phrase the conclusion conditionally. The missing WDM5 full-physics run is a concrete gap that should be addressed in the text or with an additional run. I would support publication after a major revision that tightens the central claim and quantifies the artificial-fragmentation caveat in Fig. 8."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing to know: this is a resource paper that will get cited. The new piece is the combination—the TNG galaxy formation model run unchanged in three WDM and two SIDM variants plus CDM, in matched 50 and 100 Mpc boxes, each in dark-matter-only and full-physics versions. That is a well-designed matrix, and the paper executes it cleanly: initial conditions, resolution, softening lengths, the M_lim masking, and comparison to previous EAGLE/BAHAMAS/TANGO runs are all handled carefully. The first measurements (mass functions, density profiles, concentration-mass relation, power spectra, galaxy sizes) are sensible, and the opposite trends between WDM (cores at low mass) and SIDM (cores at high mass, cuspy at 1e12–1e13 Msun) come through clearly. Credit where due: the paper is honest about what it does not do—artificial fragmentation is masked rather than eliminated, the NFW fits are acknowledged as a poor description for cored profiles, and the WDM1 model is flagged as already excluded. That is the right tone for a survey paper.\n\nThe soft spot is the abstract's claim that the TNG model \"can produce a realistic galaxy population in all scenarios.\" What the paper actually shows is that the model's output is nearly unchanged across dark matter models. That is a statement about the TNG subgrid, not about nature. The transferability concern is real: a CDM-calibrated feedback model that is insensitive to potential depth would produce exactly this kind of similarity. The paper's own comparison to EAGLE-SIDM at 1e12–1e13 Msun reveals that the baryonic response differs between TNG and EAGLE, so the similarity is not model-independent. The paper does acknowledge this in Sec 4.1 (\"it could be heavily affected by the galaxy formation model\"), which softens the blow, but the abstract overstates it. I would like to see either a direct observational test in the mass range where ADM changes halo structure, or a second galaxy formation model, before calling the population \"realistic.\"\n\nTwo smaller things. The data is only available on request, not public; for a benchmark suite, that shortens its useful life. And the 5–10% SIDM excess in the full-physics halo mass function at low masses is left unexplained—minor in the context of this paper, but it deserves a comment.\n\nNet: this deserves a serious referee. It is not a groundbreaking claim paper; it is a careful, well-documented resource with first results. I would accept it after the \"realistic\" phrasing is softened and the data availability statement is clarified. I would cite it.","headline":"A well-executed simulation resource that will become a benchmark, but the 'realistic in all scenarios' claim outruns the evidence—similarity across one code's runs is not yet a physical result.","tokens_in":30322,"tokens_out":2830,"would_cite":true,"duration_ms":28795,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper introduces the AIDA-TNG simulation suite and argues that a galaxy formation model calibrated on cold dark matter still yields a realistic galaxy population when the dark matter is warm or self-interacting.","keywords":["AIDA-TNG","galaxy formation","warm dark matter","self-interacting dark matter","cosmological hydrodynamical simulations","halo mass function","matter power spectrum","galaxy sizes"],"falsifier":"Take the WDM3 and SIDM1 boxes and re-run them with the TNG feedback parameters varied by a factor of two (for example, doubling or halving the galactic wind energy). If the resulting changes in stellar mass fractions or galaxy sizes are as large as the dark-matter-driven differences, the unchanged-model comparison cannot isolate dark matter physics, and the central claim fails; on the observational side, a survey search for the predicted roughly 20 percent larger SIDM galaxies at fixed stellar mass would provide a direct test.","tokens_in":29245,"feed_emoji":"🌌","tokens_out":9698,"duration_ms":85737,"temperature":0.7,"pith_summary":"The AIDA-TNG project runs cosmological simulations with the same baryonic galaxy formation recipe as the IllustrisTNG model, but with six dark matter variants: cold dark matter, three warm dark matter models (light particles that erase small-scale structure), and two self-interacting models (dark matter that scatters with itself and smooths dense centres). The paper's central claim is that the unmodified TNG model, calibrated on cold dark matter, produces a realistic galaxy population in every scenario. Stellar and gas mass fractions, the stellar mass function, supermassive black hole masses, and the star formation rate density are nearly identical across models, with deviations only in the most extreme warm model. The clear differences are structural: warm dark matter creates cores in low-mass haloes, self-interacting dark matter creates cores in high-mass haloes, and self-interacting models make galaxies about 20 percent larger. The suite's matched volumes and paired dark-matter-only and full-physics runs let baryonic and dark matter effects be separated, so the project can say where alternative dark matter should appear in observations.","feed_headline":"Galaxy formation recipe survives warm and self-interacting dark matter","feed_subtitle":"Stellar masses, black holes, and star formation are unchanged; halo cores and galaxy sizes reveal the dark matter.","key_machinery":"The central object is the AIDA-TNG simulation suite: 51.7 and 110.7 Mpc cosmological boxes, each run with the AREPO magnetohydrodynamic code, the IllustrisTNG galaxy formation model left entirely unchanged, and six dark matter models—CDM, three WDM masses (1, 3, 5 keV), and two SIDM cross-sections (constant $\\sigma/m=1\\,\\mathrm{cm}^2\\,\\mathrm{g}^{-1}$ and a velocity-dependent model). Two features carry the argument: matched initial conditions taken from TNG50 and TNG100, so any difference between models is caused by dark matter physics, and paired dark-matter-only and full-physics runs, which let the authors factor baryonic feedback out of the comparison. The WDM models are imposed through a transfer-function suppression of the initial power spectrum with a half-mode mass for each particle mass, while the SIDM models use the Monte Carlo scattering scheme in AREPO.","core_discovery":"The core claim is that the redshift-zero galaxy population produced by the IllustrisTNG galaxy formation model is statistically indistinguishable in its global scaling relations across CDM, WDM (1, 3, 5 keV), and SIDM (constant and velocity-dependent cross-sections), even though the model was tuned on CDM alone. The paper supports this by comparing matched cosmological boxes in dark-matter-only and full-physics versions of each model, and by measuring the halo mass function, stellar mass function, stellar and gas mass fractions, SMBH–stellar mass relation, and star formation rate density. It reports that all these quantities agree closely across scenarios, with the largest deviations in the 1 keV warm model, which is already excluded by other observations. The paper also reports that the differences that do survive are structural: WDM suppresses low-mass halo counts and lowers central densities at the low-mass end, while SIDM erodes central cusps at the high-mass end, and SIDM galaxies have stellar half-mass radii about 20 percent larger than CDM at fixed stellar mass. On scales below about 1 Mpc, the matter power spectrum is suppressed in all models, but WDM and SIDM reach that suppression from opposite directions—WDM from the initial power-spectrum cut-off, SIDM from late-time core formation.","pith_inferences":["An implication of the paper's set-up is that observational tensions between CDM and galaxy scaling relations are unlikely to be resolved by switching to these WDM or SIDM models, since the same baryon recipe reproduces the same galaxy population in all of them; tensions would have to come from structural or small-scale data.","The near-universality of baryonic effects suggests a practical shortcut the authors do not spell out: a single baryonic correction fitted in CDM could be applied to dark-matter-only predictions in any of these dark matter models, as long as structural differences are treated separately.","If the SIDM galaxy-size signal survives comparison with surveys, it provides a way to break degeneracies with baryonic feedback, which also inflates galaxy sizes; the two effects could be separated by combining size data with central dark matter densities.","The velocity-dependent SIDM model's smaller cores at cluster masses imply that cluster-scale constraints on constant cross-sections may not transfer directly to velocity-dependent models, pointing to dwarf-scale structure as the discriminating regime."],"forward_implications":["Observers can use TNG-based mock galaxy populations to test warm and self-interacting dark matter without first re-calibrating the baryon model for each scenario.","Baryonic effects on halo counts and the small-scale matter power spectrum can be treated as approximately universal across these dark matter models, so dark-matter-only predictions can be corrected by a single baryonic transfer function.","Self-interacting dark matter predicts a measurable population of galaxies roughly 20 percent larger at fixed stellar mass in the range $5\\times10^9$ to $10^{12}\\,M_\\odot$, a signature that large imaging surveys can look for directly.","Warm and self-interacting dark matter should be sought in halo structure and small-scale clustering, not in global galaxy scaling relations: SIDM cores at high halo masses, WDM cores and missing haloes at low masses."],"supporting_citations":[{"why":"Defines the TNG kinetic AGN feedback model that the paper keeps unchanged in all dark matter scenarios.","marker":"Weinberger et al. 2017"},{"why":"Supplies the TNG galaxy formation model and its calibration quantities, the baseline the paper tests in alternative dark matter.","marker":"Pillepich et al. 2018b"},{"why":"Defines the TNG50/TNG100 initial conditions, halo catalogues, and resolution levels that AIDA-TNG reuses for matched comparisons.","marker":"Nelson et al. 2019"},{"why":"Provides the TNG100 run from which the 110.7 Mpc initial conditions and resolution levels are taken.","marker":"Pillepich et al. 2018a"},{"why":"Provides the TNG300 matter power spectrum and baryonic suppression measurement that AIDA-TNG validates against.","marker":"Springel et al. 2018"},{"why":"Provides the WDM mass-function suppression formula and the spurious-halo identification method used to interpret WDM runs.","marker":"Lovell et al. 2014"},{"why":"Provides the CDM virial mass function prediction used as the DMO baseline in the halo mass function comparison.","marker":"Despali et al. 2016"},{"why":"Supplies the velocity-dependent SIDM cross-section model adopted as the vSIDM scenario.","marker":"Correa 2021"},{"why":"Implements the SIDM scattering scheme in AREPO that the simulations use for self-interactions.","marker":"Vogelsberger et al. 2012"}],"fun_headline_variants":["Galaxy formation survives WDM and SIDM, simulations show","Global galaxy properties unchanged by alternative dark matter","AIDA-TNG: baryonic outcomes robust to dark matter variations","Dark matter alternatives alter halo cores, not galaxy scaling","Stellar masses and black holes resist warm and self-interacting dark matter"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise, stated in Section 2.3, is that the TNG galaxy formation model, tuned on cold dark matter, also governs gas cooling, star formation, and black hole feedback correctly inside haloes whose dark matter is warm or self-interacting—so the similar galaxy properties are physical insensitivity, not an artifact of a rigid subgrid model.","fun_headline_variants_meta":{"raw":{"variants":["Galaxy formation survives WDM and SIDM, simulations show","Global galaxy properties unchanged by alternative dark matter","AIDA-TNG: baryonic outcomes robust to dark matter variations","Dark matter alternatives alter halo cores, not galaxy scaling","Stellar masses and black holes resist warm and self-interacting dark matter"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000656,"raw_usage":{"total_tokens":3109,"prompt_tokens":1156,"completion_tokens":1953,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":772,"completion_tokens_details":{"reasoning_tokens":1869}},"tokens_in":772,"tokens_out":1953,"duration_ms":15474,"temperature":1.0,"reasoning_tokens":1869,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T17:12:09.394024+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the WDM3 and SIDM1 boxes and re-run them with the TNG feedback parameters varied by a factor of two (for example, doubling or halving the galactic wind energy). If the resulting changes in stellar mass fractions or galaxy sizes are as large as the dark-matter-driven differences, the unchanged-model comparison cannot isolate dark matter physics, and the central claim fails; on the observational side, a survey search for the predicted roughly 20 percent larger SIDM galaxies at fixed stellar mass would provide a direct test.","supporting_citations":[{"cited_title":"R., Frenk, C","cited_arxiv_id":null,"evidence_quote":"Provides the WDM mass-function suppression formula and the spurious-halo identification method used to interpret WDM runs."}],"review_version":1}