{"id":"59447057-55dc-424f-97e0-90dd8cab5883","arxiv_id":"2502.06077","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"In IllustrisTNG, the mutual information between galaxy formation efficiency and halo assembly time exceeds that of colour, sSFR, or cluster observables, especially for low-mass central galaxies.","lead":"This paper uses a statistical tool called Mutual Information to compare how tightly galaxy properties track the assembly history of their dark matter haloes in the IllustrisTNG simulation. It finds that the galaxy formation efficiency, the ratio of stellar to halo mass, carries the strongest signal, which could help observers estimate halo ages from galaxy data.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"F⋆–zH mutual information is largely inherited from the Mh–zH relation; without a conditional-MI test the 'sensitive indicator' claim is unsupported.","rationale":"The reader's weakest_assumption is exactly the load-bearing issue. I agree with the CONDITIONAL verdict rather than moving it, because the paper is transparent about the confound and the reported MI values are plausibly computed from public TNG300-1 data with a standard GMM-MI estimator; those parts are not in question. The problem is interpretive: the abstract's stronger claim ('establishes F⋆ as a more sensitive indicator...') goes beyond what the estimator can show. Concretely, F⋆ = log10(M⋆/Mh). In each 0.25-dex M⋆ bin, M⋆ varies little, so F⋆ is nearly a monotone decreasing function of Mh. The inset to Fig. 4 explicitly states that the F⋆-zH correlation is 'primarily driven by' the Mh-zH relation. That admission converts the impressive 0.38–0.43 nats from a discovery about baryonic assembly into a restatement of the standard CDM relation between halo mass and formation time. To rescue the interpretation, the authors would need to show that F⋆ carries MI about zH beyond Mh — for example via conditional MI or by comparing I(F⋆,zH) with I(Mh,zH). Without such a test, the comparison to colour/sSFR/cluster observables is not apples-to-apples, since those variables do not contain halo mass by construction. I would keep the verdict CONDITIONAL, with the condition being exactly this demonstration, and note that the absence of a null-MI baseline makes even the 'weak correlations' hard to calibrate.","tokens_in":20730,"tokens_out":4019,"duration_ms":35939,"concrete_test":"Repeat the Fig. 4 (right-panel) analysis on TNG300-1 centrals in the same M⋆ bins, but compute the conditional mutual information I(F⋆; zH | log10 Mh) with the same GMM-MI bootstrap pipeline, and also compute I(log10 Mh; zH) without conditioning. If I(F⋆; zH | log10 Mh) ≈ 0 while I(log10 Mh; zH) ≈ I(F⋆; zH) in the M⋆ < 10^10.25 bins, then the F⋆ signal is inherited from Mh and the 'sensitive indicator' claim is not supported. Optionally, permute M⋆ within each bin while holding Mh fixed; if the MI is unchanged, M⋆ contributes nothing beyond Mh.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that F⋆ is 'a more sensitive indicator of halo assembly history than colour (g−i), sSFR, or cluster observables' (Abstract). But F⋆ is defined as log10(M⋆/Mh) (Section 2), and the analysis is performed in narrow M⋆ bins. Within a bin, F⋆ is essentially −log10 Mh plus small scatter from M⋆, so any measured I(F⋆; zH) is dominated by the well-known anti-correlation between halo mass and halo assembly time. The paper's own inset to Fig. 4 concedes this: 'The observed correlation between F⋆ and zH for lower-mass galaxies appears to be primarily driven by the correlation between the halo mass and the halo assembly time in this mass range.' Consequently the headline comparison to colour and sSFR is asymmetric: colour and sSFR are baryonic properties that do not contain Mh by construction, while F⋆ is a rescaled halo mass. The reported MI values of 0.38–0.43 nats therefore do not establish that F⋆ adds independent information about zH; they mostly confirm that Mh and zH are correlated. Since the interpretation depends on F⋆'s baryonic component carrying information beyond Mh, the central claim rests on a confound that the reported statistic cannot resolve.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper uses the GMM-MI estimator of mutual information (MI) to quantify, in the IllustrisTNG TNG300-1 simulation, how strongly several galaxy properties (assembly time zG, colour g-i, specific star formation rate, and galaxy formation efficiency F_star = log10(M_star/M_h)) correlate with host dark matter halo assembly time zH, for central and satellite galaxies in stellar mass bins from 10^9 to 10^11.5 h^-1 M_sun. It also measures MI between zH and cluster observables (magnitude gap, satellite distances, richness). The central claim is that F_star is a more sensitive indicator of halo assembly history than colour, sSFR, or cluster observables, based on MI values of 0.38-0.43 nats for low-mass central galaxies versus values below about 0.18 nats for colour and sSFR.","tokens_in":20940,"tokens_out":2880,"duration_ms":29217,"significance":"If the claims were fully supported, the paper would provide a useful observational proxy for halo assembly history, relevant to assembly-bias studies and the galaxy-halo connection. The authors adopt a careful MI estimation approach (GMM-MI with bootstrap uncertainties), use a public simulation with open data, and report error bars candidly, including large uncertainties at high stellar masses. However, the headline 'more sensitive indicator' conclusion is currently undercut by a definitional confound: F_star contains halo mass by construction, and within the narrow stellar mass bins used, F_star is essentially -log10 M_h plus a constant. The paper itself concedes in Section 4 (inset to Fig. 4) that the F_star-zH correlation is primarily driven by the halo mass-zH correlation. Because this concession is about the central claim, the paper requires additional analysis (e.g., conditional MI at fixed M_h, or an explicit M_h-based baseline) before the main conclusion can be accepted.","major_comments":[{"comment":"The central claim that F_star is 'a more sensitive indicator of halo assembly history than colour (g-i), sSFR, or cluster observables' is not supported by the reported statistic. Since F_star is defined as log10(M_star/M_h) (Section 2) and the analysis is performed in narrow stellar mass bins, within a bin F_star is approximately -log10 M_h plus a small scatter from M_star. The measured I(F_star; zH) therefore largely restates the known anti-correlation between halo mass and halo assembly time in this mass range. The authors themselves state in the text accompanying the inset of Figure 4: 'The observed correlation between F_star and zH for lower-mass galaxies appears to be primarily driven by the correlation between the halo mass and the halo assembly time in this mass range.' To establish F_star as an informative, non-redundant indicator, the authors must either report the conditional MI I(F_star; zH | M_h) or a partial-correlation test against M_h, or explicitly compare I(F_star; zH) with I(M_h; zH) and show that F_star adds information beyond halo mass. Without such a test, the 'sensitive indicator' claim is confounded.","section":"Section 4, Figure 4 inset; also Section 2"},{"comment":"The comparison between F_star and the baryonic properties (colour and sSFR) is asymmetric. Colour and sSFR do not contain M_h by construction, whereas F_star is a rescaled halo mass. Even if the reported MI values are correct, a larger MI for F_star than for colour or sSFR does not demonstrate that F_star is a better observational proxy for zH; it only demonstrates that F_star inherits the halo mass-zH correlation. The paper should compute an equally constructed quantity, such as the MI between a pure baryonic efficiency measure and zH at fixed M_h, or should normalize all MI values by the entropy of the respective galaxy property, to make the 'more sensitive' ranking meaningful.","section":"Section 4, Figures 3 and 4"},{"comment":"The MI values for different galaxy properties are compared directly in nats, but MI is not invariant to the marginal distributions of the variables. A property with a larger dynamic range or variance can produce a larger MI even if the underlying dependence is equally strong. The paper does not report normalized MI (e.g., MI divided by the entropy of the property) nor a null/shuffled baseline to establish the scale of 'significant' MI for each panel. Without this, the quantitative ranking of F_star against colour and sSFR is not yet established, particularly given the large error bars acknowledged at higher stellar masses.","section":"Section 4 and Appendix B"}],"minor_comments":[{"comment":"The phrase 'strong correlation' is used repeatedly for MI values around 0.35-0.43 nats, but the paper does not calibrate what a 'strong' MI is in this context; a null-baseline comparison would help the reader interpret these numbers.","section":"Abstract and Section 5"},{"comment":"In the text describing Figure 5, the inequality '0.8 <= r_x,y <= 0.6' is internally inconsistent; it should presumably read 0.6 <= r_x,y <= 0.8, or the intended range should be stated precisely.","section":"Section 4, Figure 5 description"},{"comment":"The horizontal axis labels in Figures 3 and 4 read 'ZF or' instead of 'z_F or' or a clear label; these are likely typos and should be corrected.","section":"Figures 3 and 4"},{"comment":"Several references are incomplete: Montero-Dorta et al. (2020) has no volume or page, and Niemiec et al. (2022) is also missing publication details. The reference list should be checked for consistency with the journal style.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The central confound is acknowledged by the authors in the text itself, so the path to a revised paper is clear: repeat the MI analysis with a conditional or partial measure that removes the M_h-zH contribution. If such an analysis is not possible or does not support the claim, the paper's main conclusion should be reframed as a measurement of the M_h-zH correlation rather than a new galaxy-property indicator. The manuscript's scope fits the journal, but the current version does not yet substantiate its headline result."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. The comparative result is genuinely new: no one else has ranked F*, colour, sSFR, and cluster observables against halo assembly time with MI in TNG300, and the mass-binned MI curves are cleanly computed with GMM-MI and bootstrap uncertainties. But the headline claim—that F* is a 'more sensitive indicator' of assembly history than colour or sSFR—does not hold up as stated, and the authors know it.\n\nThe structural problem is that F* = log10(M*/Mh). Within a narrow M* bin, F* is essentially -log10 Mh plus scatter. So MI(F*, zH) is dominated by the well-known anti-correlation between halo mass and assembly time. The comparison is asymmetric: colour and sSFR are baryonic properties that contain no halo mass by construction, while F* is a rescaled halo mass. The paper's own Fig. 4 inset concedes this: 'the correlation between F* and zH for lower-mass galaxies appears to be primarily driven by the correlation between the halo mass and the halo assembly time.' That is the whole ballgame. The abstract ignores it. A conditional MI, or MI(F* | Mh), is the obvious fix, and it is absent.\n\nWhat the paper does well: the MI estimation is careful, error bars are reported and large where they should be, and the negative results for colour and sSFR are consistent with previous work (Lin et al. 2016). The satellite versus central comparison and the cluster observable section are useful descriptive material. No code is shipped, but TNG300 is public and the methods are standard enough to reproduce.\n\nSoft spots, in proportion: the central claim is overstated, not fabricated. The MI numbers are probably right; the interpretation is what needs work. There is no null baseline, so calling 0.14-0.18 nats 'weak' is qualitative. And the F*-richness discussion in Fig. 7 drifts into narrative beyond what the MI values support. These are fixable in revision.\n\nWho this is for: people working on assembly bias proxies and age matching. It is a solid descriptive study that confirms colour and sSFR are weak proxies and adds F* to the menu, but it does not establish F* as an independent indicator. With a conditional-MI analysis and a softened claim, it would be a useful contribution. I would send it to peer review, but I'd ask for that revision.","headline":"The F*–zH MI is largely Mh–zH in disguise, and the paper's own inset admits it; the abstract's 'sensitive indicator' claim is overstated, but the descriptive MI ranking is a legitimate contribution.","tokens_in":21551,"tokens_out":1883,"would_cite":false,"duration_ms":16778,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Using mutual information on the IllustrisTNG simulation, this paper argues that galaxy formation efficiency F⋆ = log10(M⋆/Mh) is a more sensitive tracer of dark matter halo assembly time than colour, sSFR, or cluster observables, for…","keywords":["galaxy formation","dark matter halo assembly","mutual information","IllustrisTNG","stellar-to-halo mass relation","halo assembly bias","galaxy assembly time","cluster richness"],"falsifier":"Compute MI($F_\\star, z_H$) within narrow bins of halo mass, or the conditional MI($F_\\star, z_H\\mid M_h$), for central galaxies below $10^{10.25}\\,h^{-1}M_\\odot$; if it drops to the level of colour or sSFR, the claim that $F_\\star$ is a more sensitive independent indicator fails. The equivalent observational check is to build group catalogues with weak-lensing halo masses and see whether $F_\\star$ beats colour at fixed $M_h$.","tokens_in":20477,"feed_emoji":"🌌","tokens_out":11060,"duration_ms":88857,"temperature":0.7,"pith_summary":"The paper asks which galaxy property carries the most information about when a dark matter halo assembled half of its present-day mass. Using mutual information instead of linear correlation on the IllustrisTNG TNG300-1 simulation, it finds that for central galaxies below about $10^{10.25}\\,h^{-1}M_\\odot$ the galaxy formation efficiency $F_\\star=\\log_{10}(M_\\star/M_h)$ shares $0.38$–$0.43$ nats with halo assembly time $z_H$, with a Pearson correlation between $0.6$ and $0.8$. In that same mass range, colour $(g-i)$ and specific star formation rate stay below roughly $0.18$ nats, which is why the paper names $F_\\star$ the more sensitive indicator of assembly history. Satellite galaxies show negligible mutual information with their host haloes, and cluster observables such as magnitude gap and satellite distance are weak, with richness the only quantity that strengthens with stellar mass. The paper itself notes that the low-mass $F_\\star$–$z_H$ signal may be largely inherited from the correlation between halo mass and halo assembly time, which is the main caveat against reading $F_\\star$ as fully independent information.","feed_headline":"Galaxy efficiency outranks galaxy colour as halo-assembly tracer","feed_subtitle":"Mutual information on IllustrisTNG shows F⋆ shares up to 0.43 nats with halo formation time; colour and sSFR lag below 0.18 nats.","key_machinery":"The central object is the mutual information MI($x, z_H$) between a galaxy property $x$ and the halo assembly time $z_H$, estimated with the GMM-MI procedure, which fits Gaussian Mixture Models to the joint distribution and uses bootstrap resampling for uncertainties; unlike Pearson correlation, it captures any nonlinear dependence. The target variable $z_H$ is defined as the redshift at which a halo's main branch reaches half of its $z=0$ mass, tracked through subhalo merger trees. The winning property is $F_\\star = \\log_{10}(M_\\star/M_h)$, the galaxy formation efficiency, compared against galaxy assembly time $z_G$, colour $(g-i)$, sSFR, and cluster observables (magnitude gap, satellite distances, and richness) in stellar-mass bins.","core_discovery":"The central claim is that mutual information analysis of the TNG300-1 sample establishes $F_\\star$ as a more sensitive indicator of halo assembly history than colour, sSFR, or cluster observables. For central galaxies with stellar masses up to about $10^{10.25}\\,h^{-1}M_\\odot$, MI($F_\\star, z_H$) rises from $0.38$ to $0.43$ nats with a Pearson coefficient of $0.6$–$0.8$, whereas MI for $(g-i)$ peaks near $0.18$ nats and sSFR is effectively negligible. The authors interpret the signal as a co-evolutionary imprint: in low-mass haloes, gas accretion and supernova-regulated star formation tie the growth of the central galaxy to the assembly of the halo. The correlation weakens above the characteristic stellar mass, where AGN feedback and mergers dominate, and it disappears for satellites, where environmental quenching decouples galaxy properties from host-halo history. Among cluster observables, richness is the only quantity whose information about $z_H$ grows with stellar mass, consistent with satellite accretion dominating the late growth of massive haloes. The paper also reports that for lower-mass galaxies the $F_\\star$–$z_H$ correlation appears primarily driven by the known halo mass–$z_H$ correlation, which is the main internal caveat to the claim of independence.","pith_inferences":["The paper does not report mutual information conditioned on halo mass, so its headline claim is not yet fully separated from the known halo mass–assembly time relation; a conditional-MI test would settle whether $F_\\star$ adds independent information.","An observational programme could estimate $F_\\star$ from weak-lensing halo masses plus stellar masses in group catalogues and test the same ranking; if it holds, $F_\\star$ should outperform the magnitude gap as a halo-age proxy at fixed stellar mass.","The same MI machinery could rank other observables, such as morphology, central velocity dispersion, size, or environment density, against $z_H$, extending the search for a robust assembly-history tracer.","Because assembly time is defined as the half-mass redshift, the ranking of proxies might shift if a different formation threshold such as $z_{25}$ or $z_{75}$ were used; that would test whether the $F_\\star$ advantage is robust to the definition."],"forward_implications":["For central galaxies below about $10^{10.25}\\,h^{-1}M_\\odot$, $F_\\star$ provides a stronger statistical handle on halo assembly time than colour or sSFR, with mutual information reaching $0.38$–$0.43$ nats and a Pearson correlation of $0.6$–$0.8$.","Standard age-distribution matching proxies, colour and sSFR, carry less than half the information about $z_H$ in this mass range, so models that map colour or sSFR to halo age will be weak tracers of assembly history.","Satellite galaxy properties carry almost no information about host-halo assembly time, meaning environmental processing erases the assembly-time signal and assembly-bias studies should avoid satellites as halo-age proxies.","Among cluster observables, richness is the only quantity whose information about $z_H$ grows with stellar mass, pointing to satellite accretion as the dominant growth channel of massive haloes.","The decline of the $F_\\star$–$z_H$ correlation above about $10^{10.25}\\,h^{-1}M_\\odot$ marks the mass scale where AGN feedback and mergers start to dominate over the assembly-time signal."],"supporting_citations":[{"why":"Provides the TNG300-1 simulation data, subhalo catalogues, and merger trees that supply every galaxy property and halo assembly time analysed in the paper.","marker":"Nelson et al. 2019"},{"why":"Supplies the GMM-MI estimator, Gaussian Mixture Models with bootstrap, used to compute the mutual information values and uncertainties.","marker":"Piras et al. 2023"},{"why":"Introduced the stellar-to-halo mass ratio fc as an observational proxy for halo assembly time, the direct precursor of F⋆ that the paper extends with MI.","marker":"Lim et al. 2016"},{"why":"Establishes that the half-mass assembly time zH is closely tied to halo internal structure, justifying zH as the target variable.","marker":"Wang et al. 2011"},{"why":"Found that scatter in the stellar-to-halo mass relation correlates with halo assembly time, motivating the F⋆–zH connection.","marker":"Matthee et al. 2017"},{"why":"Provides the IllustrisTNG magnitude-gap analysis that the cluster-observable comparison attempts to contextualise.","marker":"Farahi et al. 2020"},{"why":"Supports the in situ versus ex situ growth interpretation used to explain richness and F⋆ trends in massive haloes.","marker":"Moster et al. 2020"}],"fun_headline_variants":["F_star beats color and sSFR as halo-assembly tracer","Mutual information ranks F_star top for halo assembly history","Galaxy efficiency reveals halo assembly, color and sSFR lag","In IllustrisTNG, F_star outpaces color and sSFR in tracing halo formation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that $F_\\star$'s strong mutual information with halo assembly time is a genuine, independent imprint of assembly history, but the paper itself notes that for lower-mass galaxies this correlation appears primarily driven by the known correlation between halo mass and halo assembly time.","fun_headline_variants_meta":{"raw":{"variants":["F_star beats color and sSFR as halo-assembly tracer","Mutual information ranks F_star top for halo assembly history","Galaxy efficiency reveals halo assembly, color and sSFR lag","In IllustrisTNG, F_star outpaces color and sSFR in tracing halo formation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000197,"raw_usage":{"total_tokens":1469,"prompt_tokens":1155,"completion_tokens":314,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":771,"completion_tokens_details":{"reasoning_tokens":232}},"tokens_in":771,"tokens_out":314,"duration_ms":3627,"temperature":1.0,"reasoning_tokens":232,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T16:49:14.198147+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute MI($F_\\star, z_H$) within narrow bins of halo mass, or the conditional MI($F_\\star, z_H\\mid M_h$), for central galaxies below $10^{10.25}\\,h^{-1}M_\\odot$; if it drops to the level of colour or sSFR, the claim that $F_\\star$ is a more sensitive independent indicator fails. The equivalent observational check is to build group catalogues with weak-lensing halo masses and see whether $F_\\star$ beats colour at fixed $M_h$.","supporting_citations":[{"cited_title":"V., Pontzen A., Lucie-Smith L., Guo N., Nord B., 2023, @doi [Machine Learning: Science and Technology] 10.1088/2632-2153/acc444 , 4, 025006","cited_arxiv_id":null,"evidence_quote":"Supplies the GMM-MI estimator, Gaussian Mixture Models with bootstrap, used to compute the mutual information values and uncertainties."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the IllustrisTNG magnitude-gap analysis that the cluster-observable comparison attempts to contextualise."}],"review_version":1}