{"id":"a335d1ec-2ea7-4142-85f8-3723385dfd75","arxiv_id":"2411.13484","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"The simple linear galaxy size model, combined with abundance matching on peak halo mass, secretly encodes halo formation time into galaxy sizes, so the observed size-split clustering pattern reflects assembly bias rather than a direct size-to-radius link.","lead":"Galaxies split by size in a standard dark matter simulation cluster differently because a simple size model secretly depends on when each halo formed, not just on its current radius. The work explains why such a model can mimic observed galaxy clustering and warns that clustering alone cannot identify what controls galaxy size.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The explanation of the high-mass clustering convergence depends on the adopted SHMR high-mass slope; if the true SHMR differs, the halo bias cancellation may not occur.","rationale":"The paper's core mechanism - that using Mpeak for abundance matching and r_Mpeak for sizes introduces an implicit dependence on a_Mpeak - is clearly demonstrated by Equation 2 and Figure 4, and is robust to the external SHMR. The control experiment in Appendix A (removing assembly bias by selecting a_Mpeak=1 and reweighting Mpeak) provides credible evidence that the clustering difference at low stellar mass is indeed due to assembly bias rather than halo bias alone. I therefore do not question the mechanism itself. The load-bearing concern is about the second half of the central claim: that this mechanism explains the observed size-split clustering pattern, especially the high-mass convergence. That explanation depends on the SHMR's high-mass slope, which is adopted from Moustakas et al. (2013) and Mortlock et al. (2011) via abundance matching with 0.2 dex scatter. Figure 3 (right panel) shows that the Mpeak distributions of small and large galaxies separate only at high stellar mass, and Section 3.2 attributes this to the shallow SHMR slope. If the true SHMR is different, the halo bias contribution would change and the convergence might disappear, meaning the model would not reproduce the observed pattern. The paper does not include a sensitivity analysis to the SHMR or scatter, so the explanation is conditional. This matches the reader's weakest_assumption, and no new concern changes the verdict.","tokens_in":1115,"tokens_out":924,"duration_ms":169788,"concrete_test":"Re-run the full model of Section 2.2 using an alternative abundance matching calibration (e.g., the Behroozi et al. 2019 SHMR/SMF instead of Moustakas et al. 2013 and Mortlock et al. 2011), keeping all other choices fixed, and recompute Figures 2 and 3. If the Mpeak distributions of small and large galaxies no longer separate at high stellar mass, or the high-mass wp gap reverses, then the explanation of the observed convergence is not robust to the assumed SHMR.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central explanation in Sections 3.2-3.3 attributes the high-mass convergence of size-split clustering to the shallow high-mass slope of the abundance-matched SHMR (Figure 1), which makes the Mpeak distributions of large and small galaxies separate at high stellar mass (Figure 3, right panel). This separation is what allows halo bias to cancel the assembly bias signal. However, the SHMR is not measured directly; it is an output of the abundance matching procedure of Section 2.2.1 that adopts the Moustakas et al. (2013) and Mortlock et al. (2011) stellar mass functions and a 0.2 dex scatter. If the true high-mass SHMR slope is steeper (e.g., Behroozi et al. 2019) or the scatter is larger, the Mpeak distributions of small and large galaxies may separate less at fixed stellar mass, weakening the halo-bias cancellation. The paper does not test this sensitivity, so the claim that the K13 model reproduces the observed size-split clustering pattern is conditional on the adopted external SHMR.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper addresses a puzzle in the galaxy–halo connection: why a simple linear galaxy-size–halo-radius model (Kravtsov 2013, as used by Hearin et al. 2019) reproduces the observed size-split galaxy clustering pattern even though it contains no explicit halo formation-time dependence. Using the VSMDPL simulation and the H19 modeling pipeline, the authors show that when stellar masses are assigned via subhalo abundance matching with M_peak and sizes are assigned with r_Mpeak, the size model implicitly encodes the halo formation scale factor a_Mpeak through the time-dependent M_vir–r_vir relation. As a result, at fixed stellar/halo mass, smaller galaxies preferentially occupy earlier-forming halos, injecting halo assembly bias into the clustering signal. The paper argues that at low stellar masses this assembly bias makes small galaxies more clustered than large ones, while at high stellar masses the shallow slope of the abundance-matched SHMR separates the M_peak distributions of small and large galaxies, so that halo bias partially cancels the assembly bias and the clustering gap shrinks. The authors further show that an r_vir-based size model produces nearly identical clustering and size–mass relations, and they discuss the difficulty of identifying the specific halo property controlling galaxy size from clustering alone.","tokens_in":16557,"tokens_out":3998,"duration_ms":43085,"significance":"If the mechanism identified here is correct, the paper provides a clean explanation of a previously puzzling result: the success of the simple K13 size model does not imply that galaxy sizes are directly governed by halo radius alone, because the abundance-matching step smuggles in a formation-time dependence. The control experiment in Appendix A (restricting to a_Mpeak = 1 and reweighting to preserve M_peak distributions) is a strong, targeted test that supports the assembly-bias interpretation, and the model is not circular: the size-model constant and abundance-matching scatter are inherited from prior work, and the clustering gap is an output, not an input. The demonstration that r_Mpeak and r_vir models are nearly degenerate in both clustering and size–mass relations is a useful caution for the field. However, the central explanation of the high-mass convergence depends on the adopted SHMR, and the paper does not quantify how sensitive this conclusion is to that external input.","major_comments":[{"comment":"The explanation of the high-mass convergence relies on the M_peak distributions of small and large galaxies separating at high stellar mass (right panel of Fig. 3). This separation is a consequence of the shallow high-mass slope of the SHMR produced by the abundance-matching procedure of §2.2.1 with the Moustakas et al. (2013) and Mortlock et al. (2011) stellar mass functions and a fixed 0.2 dex scatter. The authors do not test how a steeper high-mass SHMR (e.g., Behroozi et al. 2019) or a larger scatter would change the M_peak distributions and thereby weaken the halo-bias cancellation. Without such a sensitivity test, the claim that the K13 model reproduces the observed size-split clustering pattern is conditional on the adopted external SHMR.","section":"§3.2, §3.3, and Fig. 3"},{"comment":"Figure 2 shows no error bars on the mock clustering measurements, and the paper does not display the SDSS data against which the model is said to match 'reasonably well'; the comparison is only asserted through H19. Since the central narrative turns on the presence and disappearance of the clustering gap between small and large galaxies, the authors should provide a direct quantitative comparison or at least an estimate of the uncertainty (e.g., jackknife or bootstrap) so that the reader can judge whether the reported gap is significant and whether the claimed match to observations holds.","section":"Fig. 2 and §3"},{"comment":"The conclusion that 'small galaxies have to occupy halos that form early' is stated as a general inference in §4.1 and the summary, but the argument as presented applies within the specific assumptions of the K13/H19 size model: if one relaxes the premise that galaxy size traces halo radius at fixed stellar mass (e.g., allowing morphology or other baryonic effects to set size), the dichotomy between 'more massive halo' and 'earlier-forming halo' is not exhaustive. The authors should soften this claim or explicitly limit it to the class of models considered here.","section":"§4.1, summary"}],"minor_comments":[{"comment":"Throughout this section the term 'viral radius' should be 'virial radius' (e.g., in the sentence defining r_1/2 = 0.01 r_vir).","section":"§2.2.2"},{"comment":"The summary refers to the 'shallower SMHR' at high stellar mass; the abbreviation should be SHMR for consistency with the rest of the paper.","section":"§5"},{"comment":"The simulation name is written as 'VSDMPL' in the text but later as 'VSMDPL'; please use one consistent spelling.","section":"§2.1"},{"comment":"The statement that the predicted size evolution 'appears to be greater than one would expect' is not supported by a quantitative comparison to the cited observational constraints (Huang et al. 2017; Martorano et al. 2024); adding a brief quantitative statement would strengthen the point.","section":"§4.2"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid contribution to the galaxy–halo connection literature, and the core mechanistic explanation appears sound. The main risk is that the high-mass cancellation explanation is tied to the adopted SHMR without a sensitivity test; I would be comfortable with acceptance after the authors add a robustness check or explicitly temper the claim. I do not see grounds for rejection, and the paper's strengths (clear control experiment, non-circular model setup, useful degeneracy discussion) should be credited in the final decision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a clean, honest piece of model forensics. The genuinely new thing is the mechanism they identify: because r_Mpeak = r_vir(a_Mpeak) and the Mvir-rvir relation is time-dependent, using Mpeak for abundance matching while setting galaxy size ∝ r_Mpeak secretly injects a halo formation-time dependence into the size assignment. That explains why the K13 linear model reproduces the H19 size-split clustering without any explicit assembly-bias ingredient. The control experiment in Appendix A (restricting to a_Mpeak=1) really does support the assembly-bias interpretation, and the demonstration that r_Mpeak and r_vir models give nearly identical clustering and size-mass relations is a nice, practical point that echoes Behroozi et al. (2022).\n\nThe paper is also transparent: public simulation, public codes, no parameter tuned to the clustering signal. The free parameters come from prior work, so circularity is not a worry.\n\nSoft spots, in proportion. The most consequential is the one the stress-test note flags: the high-mass convergence of large and small galaxy clustering relies on the shallow high-mass slope of the SHMR produced by the adopted stellar mass functions and 0.2 dex scatter. If the true high-mass SHMR is steeper or the scatter larger, the Mpeak separation between size-split samples at fixed M* would shrink and the halo-bias cancellation would weaken. The paper does not explore that sensitivity. That does not overturn the low-mass assembly-bias result, since at low masses the Mpeak distributions overlap irrespective of reasonable SHMR choices, but it does make the quantitative high-mass prediction less secure than the text implies.\n\nTwo smaller issues. Figure 2 has no error bars or cosmic-variance estimate; since the comparison is between subsamples from one simulation, some errors cancel, but a bootstrap or jackknife would help. And the match to SDSS is asserted via H19 rather than shown; the reader has to trust that their mock reproduces the observed data. Both are addressable in revision.\n\nNet: this is a useful conceptual paper for the galaxy-halo connection subfield. It deserves a serious referee. I'd cite it and would send it out rather than desk reject.","headline":"Useful model forensics: explains the K13 size-split clustering via implicit assembly bias, with a caveat that the high-mass cancellation leans on the adopted SHMR.","tokens_in":17088,"tokens_out":2741,"would_cite":true,"duration_ms":28012,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that a simple linear galaxy-size–halo-radius model, when combined with peak-mass abundance matching, implicitly encodes halo formation time and thereby reproduces the observed size-split clustering pattern without…","keywords":["galaxy sizes","halo assembly bias","subhalo abundance matching","galaxy clustering","dark matter halos","stellar-to-halo mass relation","projected correlation function","K13 size model"],"falsifier":"A concrete calculation is to recompute the high-mass size-split clustering after replacing the adopted high-mass stellar mass function with one whose slope is shifted by its quoted uncertainty; if the convergence of large- and small-galaxy clustering disappears or moves by more than the clustering error bars, the mechanism's reliance on the stellar-to-halo mass relation is falsified. A direct observational check is to measure the size-split correlation function at $\\log(M_\\star/M_\\odot)\\gtrsim 11$ with precision high enough to confirm the predicted convergence.","tokens_in":1787,"feed_emoji":"🌌","tokens_out":4644,"duration_ms":112512,"temperature":0.7,"pith_summary":"This paper tries to explain why a deliberately simple recipe for galaxy sizes—make the half-light radius a fixed fraction of a halo's virial radius at peak mass—reproduces observed clustering differences between small and large galaxies at fixed stellar mass. The explanation is that the recipe is not as simple as it looks: because the virial mass–radius relation depends on cosmic time, using the radius at peak mass makes a galaxy's size sensitive to when the halo reached peak mass. Small galaxies preferentially sit in earlier-forming halos, whose clustering is boosted by halo assembly bias. At high stellar mass this assembly-bias boost is balanced by ordinary halo bias favoring large galaxies in more massive halos. The consequence is that matching clustering data forces small and large galaxies to differ in halo assembly history, not just halo mass.","feed_headline":"Simple size models hide assembly bias in halo peak timing","feed_subtitle":"Using peak-mass radii makes galaxy sizes depend on when halos assembled, explaining the observed clustering gap.","key_machinery":"The central object is the time-dependent virial mass–radius relation, $M_{\\rm peak} = (4\\pi/3)\\,r_{M_{\\rm peak}}^3\\,\\Delta_{\\rm vir}(a_{M_{\\rm peak}})\\,\\rho_{\\rm crit}(a_{M_{\\rm peak}})$, together with the size prescription $r_{1/2}=0.01\\,r_{M_{\\rm peak}}$. Because $\\Delta_{\\rm vir}$ and $\\rho_{\\rm crit}$ depend on the scale factor at peak mass, this pair of equations converts a halo's assembly time into its assigned galaxy size at fixed peak mass. The mechanism carries the argument by showing that the size split at fixed stellar mass is also a split in $a_{M_{\\rm peak}}$, which is the property that halo assembly bias acts on.","core_discovery":"On the paper's own terms, the discovery is that the linear size model of Kravtsov (2013), in which $r_{1/2}=0.01\\,r_{M_{\\rm peak}}$, implicitly depends on halo formation history when stellar masses are assigned by subhalo abundance matching on $M_{\\rm peak}$. The time-dependent spherical-overdensity relation means that among halos of the same $M_{\\rm peak}$, those that reached peak mass earlier have smaller $r_{M_{\\rm peak}}$ and hence smaller modeled galaxies. Since earlier-forming halos are more clustered at fixed mass, small galaxies are predicted to be more clustered at lower stellar masses, and at higher stellar masses that trend is offset by the larger $M_{\\rm peak}$ of large galaxies as the stellar-to-halo mass relation flattens. The paper also finds that replacing $r_{M_{\\rm peak}}$ with present-day $r_{\\rm vir}$ changes little, because tidal stripping introduces a similar assembly-history dependence. The conclusion is that any size model matching the observed size-split clustering must effectively separate galaxies by halo assembly history, and clustering alone cannot identify which halo property controls size.","pith_inferences":["One testable extension is to recompute the predicted size-split clustering after perturbing the high-mass slope of the adopted stellar-to-halo mass relation; the stellar mass where the clustering gap closes would shift, allowing existing surveys to test the mechanism.","The near-degeneracy between the $r_{M_{\\rm peak}}$ and $r_{\\rm vir}$ models suggests that any secondary halo property strongly correlated with $a_{M_{\\rm peak}}$, such as concentration, can substitute for formation time in a size model while leaving clustering predictions nearly unchanged, so clustering alone cannot break that degeneracy.","A direct observational test would measure size segregation in overdense environments: the model predicts that, at fixed stellar mass, small galaxies should be more abundant than large galaxies in dense regions, which can be checked with group catalogs and redshift surveys."],"forward_implications":["At fixed stellar mass, small modeled galaxies occupy halos with earlier $a_{M_{\\rm peak}}$ at all four mass thresholds, so the size-split samples differ in assembly history even though the size model never uses formation time.","At low stellar mass the relative halo bias between size-split samples is weak because the stellar-to-halo mass relation is steep, so the clustering gap is dominated by assembly bias.","At high stellar mass the stellar-to-halo mass relation flattens, large galaxies occupy more massive halos, and the resulting halo bias offsets assembly bias, making large and small galaxies cluster similarly.","Using present-day $r_{\\rm vir}$ instead of $r_{M_{\\rm peak}}$ gives nearly identical clustering and size–mass relations because the amount of tidal stripping is strongly correlated with $a_{M_{\\rm peak}}$.","If assembly bias is artificially removed by selecting halos with $a_{M_{\\rm peak}}=1$, large galaxies cluster slightly more than small galaxies at all stellar masses, confirming that assembly bias drives the low-mass gap."],"supporting_citations":[{"why":"Supplies the linear size model $r_{1/2}=0.01\\,r_{M_{\\rm peak}}$ that the paper explains and tests.","marker":"K13"},{"why":"Provides the observational size-split clustering pattern and the model setup that this work reproduces and interprets.","marker":"H19"},{"why":"Justifies using peak halo mass as the abundance-matching property and the scatter treatment adopted here.","marker":"Reddick et al. (2013)"},{"why":"Establishes halo assembly bias as the physical effect that makes earlier-forming halos more clustered at fixed mass.","marker":"Wechsler et al. (2006)"},{"why":"Provides an alternative size model with explicit assembly-history dependence, used as the comparison motivating the explanation of the simple model.","marker":"Jiang et al. (2019)"},{"why":"Introduces a growth-rate-based size model and supports the conclusion that clustering alone cannot distinguish which halo property controls size.","marker":"Behroozi et al. (2022)"},{"why":"Supplies the low-redshift stellar mass functions used in the abundance-matching step.","marker":"Moustakas et al. (2013)"},{"why":"Supplies the higher-redshift stellar mass functions used for the $z\\geq 1$ predictions.","marker":"Mortlock et al. (2011)"}],"fun_headline_variants":["Size model ties galaxy clustering to halo age","Assembly bias slips into simple size prescriptions","Growth history biases galaxy size clustering","Clustering alone can't pin down size controls","Peak mass and formation time decide galaxy sizes"],"cache_read_input_tokens":19328,"weakest_assumption_plain":"The load-bearing premise is that the abundance-matched stellar-to-halo mass relation, especially its high-mass slope, is accurate; if the high-mass slope is wrong, the $M_{\\rm peak}$ distributions of large and small galaxies at fixed stellar mass shift, and the predicted cancellation between halo bias and assembly bias at high stellar mass fails.","fun_headline_variants_meta":{"raw":{"variants":["Size model ties galaxy clustering to halo age","Assembly bias slips into simple size prescriptions","Growth history biases galaxy size clustering","Clustering alone can't pin down size controls","Peak mass and formation time decide galaxy sizes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000382,"raw_usage":{"total_tokens":2064,"prompt_tokens":1023,"completion_tokens":1041,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":639,"completion_tokens_details":{"reasoning_tokens":976}},"tokens_in":639,"tokens_out":1041,"duration_ms":8625,"temperature":1.0,"reasoning_tokens":976,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T16:21:51.210052+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete calculation is to recompute the high-mass size-split clustering after replacing the adopted high-mass stellar mass function with one whose slope is shifted by its quoted uncertainty; if the convergence of large- and small-galaxy clustering disappears or moves by more than the clustering error bars, the mechanism's reliance on the stellar-to-halo mass relation is falsified. A direct observational check is to measure the size-split correlation function at $\\log(M_\\star/M_\\odot)\\gtrsim 11$ with precision high enough to confirm the predicted convergence.","supporting_citations":[],"review_version":1}