{"id":"1f497167-327a-4823-8b4c-df54cf18725b","arxiv_id":"2508.13274","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"In Kepler multi-planet systems, inner planets tend to be smaller than outer planets, a trend the authors argue is not solely an observational artifact.","lead":"This paper analyzes the order of planet sizes in multi-planet systems from NASA's exoplanet catalog, finding that inner planets are usually smaller than outer ones. It also reports that this pattern depends on the host star's metallicity, while planets near orbital resonances show no special size pattern.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The strongest evidence against a pure bias explanation is the synthetic single-planet pairing in §6.3, but it is not a Kepler completeness forward model; the claim that the ordering is 'not solely a product of observational biases' rests on an untested selection null.","rationale":"The reader's weakest assumption points to the hand-chosen de-biased region and the synthetic pairing; my stress test sharpens that into a specific, load-bearing gap: neither control is an injection-recovery or forward-model test of the Kepler selection function. The simple counts in Tables 1–3 are real and consistent with earlier work (e.g., Ciardi et al. 2013), but the transition from 'a trend exists in the archive' to 'the trend is not solely a product of observational biases' requires a null model that actually includes the selection effects that produce detected two- and three-planet systems. The paper's own summary (Section 7, item 4) concedes that debiasing significantly affects three-planet systems and outer pair ordering, which reinforces the need for a full selection model. I do not see an internal inconsistency in the counting or the Anderson-Darling comparisons, and the metallicity and resonance results are honestly labeled as small-statistics findings. The central concern is therefore a correctness risk in the bias control, not a disagreement with the field's consensus. Since the reader already issued a CONDITIONAL verdict, my recommendation is unchanged: the paper should be accepted conditional on a forward-model bias test and explicit multiple-testing treatment for the pair subsamples.","tokens_in":24429,"tokens_out":5184,"duration_ms":58667,"concrete_test":"Use the Kepler DR25 (or Petigura et al. 2013) completeness map to forward-model the selection function: (i) generate synthetic multi-planet systems under a null hypothesis of random pairwise radius ordering, with period and radius marginals matched to the observed sample; (ii) assign mutual inclinations from the observed Kepler multiplicity distribution; (iii) compute each planet's detection probability from transit geometry and the completeness map, then apply the same cuts (R > 2 R_Earth, P < 50 d) and pair-counting used in §3–§5; (iv) repeat for alternative intrinsic orderings (inner-smaller, outer-smaller). Compare the recovered fraction of '12' configurations to Table 1's 74.6% (full) and 69.2% (de-biased). If the random-ordering simulation reproduces the observed excess, the central claim's bias-free interpretation fails; if it does not, the concern is retired.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim in Section 6 — that the inner-smaller ordering is 'not solely a product of observational biases' — leans on two controls: the hand-selected 'de-biased' box (R > 2 R_Earth, P < 50 d; §3.1) and the synthetic single-planet pairing (§6.3). Neither inverts the Kepler selection function. The de-biased cuts remove small planets but do not correct the within-pair detectability gradient: for two transiting planets, the outer planet has a longer period, fewer transits, and a lower geometric transit probability, so a population with random (or even mildly inner-larger) radii will preferentially yield detected pairs with larger outer planets. The paper itself concedes the box 'may still experience some minor selection effects' (§3.1), and the Fisher tests in §5.1 only show that applying the cut does not significantly change the configuration mix — not that the mix is unbiased. The synthetic test in §6.3 is likewise not a completeness model: it pairs planets drawn from single-planet systems, preserving the marginal period/radius distribution of single-planet detections, but single- and multi-planet samples have different selection functions (multiplicity-dependent completeness, mutual inclination, pipeline efficiency), and the synthetic pairs are never passed through a detection pipeline. Demonstrating that observed ordering differs from this synthetic null does not establish that an unbiased population would produce the observed ordering. The 'intrinsic' part of the claim therefore rests on an untested selection model.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper analyzes the relative radii ordering of planets within multi-planet systems, using NASA Exoplanet Archive data, and focuses on Kepler systems with two to four planets. Systems are classified by the sequence of planet sizes from the innermost to outermost orbit (e.g., \"12\" vs \"21\" for pairs), and the analysis is repeated on a hand-defined \"de-biased\" sample with R > 2 R_Earth and P < 50 days. The paper reports central counts of 273 vs 93 for \"12\" vs \"21\" in the full two-planet sample and 153 vs 68 in the de-biased sample, finds that the trend is strongest for inner pairs in three-planet systems, claims a metallicity dependence using a split at [Fe/H] = -0.2, finds no significant difference between resonant and non-resonant pairs, and interprets the results as evidence that larger planets form farther out and migrate inward. The central conclusion is that the inner-smaller ordering is intrinsic and not solely a product of observational biases.","tokens_in":24750,"tokens_out":5513,"duration_ms":59545,"significance":"If the intrinsic-ordering claim were established, the paper would introduce a useful and underused observable: the relative size ordering within a system, which can constrain formation, migration, and subsequent dynamical evolution. The paper has clear strengths: transparent contingency tables with counts, use of public data, simple and reproducible statistical tests, and a falsifiable synthetic null. The observed trend in the Kepler sample is genuinely interesting and worth reporting. However, the load-bearing robustness claim rests on incomplete bias controls: the de-biased box does not remove the within-pair detectability gradient, the synthetic single-planet pairing is not a Kepler completeness forward model, and the metallicity split is chosen post hoc. The paper is therefore best viewed as a solid observed-trend study whose interpretation as an intrinsic property needs substantially more support.","major_comments":[{"comment":"The de-biased sample does not establish that the trend is unbiased. The cuts R > 2 R_Earth and P < 50 days remove small planets but do not correct the within-pair detection gradient: for two transiting planets, the outer planet has a longer period, fewer transits, and a lower geometric transit probability, so a population with random or even mildly inner-larger radii will preferentially yield detected pairs with larger outer planets. The Fisher exact tests in §5.1 (Table 8) only show that applying the cut does not significantly change the configuration mix compared with the full sample (p-values 0.2219 to 0.8918), not that the mix equals the intrinsic population. Since the central claim in §6 that the trend is \"not solely a product of observational biases\" rests on this control, that claim is currently unsupported. The paper itself concedes in §3.1 that the region \"may still experience some minor selection effects,\" but the relevant point is that the selection effect is not minor for the ordering statistic.","section":"§3.1, §5.1, §6"},{"comment":"The synthetic single-planet pairing is not a Kepler completeness forward model. The synthetic pairs are drawn from single-planet detections, preserving the marginal period and radius distributions, but the pairs are never passed through a detection pipeline, and single-planet and multi-planet samples have different selection functions (multiplicity-dependent completeness, mutual inclination, pipeline efficiency). Demonstrating that the observed ordering differs from this synthetic null only shows non-random pairing relative to that particular null; it does not establish that an unbiased population would produce the observed ordering. The text in §6.3 states that the synthetic sample \"was constructed to replicate these biases,\" but the construction replicates only the marginal distributions of single-planet detections, not the joint detection probability for pairs. This is load-bearing for the \"intrinsic\" part of the central claim.","section":"§6.3, Table 13, Figure 15"},{"comment":"The metallicity split at [Fe/H] = -0.2 is chosen post hoc from the same data used to test metallicity dependence. The text in §4.3 states that the value \"divides the two-planet sample into two roughly equal parts,\" and this same split is then used in §6.1 to claim a significant metallicity dependence of the radius-ratio distribution (Table 12). This is circular for the metallicity claim. The analysis should either use a pre-specified threshold, demonstrate robustness across a range of thresholds, or otherwise treat the split as a discovery that requires independent confirmation. In addition, the many pairwise Anderson-Darling and Fisher tests in Tables 9-12 are reported without any multiple-testing correction, so some of the \"significant\" results, including the metallicity contrast, may be chance findings.","section":"§4.3, §6.1, Figures 13-14"},{"comment":"The analysis compares planetary radii without propagating their uncertainties. For typical Kepler radius uncertainties of several percent to ten percent, many adjacent planets in a system may be consistent with equal sizes, and the configuration labels \"12\" versus \"21\" as well as the ratios R_in/R_out can flip under plausible radius errors. The central count statistics, such as the 273 vs 93 in Table 1, therefore mix real ordering signal with measurement noise. The robustness of the trend should be checked by Monte Carlo resampling of radii within published uncertainties, or by restricting the ordering analysis to pairs with a radius difference that is significant at, say, the 2-sigma level. Without such a test, the quantitative strength of the trend is uncertain.","section":"§3.1, Tables 1-3, Figures 3-4"}],"minor_comments":[{"comment":"Please clarify whether the \"full\" sample is restricted to Kepler detections or includes all missions in the NASA Exoplanet Archive. Section 3.1 describes a \"wide and heterogeneous sample\" while the abstract and Section 6 refer specifically to \"Kepler multi-planet systems.\" If the full sample contains non-Kepler planets, the comparison between full and de-biased samples mixes different selection functions.","section":"§3.1 and abstract"},{"comment":"The text says that shuffling the two-planet sample 100 times results in 34,600 synthetic pairs, but 366 pairs repeated 100 times gives 36,600 pairs; please correct the arithmetic.","section":"§5.2"},{"comment":"The description of the bold p-values is inconsistent: the text says bold values indicate failure to reject the null hypothesis, while the table caption says bold values indicate rejection of the null hypothesis. Please align the notation with the intended meaning.","section":"Table 11 and surrounding text"},{"comment":"The figure caption and the text disagree about which region is the de-biased sample: Figure 1 shows a shaded rectangle R > 2 R_Earth and P < 50 days, but the following paragraph says \"the planet above the dashed green line is part of the de-biased sample,\" which describes a different selection. Please correct the wording.","section":"Figure 1 and §3.1"},{"comment":"There are several typographical errors, including \"de-baised\" for \"de-biased\" and \"refereed\" for \"referred,\" and the phrase \"we select planes in a region\" should be \"we select planets in a region.\"","section":"§3.1"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within the journal's scope and the observed ordering trend is interesting. The main concern is the gap between the data analysis and the strong intrinsic-ordering claim; a revision that reframes the central claim as an observed trend in a biased sample, or adds a genuine completeness/sensitivity model, would make the paper much stronger. The novelty of the ordering observable itself is not in question."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Hi,\n\nQuick take on the Lozovsky & Perets paper (arXiv:2508.13274). The observational trend is likely real: inner planets in Kepler multi-planet systems tend to be smaller than outer ones. The paper gives the most complete catalog of that ordering for 2- and 3-planet systems so far, including the 123/132-type configuration counts and the ab/ac/bc pair decomposition. It also reports a useful null result for resonance proximity and a metallicity split, both preliminary. The writing is clear, the counts are simple enough to check by hand, and the authors are honest about some of their cuts.\n\nWhat I don't buy is the central robustness claim in Section 6 and the Summary: that the trend is \"not solely a product of observational biases.\" The controls do not remove the within-pair detectability gradient. For a two-planet transiting system, the outer planet has a longer period and a lower geometric transit probability, so a detected pair is preferentially a pair with a larger outer planet. The \"de-biased\" box (R > 2 R_Earth, P < 50 days) removes small planets but keeps that gradient. The synthetic single-planet pairing in Section 6.3 is a nice sanity check, but it never sends the synthetic pairs through a transit-detection pipeline; it pairs planets that are already known detections. So the fact that observed ordering differs from that synthetic null does not tell us what an unbiased population would produce. The paper's own caveat that the box \"may still experience some minor selection effects\" understates the issue.\n\nOther soft spots: radius error bars are ignored, which matters for individual classifications near the 12/21 boundary; the metallicity threshold [Fe/H] = -0.2 is chosen from the same data it is used to test; and the many sub-sample comparisons are not corrected for multiple testing. None of these are fatal, but they push the paper into \"interesting observed trend, selection contribution unresolved\" territory.\n\nWho should read it: exoplanet population and formation folks. It consolidates and extends earlier work like Ciardi et al. 2013, and it gives theorists a concrete set of ordering statistics to reproduce. It deserves a serious referee, but the referee should ask for a forward-model bias test or a forthright downgrade of the \"intrinsic\" claim. I'd send it out with a minor-to-major revision request rather than accept as is.","headline":"Useful catalog of size-ordering trends in Kepler multi-planet systems, but the paper overreaches when it claims the inner-smaller trend is not primarily a selection artifact; the bias controls don't actually test the selection function.","tokens_in":25294,"tokens_out":4142,"would_cite":false,"duration_ms":42325,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Kepler multi-planet systems show a persistent tendency for inner planets to be smaller than outer planets, even after accounting for detection biases.","keywords":["exoplanet systems","planet size ordering","Kepler multi-planet systems","planet formation","orbital migration","planetary resonances","stellar metallicity","observational biases"],"falsifier":"Forward-model Kepler's detection efficiency: inject synthetic multi-planet populations with no intrinsic radial size gradient, run them through the selection function, and check whether the observed '12' versus '21' counts and the $R_{\\mathrm{in}}/R_{\\mathrm{out}}$ distributions can be reproduced by bias alone. If they can, the claim that the ordering is intrinsic collapses; if the inner-smaller trend disappears when the same ordering analysis is applied to a sample with an independent selection function, the de-biasing is insufficient to establish it.","tokens_in":24195,"feed_emoji":"🪐","tokens_out":8627,"duration_ms":82471,"temperature":0.7,"pith_summary":"This paper argues that within multi-planet exoplanet systems, planet sizes are not arranged by chance: inner planets tend to be smaller than their outer neighbours, and this 'smaller-inner' ordering is present even after restricting the sample to a region where transit detection biases are weak. The evidence comes from Kepler systems with two to four planets in the public exoplanet catalog, classified by the relative-size sequence of their planets and by the radius ratios of every planet pair. If the trend is real, the ordering of planet sizes inside a system becomes a new observational handle on planet formation and evolution, favouring models in which larger planets assemble farther out and migrate inward while photoevaporation shrinks close-in planets. The paper also finds that the ordering depends on host-star metallicity and that resonant pairs do not have a distinctive size-ratio distribution, which it reads as possible evidence that resonant chains formed early and were later destabilized.","feed_headline":"Kepler data show inner planets are the smaller ones in their systems","feed_subtitle":"The size ordering survives de-biasing and ties planet formation, migration, and stellar metallicity together.","key_machinery":"The argument runs on ordinal size configurations combined with pair radius-ratio distributions. Each system is assigned a sequence such as '12' for a smaller inner planet or '321' for a largest innermost planet, and the frequencies of these configurations are compared across the full and de-biased samples, stellar types, metallicities, and planet multiplicities. For each pair, the radius ratio $R_{\\mathrm{in}}/R_{\\mathrm{out}}$ and period ratio $P_{\\mathrm{in}}/P_{\\mathrm{out}}$ carry the quantitative signal, tested with Anderson–Darling comparisons of distributions and Fisher exact tests on small configuration counts. The de-biased sample is defined by hand-chosen cuts of $R > 2\\,R_\\oplus$ and $P < 50$ days, anchored to Kepler's roughly 90 percent detection-completeness region. Synthetic samples—random shuffles within multi-planet systems and random pairings of single-planet hosts matched in stellar mass and Hill-stability—serve as null models for what chance or bias alone would produce.","core_discovery":"The central claim is that inner planets in Kepler multi-planet systems are systematically smaller than outer planets, and that this ordering survives the authors' de-biasing cut (radii above two Earth radii and periods below fifty days), so it is not merely a product of transit detection biases. For two-planet systems, the configuration with a smaller inner planet outnumbers the reverse by 273 to 93 in the full sample and 153 to 68 after de-biasing. In three-planet systems the effect is strongest for the innermost pair and weakens for the outermost pair, which instead shows similar-sized planets in the 'peas in a pod' style. The observed radius-ratio distributions differ from a synthetic homogeneous sample and from synthetic pairs built from single-planet systems, supporting an intrinsic origin. The paper also reports a metallicity dependence of the inner-to-outer radius-ratio distribution, and no significant difference between resonant and non-resonant pairs, a null result it argues is in tension with simple resonant-capture expectations.","pith_inferences":["If the ordering is genuinely physical, mass ordering from transit-timing or radial-velocity measurements could be tested as a cleaner surrogate, because radii can be inflated by atmospheres and blur the formation signal.","A natural extension is to split resonant pairs by resonance order (2:1 versus 3:2) or by direct libration confirmation; the paper's null result may conceal a signal specific to certain resonances.","One discriminating prediction: if photoevaporation is a main driver, the inner-smaller trend should weaken or strengthen with stellar age and irradiation in a way the paper does not test, since envelope loss accumulates over time.","The synthetic single-planet comparison implies that blindly pairing single-planet hosts yields the opposite ordering; explaining why single-planet hosts differ from multi-planet hosts may itself be a clue about divergent formation pathways."],"forward_implications":["The inner-smaller ordering becomes a testable constraint for planet formation models, which the paper notes have not yet produced predictions for ordering.","Because the ordering's strength varies with pair location and multiplicity, formation and evolution codes will need to reproduce not just individual planet sizes but their relative arrangement within a system.","The metallicity-dependent radius-ratio distribution ties final system architecture to protoplanetary disk composition, giving observers a way to connect initial conditions to outcomes.","The null result for resonant pairs implies that if resonance capture shaped these systems, later destabilization must be common enough to erase any expected size-ratio signature.","Transit-selected multi-planet samples are biased toward nearly coplanar, dynamically quiet systems, so the trend may describe that subset rather than the full planetary population."],"supporting_citations":[{"why":"Provides the catalog of planetary radii, periods, stellar temperatures, and metallicities from which all samples are drawn.","marker":"NASA Exoplanet Archive (2022)"},{"why":"Supplies the Kepler 90 per cent detection-completeness threshold used to justify the de-biased sample region.","marker":"Petigura et al. (2013)"},{"why":"Earlier finding that the larger planet in a pair usually has the longer period, the direct observational precedent for size ordering.","marker":"Ciardi et al. (2013)"},{"why":"Proposed the 'peas in a pod' pattern of similar-sized, regularly spaced planets that this study's ratio analysis complements and partially contrasts.","marker":"Weiss et al. (2018)"},{"why":"Provided the transit-timing-variation mass measurements used in previous same-system mass studies.","marker":"Hadden & Lithwick (2017)"},{"why":"Found planets in the same system tend to have similar masses, motivating the question of how sizes and masses are arranged.","marker":"Millholland et al. (2017)"},{"why":"Reported that planet sizes in two-planet systems are ordered non-randomly, supporting the intrinsic-ordering interpretation.","marker":"Kipping (2017)"},{"why":"Justifies the 0.02 proximity tolerance used to define near-resonant pairs in the resonance analysis.","marker":"Steffen & Hwang (2015)"},{"why":"Proposed photoevaporation as the mechanism that shrinks close-in planets, giving a physical route to smaller inner planets.","marker":"Owen & Wu (2013)"}],"fun_headline_variants":["Inner planets are small in Kepler systems—no bias artifact","Even after de-biasing, inner Kepler planets stay small","Metallicity links to planet size order in Kepler systems","Resonant pairs don't perturb the inner-small pattern","Small inner planets: a robust signature of formation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The trend's robustness rests on the de-biased sample, a hand-picked window of radii above two Earth radii and periods under fifty days that the paper concedes may still have minor selection effects; if that window preferentially hides small inner planets or large outer planets, the apparent inner-smaller ordering could be manufactured by the cut itself.","fun_headline_variants_meta":{"raw":{"variants":["Inner planets are small in Kepler systems—no bias artifact","Even after de-biasing, inner Kepler planets stay small","Metallicity links to planet size order in Kepler systems","Resonant pairs don't perturb the inner-small pattern","Small inner planets: a robust signature of formation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000441,"raw_usage":{"total_tokens":2269,"prompt_tokens":1011,"completion_tokens":1258,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":627,"completion_tokens_details":{"reasoning_tokens":1179}},"tokens_in":627,"tokens_out":1258,"duration_ms":12581,"temperature":1.0,"reasoning_tokens":1179,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T17:15:16.169748+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Forward-model Kepler's detection efficiency: inject synthetic multi-planet populations with no intrinsic radial size gradient, run them through the selection function, and check whether the observed '12' versus '21' counts and the $R_{\\mathrm{in}}/R_{\\mathrm{out}}$ distributions can be reproduced by bias alone. If they can, the claim that the ordering is intrinsic collapses; if the inner-smaller trend disappears when the same ordering analysis is applied to a sample with an independent selection function, the de-biasing is insufficient to establish it.","supporting_citations":[{"cited_title":"2017, Monthly Notices of the Royal Astronomical Society, 473, 784, 10.1093/mnras/stx2383","cited_arxiv_id":null,"evidence_quote":"Reported that planet sizes in two-planet systems are ordered non-randomly, supporting the intrinsic-ordering interpretation."}],"review_version":1}