{"id":"c8ae6659-aba3-4ae5-ad92-250ce402d66a","arxiv_id":"2509.09762","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A quantitative comparison of the Bern planet formation simulation with the HARPS/Coralie survey finds about 70% too many planets, a too-deep mass desert, too-round orbits, and planets that end up too close to their stars.","lead":"This paper compares a computer simulation of how planets form (the Bern model) with a real survey of planets found around nearby Sun-like stars. It finds the simulation produces too many planets, puts them too close to their stars, and makes their orbits too round, suggesting missing physical effects in the model.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Eccentricity comparison may be dominated by the RV fitting bias the paper itself cites: the claimed factor-of-two deficit (0.07 vs 0.15) could shrink or reverse under an observer-twin pipeline test.","rationale":"The reader's weakest assumption focuses on the M11 completeness map as the load-bearing premise for all discrepancy percentages. I agree that the map is important and somewhat fragile, especially because it is derived from the same unpublished survey, but the paper is transparent about its limitations and the map is the standard tool for this kind of comparison. The more scientifically load-bearing soft spot is the eccentricity comparison, where the paper's own cited literature undermines the reliability of the observed eccentricities. Since the central claims are numerous and largely independent, the eccentricity concern does not invalidate the paper's overall conclusion that the nominal model mismatches the survey, nor the qualitative missing-physics suggestions. It does, however, affect the strength of one of the four headline numbers and one of the key conclusions about missing eccentricity excitation. The paper's appropriate hedging and the breadth of independent discrepancies keep the CONDITIONAL verdict appropriate. I therefore recommend UNCHANGED, but with a concrete observer-twin test to settle the eccentricity comparison before that specific claim is relied upon.","tokens_in":42731,"tokens_out":1417,"duration_ms":19763,"concrete_test":"Generate mock RV time series for the biased synthetic planets assuming the HARPS/Coralie sampling cadence and noise per star, run the same Keplerian fitting procedure used for the real survey, and compare the recovered eccentricities to both the true synthetic values and the observed catalog values. Specifically, recompute the median eccentricity of the 'detected' mock planets using the fitted e values; if the fitted median rises above about 0.12, the claimed factor-of-two discrepancy is largely a pipeline artifact.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's own Sect. 3.7 concedes the classic RV eccentricity overestimation bias: eccentricities are easily overestimated for low-S/N orbits (Lucy & Sweeney 1971; Hara et al. 2019), and Zakamska et al. (2011) found about 38% of RV planets have e<0.05 versus 17% in standard analyses. Yet the central eccentricity claim compares the synthetic median e=0.07, which has no observational error, against the catalog eccentricities of HARPS/Coralie planets that were fitted from noisy RV time series. A fair comparison requires applying the same detection and fitting pipeline to synthetic signals, not just applying the detection map. This matters because the 'too dynamically cold' conclusion is one of the four headline quantitative discrepancies (Abstract, Sect. 3.7) and is used in Sect. 5 as evidence for missing eccentricity-excitation physics. The paper itself notes that the Kepler multi-planet eccentricities are about 0.03-0.05, which is close to the synthetic value, so a re-fit of the RV sample could plausibly bring the observed median down by a factor of two, making the discrepancy much weaker or even inverted. This is a load-bearing concern about the correctness of a headline quantitative claim, not just about precision.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper compares synthetic planet populations from the Bern model (NG76 and NG76longshot) against the HARPS/Coralie RV survey of Mayor et al. (2011), updated to 2015. A synthetic detection bias based on the M11 completeness map is applied to the synthetic populations, and 1000 mock observations of 822-star samples are compared to the observed sample via KS tests. The nominal population reproduces several qualitative features: the bimodal mass function, close-in sub-Neptunes versus distant giants, mean multiplicity ~1.6, period-ratio pile-ups, and broad metallicity and eccentricity trends. The headline discrepancies are a ~70% overproduction of detectable planets, a planetary desert too deep by ~60%, a ~40% relative excess of giants, median eccentricity 0.07 versus 0.15, and planets too close to their stars. A tuned population NG192 (2 km planetesimals, reduced migration, modified gas accretion) nearly matches the mass function (KS distance 1.35 vs 1.36) but still fails on orbital distances and has even lower eccentricities. The paper concludes that missing physics, such as wider formation orbits, eccentricity excitation, and slower gas accretion, is needed.","tokens_in":1653,"tokens_out":1747,"duration_ms":82851,"significance":"The nominal population was not tuned to the HARPS/Coralie survey, so this is a valuable, independent stress test of the Bern model. The Monte Carlo mock-observation procedure is clearly described, and the internally consistent discrepancy list provides a concrete benchmark for the community. The tuned-population experiment illustrates parameter degeneracies and correctly identifies that no single parameter change fixes both the mass and period distributions. The main weaknesses are that all quantitative claims inherit the assumptions of a single mean completeness map, and that the eccentricity comparison uses catalog eccentricities that are affected by RV fitting bias, as the paper itself partly notes.","major_comments":[{"comment":"The factor-of-two eccentricity deficit is a headline discrepancy, but the comparison is not apples-to-apples. The synthetic eccentricities are noiseless model values, while the HARPS/Coralie eccentricities come from noisy, sparsely sampled RV fits. The paper cites Zakamska et al. (2011), who found that about 38 percent of RV planets have e<0.05 versus 17 percent in standard catalogs. Since the synthetic median is 0.07, an end-to-end test that injects synthetic RV signals and recovers eccentricities with the same pipeline could substantially reduce or even reverse the claimed deficit. This matters because the dynamically cold conclusion is used as evidence for missing eccentricity-excitation physics. Please provide such a test or re-derive the observed eccentricity distribution with an upper-limit-aware method.","section":"Sec. 3.7.1, Fig. 10; Abstract"},{"comment":"All quantitative discrepancy percentages (70 percent overproduction, 60 percent desert depth, 40 percent giant excess, and the eccentricity comparison) are computed under one mean detection-completeness map that, as stated in Sec. 2.3, ignores eccentricity and system architecture and is averaged over stars. The map is also from the same unpublished M11 analysis used as the observed sample. The quantitative claims would be much more robust with a sensitivity analysis, e.g., applying an eccentricity-aware or star-by-star completeness correction and checking how much the percentages change. Without this, the direction and magnitude of some discrepancies could plausibly change.","section":"Sec. 2.3, Fig. 3; Sec. 3"},{"comment":"The claim that NG192 nearly matches the observed mass function is based on a KS distance of 1.35 at the 95 percent level versus a threshold of 1.36. Because NG192's parameters were selected using the HARPS/Coralie mass distribution, this near-threshold value is a fitting residual, not an independent validation. The paper's main conclusion that mass and period cannot be simultaneously matched is still valid, but the near-threshold wording risks overinterpretation. A holdout split or a clear statement that this is a posterior fit would be more appropriate.","section":"Sec. 4, Fig. 13"}],"minor_comments":[{"comment":"The caption says the median is indicated by the vertical dashed red line and then refers to the synthetic value also as the vertical dashed red line. This is ambiguous; please clarify which line is which.","section":"Fig. C.1 caption"},{"comment":"The phrase 'we detect 290+30-28 planets' could be misread as an actual detection. Consider writing 'the mock observations yield...' or 'the biased synthetic sample contains...'.","section":"Sec. 3.1"},{"comment":"The desert-depth metric uses the 20-200 M_earth range chosen from the synthetic cumulative distribution. Since the observed desert may be located elsewhere, please report the sensitivity of the 57-60 percent number to the adopted mass boundaries.","section":"Sec. 3.2"},{"comment":"Given that the quantitative claims rest on the M11 completeness map, please make that map available in machine-readable form, since the M11 survey paper remains unpublished.","section":"Sec. 2.3"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is transparent and the nominal-population comparison is a fair test, but the two load-bearing concerns above are real and require additional analysis. I would not reject: the eccentricity and completeness-map issues are fixable with additional work, and the central message about missing physics is likely to survive in some form. However, the near-threshold KS fit for NG192 should not be presented as validation, and the abstract's unqualified discrepancy percentages should be softened unless the robustness tests are added. The unpublished M11 map is a reproducibility issue that the editor may want to raise with the authors."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The punchline first: this is the first quantitative comparison of the Generation III Bern model against the full HARPS/Coralie sample that actually does the work — 1000 mock observations, KS tests, 95% confidence intervals — and it produces specific, checkable discrepancy numbers instead of \"the model looks roughly similar.\" The nominal NG76 population wasn't tuned to HARPS, so the main comparison is a genuine test, and the paper is properly honest about its own caveats.\n\nWhat it does well: the headline discrepancy list (70% overproduction, ~60% desert depth, ~40% giant excess, median e 0.07 vs 0.15) is internally consistent with the figures. The optimized NG192 population is explicitly labeled as fitted, not predictive, and its near-miss KS result (1.35 vs a 1.36 rejection threshold) is reported squarely. The parameter study in Appendix B is a genuinely instructive example of how coupled the model parameters are — you can fix the mass function but the orbital distances stay wrong.\n\nThe soft spots, in proportion. The biggest is the M11 completeness map: a mean detection probability averaged over the sample stars, ignoring eccentricity and system architecture, and taken from the same unpublished survey that provides the observed sample. Every quantitative percentage inherits uncertainty from that map, and the paper doesn't quantify how sensitive the discrepancy list is to it. Second, the stress-test on eccentricity lands. The paper compares synthetic eccentricities with no observational error against catalog values fitted from noisy RV time series. It cites Zakamska et al.'s finding that proper reanalysis puts 38% of RV planets at e<0.05 versus 17% in standard fits, and it notes Kepler multi-planet systems sit at e≈0.03–0.05 — right at the synthetic value. Yet the factor-of-two deficit is still one of the four headline claims, without an observer-twin re-fit of synthetic signals to show the bias can't explain it. That's the softest number in the paper, and the missing eccentricity-excitation conclusion leans on it. The fixed 1 M⊙ stellar mass versus the sample's 0.91 M⊙ mean is acknowledged and is probably minor.\n\nWho gets value: anyone doing population synthesis, RV demographics, or model-observation comparisons. It's a good template for how to run this kind of benchmark, and the discrepancy list is a useful target list for modelers. It deserves a serious referee. My recommendation: send it out, and ask the referee to push on the observer-twin eccentricity test and on sensitivity of the headline percentages to the completeness map. The framework is reproducible and the core result — the nominal model is informative but wrong in specific, tractable ways — will survive that scrutiny.","headline":"The first genuinely quantitative benchmark of the Gen III Bern model against HARPS/Coralie, with honest Monte Carlo machinery and a specific discrepancy list; the eccentricity deficit is the softest headline number because it compares against RV fits the paper itself shows are biased.","tokens_in":43592,"tokens_out":5710,"would_cite":true,"duration_ms":57074,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A planet-formation model not tuned to any survey overproduces detectable planets by ~70% and yields orbits too close and too circular, the paper finds.","keywords":["planet formation","population synthesis","core accretion","radial velocity survey","HARPS/Coralie","planetary mass function","eccentricity distribution","planetary migration"],"falsifier":"Recompute the comparison star-by-star: inject each synthetic planet into the actual HARPS/Coralie noise and detection pipeline rather than using the averaged circular-orbit completeness map. If the corrected count drops from 290 toward 169 and the eccentricity distributions agree, the missing-physics conclusion weakens; if the 70% excess and the median-eccentricity gap persist, the conclusion stands.","tokens_in":1353,"feed_emoji":"🪐","tokens_out":1633,"duration_ms":79957,"temperature":0.7,"pith_summary":"This paper asks whether a modern planet-formation simulation can reproduce, as a statistical population, the planets found by the HARPS/Coralie radial-velocity survey. The authors run their nominal Generation III Bern model—not tuned to any particular survey—apply the survey's detection completeness to the synthetic planets, and compare 1000 mock observations with the actual 169 detected planets. The central finding is that the model captures several qualitative features (two planet groups, bimodal mass function, multiplicity near 1.6, some correlations) but fails quantitatively: too many planets, a too-deep desert, too many giants, too-low eccentricities, a too-weak metallicity effect, and planets too close to their stars. The paper then shows that modest parameter changes (larger planetesimals, slightly weaker migration, slower disc-limited gas accretion) nearly fix the mass function, but the orbital distances and dynamical excitation remain wrong, implying missing physical processes rather than simple parameter tuning.","feed_headline":"Untuned Bern model yields 70% too many planets","feed_subtitle":"Detection bias alone can't erase the gap: fixing planet masses leaves orbits too close and eccentricities too low.","key_machinery":"The load-bearing object is the Generation III Bern model: a global population-synthesis code in which protoplanets grow by core accretion from planetesimals, accrete gas, migrate under type I and type II disc migration, and interact dynamically through N-body physics, all starting from observationally motivated disc initial conditions. The comparison mechanism is the survey's mean completeness map—a detection-probability grid in minimum mass and period built by injecting circular-orbit signals into the actual HARPS/Coralie data—applied uniformly to every synthetic planet, with 1000 Monte-Carlo mock observations used to build confidence intervals and run KS tests on mass, period, mass-period,","core_discovery":"The paper's claim, stated on its own terms, is that the nominal generation-III Bern model population, once passed through the HARPS/Coralie detection bias, is statistically inconsistent with the observed sample: it predicts 290 planets where 169 are found (~70% excess), a planetary desert between about 20 and 200 Earth masses that is ~60% too empty, a ~40% relative excess of giant planets, a median eccentricity of 0.07 versus an observed 0.15, a too-weak dependence of planet occurrence on stellar metallicity, and planets systematically closer to their stars. Extending the N-body integration to 100 Myr does not cure the dynamical discrepancies. The authors then construct an adjusted populatio","pith_inferences":["A star-by-star completeness calculation that also accounts for eccentricity and per-star noise could shrink or shift the claimed 70% excess and factor-two eccentricity gap; the averaged circular-orbit map smooths over exactly the regions where the model's overabundance might be concentrated.","If wide initial orbits are indeed the missing ingredient, then a synthesis with a larger initial disc radius or with pebble accretion should populate the 300–3000 day giant-planet region without also filling the desert; that is a directly testable prediction.","The model's failure to produce hot Jupiters through disc migration while finding ~20 planets destined to hit the star on eccentric orbits suggests that adding tidal circularisation could close the hot-Jupiter gap without invoking new initial conditions.","The same biasing-the-synthesis approach could be applied to transit surveys, testing whether the missing-physics conclusion is detection-technique-specific or fundamental to the formation model."],"forward_implications":["The too-deep desert and the over-massive giants both point to the same model element: disc-limited gas accretion rates that are too high in the detached phase.","Because the period-ratio distribution changes little when the N-body integration is extended to 100 Myr, late dynamical instabilities are not what breaks 2:1 resonances in the RV-accessible regime.","Migration strength is tightly constrained: reducing it enough to push planets outward would overproduce giant planets relative to sub-Neptunes.","The optimised population gets total planet count and mass distribution nearly right but leaves distances and eccentricities wrong, so the deficit is in missing processes, not just parameter values.","The metallicity correlation is reproduced in shape but too weak, and the model cannot form enough metal-poor, close-packed sub-Neptune systems of the kind observed."],"fun_headline_variants":["Bern model births 70% too many planets","Planet model fails reality check: 70% excess","Simulated planets too close, too circular, too many","HARPS data expose model's desert too deep","Planet formation model overproduces by 70%"],"cache_read_input_tokens":44928,"weakest_assumption_plain":"The load-bearing premise is that the survey's mean completeness map—an average detection probability derived from the same unpublished survey and applied uniformly to all 822 stars and all synthetic planets—accurately represents what HARPS/Coralie would detect, including for eccentric orbits and varied system architectures.","fun_headline_variants_meta":{"raw":{"variants":["Bern model births 70% too many planets","Planet model fails reality check: 70% excess","Simulated planets too close, too circular, too many","HARPS data expose model's desert too deep","Planet formation model overproduces by 70%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000164,"raw_usage":{"total_tokens":1153,"prompt_tokens":886,"completion_tokens":267,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":630,"completion_tokens_details":{"reasoning_tokens":189}},"tokens_in":630,"tokens_out":267,"duration_ms":3348,"temperature":1.0,"reasoning_tokens":189,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T18:45:16.683313+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the comparison star-by-star: inject each synthetic planet into the actual HARPS/Coralie noise and detection pipeline rather than using the averaged circular-orbit completeness map. If the corrected count drops from 290 toward 169 and the eccentricity distributions agree, the missing-physics conclusion weakens; if the 70% excess and the median-eccentricity gap persist, the conclusion stands.","supporting_citations":[],"review_version":1}