{"id":"e4dc2586-de34-4d2a-a186-d211e8a85ae5","arxiv_id":"2504.18643","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":9,"one_line_summary":"Five hydrodynamics codes and two radiative transfer codes produce consistent planet-driven kinematic signatures in synthetic observations, allowing reliable planet location retrieval.","lead":"This paper checks whether different computer codes used to simulate protoplanetary disks give the same answers. It finds they agree well enough that any tested combination can be used to locate a hidden planet from its gas motions. A smart generalist should care because the result strengthens confidence in how astronomers interpret ALMA observations of planet-forming disks.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The abstract's universal claim that any tested code combination can be reliably used rests on a single benchmark model; the large gap-depth discrepancy in Phantom shows regime-dependent sensitivity that the one-case retrieval test cannot rule out.","rationale":"The reader's weakest assumption correctly identifies the single-model generalizability as the primary limitation of the paper's broad conclusion. The central claim is universal in form, but the data set is one point in a multidimensional parameter space. The Phantom result is a concrete warning: a major physical observable (gap depth and the associated azimuthal velocity perturbation) differs by a factor of several across codes, and the paper's defense that retrieval is unaffected relies on a single case. This is not an internal inconsistency in the tested configuration, but it is a load-bearing gap in the argument because the abstract's summary is precisely the universal generalization. The concrete test of a second benchmark configuration (e.g., Saturn-mass planet or finite cooling) would directly probe whether the consistency is robust to regime changes. If it is not, the conclusion should be qualified. The existing verdict of CONDITIONAL is appropriate; no change in verdict is needed, but the condition should explicitly include a requirement for additional benchmark points or a suitably qualified abstract.","tokens_in":15730,"tokens_out":8140,"duration_ms":84507,"concrete_test":"Run a second benchmark configuration with a lower-mass planet (e.g., Mp = 3e-4 M*, a Saturn-like planet) or with a simple beta-cooling prescription, keeping all other numerical and disk parameters identical, and repeat the full nine-combination matrix: hydrodynamics, radiative transfer, and DISCMINER retrieval. If the scatter in retrieved radial/azimuthal locations exceeds the claimed few percent, or the peak residual velocity scatter exceeds 20%, the abstract's summary must be restricted to the tested regime. Additionally, report the DISCMINER detection significance for all nine combinations in the current paper to confirm that every retrieval is a robust detection above noise.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is a universal statement: 'any combination of the tested hydrodynamics and radiative transfer codes can be used to reliably model and interpret planet-driven kinematic perturbations.' The evidence, however, is a single base model: one Jupiter-mass planet at 100 au, isothermal equation of state, kinematic viscosity nu=1e-5, no magnetic fields, and no migration. The paper itself admits this in Section 5: 'we only considered a single base hydrodynamics model.' The Phantom simulation already reveals a major observable discrepancy: surface density suppression in the gap is only about 10% versus about 70% in grid-based codes, and the azimuthal velocity perturbation across the gap is about 1% versus about 6%. The authors argue this does not affect planet-location retrieval because the spiral-driven perturbations that produce velocity kinks are similar, but this is supported by only a single retrieval case. If other regimes place more weight on the gap-induced velocity perturbation (e.g., lower planet mass, where the spiral perturbation is weaker relative to the gap velocity change, or finite cooling that modifies the spiral temperature structure), the code-to-code consistency in retrieved locations could degrade. Without a second benchmark point, the summary conclusion overreaches the tested domain.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper benchmarks five hydrodynamics codes (FARGO3D, Idefix, Athena++, PLUTO, and Phantom) and two radiative transfer codes (mcfost and RADMC-3D) on a single planet-disk interaction model: a Jupiter-mass planet at 100 au, isothermal equation of state, kinematic viscosity nu=1e-5, no migration, and no magnetic fields. The authors generate synthetic 12CO J=3-2 cubes from nine code combinations, retrieve the planet's radial and azimuthal location with DISCMINER, and compare the results. They report strong consistency: disk temperatures agree within about 3% between mcfost and RADMC-3D, brightness temperatures along the iso-velocity contour agree within ±1.5 K, retrieved planet locations show only a few percent scatter around the true location, and peak residual velocities vary by roughly 20% among models. They conclude that any combination of the tested hydrodynamics and radiative transfer codes can be used to reliably model and interpret planet-driven kinematic perturbations, while also noting in Section 5 that only a single base hydrodynamics model was considered.","tokens_in":15928,"tokens_out":5365,"duration_ms":55619,"significance":"The benchmark is valuable to the protoplanetary disk community because it is the first systematic, controlled comparison of the specific codes used for forward modeling of planet-driven kinematic structures in exoALMA-era observations. The study is carefully set up: code versions are pinned, the initial conditions are fully specified, and the planet is injected with known mass and orbital parameters, so the retrieval comparison has a well-defined ground truth. The quantitative agreement among grid-based codes and between mcfost and RADMC-3D is a useful reference result, and the paper is commendably explicit about the Phantom mass loss and shallow gap. If the conclusions are appropriately qualified to the tested configuration, this will be a useful citable benchmark for the community. The main weakness is not in the measurements themselves but in the breadth of the summary claim, which goes beyond what a single benchmark model and nine of ten possible code combinations can establish.","major_comments":[{"comment":"The abstract and Section 5 state that 'any combination of the tested hydrodynamics and radiative transfer codes' can be used reliably, but the Phantom + RADMC-3D combination was not tested; Section 3 explicitly lists it as excluded because no RADMC-3D module reads Phantom outputs. Since both Phantom and RADMC-3D are among the 'tested' codes, the blanket statement is not literally supported by the nine tested combinations. Please rephrase to 'the nine combinations tested here' or provide a justification for why the missing combination cannot affect the conclusion.","section":"Abstract and Section 3"},{"comment":"The universal conclusion overreaches the evidence from a single base model. The Phantom simulation shows only about 10% surface density suppression in the gap versus about 70% in the grid-based codes, and the azimuthal velocity perturbation across the gap is about 1% versus about 6% (Figure 5). The DISCMINER retrieval exercises the spiral-arm velocity kinks, which are consistent, but it does not test the gap-induced azimuthal velocity perturbation as an observable. The paper's own final paragraph acknowledges this limitation. The abstract and summary should therefore be narrowed, for example to 'any of the tested code combinations can be used to model and interpret planet-driven spiral kinematic perturbations in this configuration,' and should state explicitly that consistency for gap-depth-based diagnostics is not established by this benchmark.","section":"Section 5 and Figure 5"},{"comment":"The statement that peak residual velocities vary by 'about 20%' appears inconsistent with the reported minimum, maximum, and mean values (0.054 km/s, 0.088 km/s, and 0.074 km/s). The minimum is about 27% below the mean, and the full range is about 46% of the mean. Please report a standard scatter metric (standard deviation or max absolute deviation from the mean) and adjust the abstract and Section 4 wording accordingly, or explain what '20%' refers to.","section":"Section 4 and Abstract"}],"minor_comments":[{"comment":"There is a typo in 'hydrodynmics codes'; it should read 'hydrodynamics codes.'","section":"Section 3.2.1"},{"comment":"The notation for carbon monoxide should be consistent: use $^{12}$CO with a superscript in all instances rather than the plain text '12CO'.","section":"Throughout"},{"comment":"In the sentence about the stellar photon count, '10 9 photons' is missing a superscript; it should read $10^9$ photons.","section":"Section 3.1"},{"comment":"The radial retrieval is reported as systematically biased to a mean of 109 au versus the true 100 au. This is within one beam, but since the bias appears in all models, a brief discussion of its origin beyond the cited prior work would strengthen the paper.","section":"Section 4"},{"comment":"The caption refers to a 'yellow dashed square' marking the velocity kink, but in the figure as rendered the square may be difficult to locate; please ensure the marker is clearly visible in the published version.","section":"Figure 7"}],"recommendation":"major_revision","confidential_remarks":"The science in this benchmark appears sound, and the quantitative comparisons are a useful community resource. The main issue is scope: the abstract's 'any combination' claim is too strong given that one of ten possible hydrodynamics/radiative-transfer combinations was not tested and only a single physical setup was simulated. These are fixable with rewording and additional caveats, so I recommend major revision rather than rejection. The paper's fit to the journal is good; the required changes are primarily in the presentation of the conclusions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, the useful news: this is the benchmark the kinematic planet-detection field actually needed. Prior code comparisons stopped at gap structure and torques; this one runs five hydrodynamics codes and two radiative transfer codes all the way through synthetic 12CO cubes to DISCMINER planet-location retrieval, with a known injected planet as ground truth. For the tested configuration the agreement is genuine: temperatures within ~3%, brightness temperatures within ~1.5 K, retrieved locations within a few percent of each other and within about a beam of the truth. The paper reports the Phantom outlier honestly—gap surface density suppressed ~10% versus ~70% in grid codes, azimuthal velocity perturbation ~1% versus ~6%—and it pins code versions and describes the setup well enough to be repeatable.\n\nSoft spots, in proportion. The abstract's closing claim, that any combination of the tested codes can be used reliably, is broader than the evidence: one base model, one planet mass, isothermal EOS, fixed viscosity, no migration, no cooling, no magnetic fields. The authors do acknowledge this in the final section, but the summary sentence still oversells. And since Phantom+RADMC-3D was never run, 'any combination' is not literally true on the paper's own terms; the missing configuration files are a small reproducibility gap for a benchmark. Minor issues, easy to fix.\n\nThe bigger soft spot is the Phantom gap-depth discrepancy. The physical argument that it does not affect retrieval—velocity kinks trace spiral perturbations, and those agree across codes—carries weight, but it rests on a single retrieval case. At lower planet masses, where the spiral perturbation is weaker relative to the gap-induced azimuthal velocity change, the code-to-code retrieval scatter could grow. A second benchmark point (a Neptune-mass planet, or a stratified-temperature run) would turn that argument from plausible into demonstrated. As written, the conclusion is solid for Jupiter-mass, isothermal, non-migrating disks, and an open question elsewhere.\n\nCitation pattern is fine. The de Val-Borro lineage is credited; the self-citation is mostly code papers and exoALMA companion papers, which is appropriate for a benchmark. No circularity problem: the injected planet is ground truth, not a fitted result.\n\nThis paper is for anyone forward-modeling kinematic planet detections in ALMA data, and it deserves serious refereeing. My recommendation: send it to peer review, and ask the authors to temper the abstract (or add the second benchmark point) as the main revision. I would not desk-reject this.","headline":"A genuinely useful benchmark that does what it claims for the tested configuration, but the abstract's 'any combination' conclusion outruns the single-model evidence.","tokens_in":16643,"tokens_out":6022,"would_cite":true,"duration_ms":52509,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Forward modeling of planet-driven disk kinematics is robust to the choice of code, a benchmark of five hydrodynamics and two radiative transfer codes shows.","keywords":["protoplanetary disks","planet-disk interactions","hydrodynamical simulations","radiative transfer simulations","kinematic perturbations","velocity kinks","forward modeling","12CO channel maps"],"falsifier":"Run the same five hydrodynamics and two radiative transfer codes for a substantially different regime, such as a 10 Earth-mass planet or a disk with vertical temperature stratification and finite cooling; if retrieved planet positions scatter by more than a beam or brightness temperatures differ by more than the 1.5 K noise level across code combinations, the claim that any tested combination can be used reliably would be refuted.","tokens_in":15476,"feed_emoji":"🪐","tokens_out":4143,"duration_ms":39571,"temperature":0.7,"pith_summary":"This paper tries to establish that the forward modeling pipeline used to interpret planet-driven kinematic perturbations in protoplanetary disks is reliable regardless of which community code is used. It benchmarks four grid-based hydrodynamics codes and one smoothed particle hydrodynamics code, two radiative transfer codes, and nine code combinations, then retrieves the embedded planet's location from synthetic 12CO cubes. The result is that retrieved radial and azimuthal planet positions agree with the true location within a few percent scatter, peak residual velocities vary by less than 20 percent, and brightness temperatures differ by at most about 1.5 K. If the conclusion holds, researchers can trust planet-location results obtained with any of these code combinations, which matters because exoplanet detections in disks increasingly rely on such kinematic signatures.","feed_headline":"Benchmark: any disk-code combo finds the planet at the same spot","feed_subtitle":"Nine hydrodynamics-plus-radiative-transfer pipelines retrieve a Jupiter's position within a few percent spread.","key_machinery":"The central object is the planet-driven spiral perturbation in gas density and velocity, which produces a localized velocity kink in 12CO channel maps. The benchmark chain consists of hydrodynamics codes that evolve the disk, radiative transfer codes that convert the density and velocity fields into temperatures and synthetic line cubes, and DISCMINER's folded velocity residual maps, which locate the peak residual velocity associated with the kink. The argument works because the strength of the kink is set by the spiral-arm velocity perturbations, which agree across all five hydrodynamics codes, rather than by the depth of the planet-opened gap, which does not.","core_discovery":"The central claim is that any combination of the tested hydrodynamics and radiative transfer codes can be used reliably to model and interpret planet-driven kinematic perturbations in protoplanetary disks. Across the nine hydrodynamics-plus-radiative-transfer combinations, the disk temperature agrees to within about 3 percent or better between mcfost and RADMC-3D models everywhere in the domain, and synthetic 12CO channel maps show brightness temperature differences within 1.5 K, which is around the typical noise level of exoALMA-quality molecular line images. DISCMINER retrieves the planet's radial location at 105 to 116 au for a true value of 100 au and azimuthal location within about 4 degrees of the true value of -45 degrees, with peak residual velocities between 0.054 and 0.088 km/s. The only notable disagreement is that the Phantom simulation opens a much shallower gap than the grid-based codes, but this does not affect the kinematic retrieval, because the spiral-arm density and velocity perturbations that produce the velocity kinks remain consistent.","pith_inferences":["The benchmark covers only one model—a Jupiter-mass planet at 100 au in an isothermal disk with no magnetic fields, cooling, or migration—so the blanket conclusion that any combination works is an extrapolation until other regimes are tested.","If lower-mass planets or vertically stratified, cooling, or magnetized disks produce larger code-to-code differences, the kinematic retrieval may still be robust, but the exact scatter should be re-measured rather than assumed.","The Phantom simulation's shallow gap, caused partly by mass loss through the inner boundary, suggests that gap-based planet mass estimates from SPH runs could be biased even though kinematic location retrieval is not.","Comparing the retrieved peak location with both spiral arms, rather than the global peak that favors the outer arm, might reduce the systematic radial offset and is a testable extension of the DISCMINER workflow."],"forward_implications":["Planet location retrieval from kinematic signatures in exoALMA-quality data can proceed with any tested code combination, without needing to re-calibrate the result to a particular hydrodynamics or radiative transfer code.","Brightness temperature differences between code combinations stay below the typical 1.5 K noise, so synthetic cubes from different codes would be observationally indistinguishable at current sensitivity.","The shallow gap produced by the SPH code does not translate into a biased planet location, so kinematic planet-finding appears decoupled from gap-depth uncertainties.","Simulations and radiative transfer calculations from different groups can be combined and compared directly in future forward modeling surveys.","The known radial bias toward larger radii in DISCMINER retrieval persists across all code combinations, indicating it is a property of the retrieval method rather than of any particular code."],"supporting_citations":[{"why":"Provides the prior planet-disk code comparison benchmark and the wave-killing damping method used for the radial boundary conditions.","marker":"de Val-Borro et al. 2006"},{"why":"Supplies the FARGO3D code, one of the four grid-based hydrodynamics codes benchmarked.","marker":"Benítez-Llambay & Masset 2016"},{"why":"Supplies the Phantom SPH code and the viscosity implementation used in the smoothed particle hydrodynamics run.","marker":"Price et al. 2018"},{"why":"Establishes the DISCMINER retrieval method and previously noted the radial bias that appears again in the benchmark results.","marker":"Izquierdo et al. 2021"},{"why":"Defines the folded velocity residual map technique used to locate the planet in the synthetic cubes.","marker":"Izquierdo et al. 2023"},{"why":"Provides the comparison between isothermal and radiative-transfer-updated simulations that justifies the simplified temperature treatment for velocity perturbations.","marker":"Pinte et al. 2019"},{"why":"Supports the grid resolution choice of about 12 cells per scale height as sufficient for numerical convergence.","marker":"Chen & Dong 2024"},{"why":"Sets the fiducial exoALMA noise level and beam size used to degrade and interpret the synthetic 12CO cubes.","marker":"Teague et al. 2025"},{"why":"Provides the inferred temperature structure that motivates the disk aspect ratio used in the benchmark model.","marker":"Galloway-Sprietsma et al. 2025"}],"fun_headline_variants":["Any disk-code combo pinpoints planet","Disk-code benchmark: same planet spot","Nine code combos, one planet position","All tested disk codes agree on planet","Planet retrieval robust across disk codes"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The conclusion rests on a single benchmark model being representative: one Jupiter-mass planet on a fixed circular orbit at 100 au in an isothermal disk with no magnetic fields, no cooling, and no migration, a limitation the paper itself acknowledges.","fun_headline_variants_meta":{"raw":{"variants":["Any disk-code combo pinpoints planet","Disk-code benchmark: same planet spot","Nine code combos, one planet position","All tested disk codes agree on planet","Planet retrieval robust across disk codes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000917,"raw_usage":{"total_tokens":3995,"prompt_tokens":1063,"completion_tokens":2932,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":679,"completion_tokens_details":{"reasoning_tokens":2870}},"tokens_in":679,"tokens_out":2932,"duration_ms":23647,"temperature":1.0,"reasoning_tokens":2870,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:12:52.792536+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same five hydrodynamics and two radiative transfer codes for a substantially different regime, such as a 10 Earth-mass planet or a disk with vertical temperature stratification and finite cooling; if retrieved planet positions scatter by more than a beam or brightness temperatures differ by more than the 1.5 K noise level across code combinations, the claim that any tested combination can be used reliably would be refuted.","supporting_citations":[{"cited_title":"exoALMA V: Gaseous Emission Surfaces and Temperature Structures","cited_arxiv_id":"2504.19902","evidence_quote":"Provides the inferred temperature structure that motivates the disk aspect ratio used in the benchmark model."}],"review_version":1}