{"id":"9e21ee17-894e-4c62-a4c5-3c7d2b1aba8b","arxiv_id":"2412.09705","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A new survey of 47 Tuc finds no hot Jupiter transits and, combined with HST data, caps hot Jupiter occurrence at 0.11%.","lead":"This paper searched 19,930 stars in the globular cluster 47 Tucanae for transiting hot Jupiters and found none. Combined with the 2000 Hubble search of the cluster core, it sets the tightest upper limit yet on hot Jupiter occurrence, about four times below the rate seen by Kepler.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 0.11% upper limit assumes detection efficiency measured without the detailed second-pass vetting, which is explicitly admitted as likely not 100% efficient and is never calibrated on injected transits.","rationale":"The reader's weakest assumption—that the measured detection efficiency equals the real transit-recovery efficiency—is correct and is the most load-bearing point for the paper's headline number. The paper itself provides two explicit acknowledgments of this weakness: footnote 19 (injections into detrended lightcurves) and Section 8 (second-pass vetting likely not 100% efficient). Neither is propagated into the reported efficiency or the uncertainty on the upper limit. The concern is concrete because it is quantifiable: the missing second-pass efficiency can in principle be measured by running injected transits through the same detailed vetting. My assessment of the quantitative impact is that the 'strongest limit to date' claim is robust—G00 alone yields f_HJ < 0.16%, so even a factor-of-two overestimate of MISHAPS sensitivity leaves the combined limit below 0.16%. However, the specific value 0.11% and the factor-of-four comparison to the Kepler field are not robust. For this reason the verdict CONDITIONAL is appropriate; no change to the reader's verdict is needed.","tokens_in":51884,"tokens_out":9364,"duration_ms":100006,"concrete_test":"Blind a sample of several hundred injected transit lightcurves that pass both the S/N threshold and the Zooniverse first-pass stage, then run them through the full detailed second-pass vetting procedure (500x500 cutout difference imaging, period searches, image stacking) exactly as was done for the real candidates in Section 6. Record the retention fraction; call it epsilon_detailed. Then recompute N1 = N* * epsilon_total * epsilon_detailed and the combined 95% upper limit f_HJ = 3 / N1. If the resulting limit exceeds 0.18%—the MW17 Kepler-field equivalent—the paper's stated scientific conclusion changes; if it remains below 0.11%, the concern is empirically resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central numerical claim, f_HJ < 0.11%, is computed as 3/N1 (Eq. 22), where N1 = N* * epsilon_total (Eq. 21). The efficiency epsilon_total (Eq. 17) is built only from the algorithmic detection efficiency (epsilon_det) and the Zooniverse first-pass approval fraction (epsilon_Zoo). The detailed second-pass vetting described in Section 6 (target-centered cutout photometry, period searches, stacked difference images) is applied only to the 39 real candidates that survived the first pass; it is never applied to the ~40,000 injected transit lightcurves. Section 8 explicitly states: 'it is likely that the remaining vetting steps we take also are not 100% efficient.' If that second pass rejects even a small fraction of genuine transits—for example, because the cutout photometry fails for faint or blended targets—then epsilon_total is overestimated, N1 is too large, and the reported upper limit is too small. A second independent bias is flagged in footnote 19: transits are injected into already-detrended lightcurves, so the search efficiency does not account for possible absorption of real transit signals by the TFA detrending. Both effects act in the same direction, making the survey appear more sensitive than it is. Quantitatively, if the unmeasured second-pass retention fraction is 80%, MISHAPS N1 drops from 830 to ~664 and the combined limit worsens from 0.110% to ~0.117%; at 50% retention it becomes ~0.130%. The conclusion that the Kepler-field equivalent rate (MW17: 0.18%) is ruled out survives even if MISHAPS contributed nothing (G00 alone gives 0.16%), but the abstract's 'factor of ~4 below the Kepler field' would degrade to a factor of ~2.5-3, and the exact claimed strength of 0.11% is not robust.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents the first results of the MISHAPS ground-based survey for transiting hot Jupiters in the globular cluster 47 Tucanae, using ~24 nights of DECam r/z time-series photometry. The authors analyze 19,930 likely cluster members selected by Gaia proper motions and a color-magnitude cut, search for single and partial transits with a boxcar algorithm, and characterize their detection efficiency with ~40,000 injected transits that pass through the algorithmic search and Zooniverse first-pass human vetting. They report no surviving planet candidates, reject 35 initial transit candidates through detailed vetting, identify 4 eclipsing binaries, and derive a 95% upper limit of f_HJ < 0.43% for their survey alone over 0.75-2.0 R_Jup and 0.5-10 days. Combining with the G00 HST survey over the overlapping range 0.8-2.0 R_Jup and 0.5-8.3 days, they quote f_HJ < 0.11%, which they describe as the strongest limit to date and a factor of ~4 below the Kepler-field occurrence rate.","tokens_in":52252,"tokens_out":6018,"duration_ms":61952,"significance":"The survey addresses a genuinely open question: whether hot Jupiter formation is suppressed in the low-metallicity, high-stellar-density environment of a globular cluster. The pipeline is careful and transparent in several respects: injection-recovery simulations are performed over the actual stellar sample; the first-pass human Zooniverse vetting efficiency is explicitly measured rather than assumed to be 100%; proper-motion and color cuts remove foreground and SMC contamination; and the injection-recovery products are publicly released. The 4 new eclipsing binaries are a useful byproduct. However, the central quantitative claim, the combined 0.11% upper limit, depends on two efficiency terms that the paper itself flags as uncalibrated (the second-pass detailed vetting and the post-detrending injection procedure) and on an adopted value of G00's sensitivity that the paper's own Section 8 shows to be sensitive to the assumed planet population. Because both uncalibrated effects act in the same direction, the quoted limit is likely too stringent as a stated 95% confidence bound.","major_comments":[{"comment":"The total efficiency used in the N1 calculation includes only the algorithmic detection efficiency and the Zooniverse first-pass approval fraction; the detailed vetting described in Section 6 (target-centered cutout photometry, period searches, stacked difference images) is applied only to the 39 real candidates and never to the injected transits. Section 8 explicitly states that 'the remaining vetting steps we take also are not 100% efficient.' Any real transit rejected in the second pass reduces the true N1 and weakens the upper limit, so the reported f_HJ < 0.11% is biased low. The authors should either calibrate the second-pass efficiency by injecting synthetic transits through that full procedure, or apply and propagate a conservative correction factor (e.g., a range of assumed retention fractions).","section":"Section 8 and Eq. (17)-(22)"},{"comment":"The transit injections are added after the TFA detrending step, so the computed efficiency does not account for the possibility that TFA partially absorbs real transit signals when they are present in the original lightcurves. The paper acknowledges this in footnote 19 as a future fix. Since this effect also makes the survey appear more sensitive than it actually is, it directly impacts the central upper limit. At minimum, the authors should estimate the size of this effect, for example by injecting before detrending on a subset of lightcurves and comparing the recovered efficiency, or by citing published estimates of TFA's suppression of transit signals.","section":"Section 5.1, footnote 19"},{"comment":"The combined limit uses G00's N1 = 1889, derived from G00's expected yield of 17 planets at an assumed 0.8-1.0% occurrence rate. The paper itself notes in Section 8 that MW17's recalibration of G00's sensitivity implies an effective N1 about two-thirds as large, and the authors compute that a reweighted combination gives a combined N1 of 1776 and f_HJ < 0.17% rather than 0.11%. Because the headline claim 'strongest limit to date, factor of ~4 below the Kepler field' depends on the choice of G00's N1, the authors must present the combined limit under both calibrations and either justify the original G00 value as the appropriate one for a uniform period-radius definition or lead with the more conservative value.","section":"Section 7, Eq. (24) and Table 5"},{"comment":"The reported 95% upper limit propagates only Poisson counting statistics (3/N1). Systematic uncertainties in the stellar radius estimates (§3.4.3, which feed the transit-depth and transit-probability calculations), the spline photometric transforms in Table 3 (particularly the ±0.08 mag residual in the (r-z)PS1 to (g-i)PS1 transform), the fixed choice of 15 TFA trend stars (§4.2), and the adopted G00 N1 are not propagated into the final limit. Since the paper's main result is a quantitative bound, the authors should provide a systematic error budget or demonstrate that the limit is robust to these choices; without this, the 0.11% figure is presented with overstated precision.","section":"Section 7, Eq. (22) and Eq. (25)"}],"minor_comments":[{"comment":"The abstract states the limit is 'a factor of ~4 below the occurrence rate in the Kepler field', but Section 8 compares against MW17's 0.18% rate, which would be a factor of 1.6. The factor of ~4 appears to refer to Fressin et al.'s 0.43% rate over a longer period range. The authors should specify which comparison is being made in the abstract to avoid the apparent inconsistency.","section":"Abstract and Section 8"},{"comment":"The paper acknowledges using the same limb-darkening coefficients in z as in r, but leaves the impact unquantified. A sentence estimating the resulting error in transit depth or recovery efficiency would clarify whether this is truly negligible for the reported limits.","section":"Section 5.1"},{"comment":"The definition of C_j states that a night counts if the classification is 'partial or full transit unanimously for all users', but with Nuser=2 it is not explicit whether both users must classify the same night as a transit, or whether one user's transit classification plus the other's abstention counts. Please clarify the unanimity rule.","section":"Section 5.4, Eq. (15)"},{"comment":"Several figure captions read only 'Same as previous' without identifying which candidates are shown in the figure. The captions should list the candidate IDs so the figures are self-contained.","section":"Figures 22-25"},{"comment":"The sentence 'The search returns an estimated depth of 0.018 and duration of 1.0 hr for this eclipse for this eclipse' contains a duplicated phrase; also, 'V-shaped bottom' should be introduced as a technical term or placed in quotes.","section":"Section 6.2.2"}],"recommendation":"major_revision","confidential_remarks":"The paper is unusually honest about its own limitations, explicitly flagging in footnote 19 and Section 8 the two main efficiency biases and the G00 sensitivity ambiguity. This honesty supports a major-revision rather than a rejection: the survey itself and the injection-recovery framework are solid, and the issues are fixable within the paper's scope by recalibrating the second-pass vetting efficiency, quantifying the detrending effect, and presenting the combined limit under alternative G00 calibrations. My main concern for the editor is communication: the abstract and conclusions currently present the 0.11% limit without the caveats that the body of the paper itself contains, which could mislead readers about the robustness of the headline number."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Greg – read the MISHAPS I paper. Worth your time if you care about globular cluster planets or upper-limit methodology.\n\nWhat's actually new: first wide-field DECam survey of 47 Tuc's outer regions, 19,930 stars with proper-motion and CMD membership cuts, a single-transit search strategy, and an injection-recovery pipeline that includes a quantified human-vetting efficiency for the first Zooniverse pass. No planets found, 35 candidates ruled out without follow-up, 4 EBs found. Combining their independent N1=830 with G00's HST inner-cluster N1=1889 gives f_HJ<0.11% for 0.8–8.3 days and 0.5–2 Rjup, the strongest limit yet and a factor of ~2.3 stronger than G00 alone.\n\nPaper is transparent: data products on Astro Data Lab and Zenodo, detailed appendices, and the authors flag the two main caveats themselves. No circularity in the main result.\n\nThe soft spots are real but not fatal. First, transits are injected into already-detrended lightcurves (footnote 19), so the search efficiency misses any transits TFA would have partially absorbed; this makes the limit look better than it is. Second, the second-pass detailed vetting – cutout photometry, period searches, stacked difference images – was applied to the 39 real candidates but never to injected transits. Section 8 admits those steps are likely not 100% efficient. If the second pass is 80% efficient, the combined limit worsens to ~0.117%; at 50%, ~0.13%. The qualitative conclusion survives – even G00 alone gives 0.16% against MW17's 0.18% Kepler-field prediction – but the abstract's 'factor of ~4 below' is optimistic at the high end; a factor of ~2.5–3 is more defensible. Also, the Kepler comparison uses slightly different radius/period ranges than the headline limit, and the radii/photometric transform systematics aren't propagated. Minor for an upper-limit paper.\n\nVerdict: solid, honest work that deserves a serious referee. I'd recommend acceptance after a minor revision asking for a sensitivity analysis on second-pass efficiency and a statement that the quoted limit does not include that uncertainty.\n\nReading group: yes.","headline":"A careful, honest upper-limit paper that strengthens the 47 Tuc hot Jupiter constraint to 0.11%, albeit with uncalibrated second-pass vetting and post-detrending injections keeping the exact number soft.","tokens_in":53053,"tokens_out":3115,"would_cite":true,"duration_ms":32648,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that hot Jupiters in 47 Tucanae are at least four times rarer than in the Kepler field, with a combined 95% upper limit of f_HJ < 0.11%.","keywords":["Exoplanet astronomy","Transit photometry","Hot Jupiters","Globular star clusters","47 Tucanae","Occurrence rate","Detection efficiency"],"falsifier":"An injection-recovery test that adds synthetic transits before detrending and independently audits the second vetting step; if the true recovery fraction falls materially below the paper's measured efficiency, the combined $N_1 = 2719$ and the 0.11% upper limit would be too optimistic.","tokens_in":51694,"feed_emoji":"🔭","tokens_out":6488,"duration_ms":61770,"temperature":0.7,"pith_summary":"This paper tries to establish that hot Jupiters are genuinely scarce in the globular cluster 47 Tucanae, not just undetected. Searching 19,930 outer-cluster stars with a wide-field ground-based camera on a 4-meter telescope, the authors find no convincing planets, rule out all 35 transit candidates as false positives without follow-up, and combine their sensitivity with the 2000 Hubble search of 34,091 inner-cluster stars. The resulting 95% upper limit is $f_{\\rm HJ} < 0.11\\%$ for planets with periods 0.8 to 8.3 days and radii 0.5 to 2.0 Jupiter radii, about four times lower than the hot Jupiter rate measured in the Kepler field. If the limit holds, it constrains how giant planets form in metal-poor, densely packed stellar environments and tests whether enhanced $\\alpha$-element abundances can compensate for low iron.","feed_headline":"Fewer hot Jupiters in 47 Tuc: new limit is 0.11%","feed_subtitle":"Ground-based plus Hubble data show the cluster's hot Jupiter rate is roughly four times below the Kepler field's.","key_machinery":"The central machinery is a single-transit search rather than a phased multi-transit search. A sliding boxcar scans each night's detrended lightcurve for a transit-shaped dip and requires a signal-to-noise ratio of at least 7; the telescope's aperture makes a Jupiter-radius transit detectable over several magnitudes of the cluster main sequence, so even one partial transit can be found. Detection efficiency is calibrated by injecting about 40,000 synthetic transits into the real lightcurves, measuring recovery through the automated search and the first human vetting step, which the paper finds to be near 90% efficient, and folding in the geometric transit probability. The quantity $N_1 = N_\\star \\times \\epsilon_{\\rm total}$, the expected number of planets if every star had one, converts a null result into an occurrence limit through $f_{\\rm HJ} < 3/N_1$.","core_discovery":"On its own, the new survey's 19,930 stars yield $N_1 = 830$ and a 95% upper limit $f_{\\rm HJ} < 0.36\\%$ over the same period and radius range as the earlier Hubble search. Because the two surveys cover independent samples, the new one in the cluster's outskirts and Hubble's in the core, their expected yields add, giving $N_1 = 2719$ and a combined limit $f_{\\rm HJ} < 0.11\\%$ for hot Jupiters with $0.8 \\leq P \\leq 8.3$ days and $0.5 \\leq R \\leq 2.0\\,R_{\\rm Jup}$. The paper argues this is the strongest limit to date and concludes that the occurrence rate of hot Jupiters in 47 Tuc is roughly four times below that of the Kepler field.","pith_inferences":["Because synthetic transits were added after detrending, a signal that detrending would partially erase could make the measured recovery efficiency optimistic; the paper itself flags this as a future fix.","If the sensitivity is as claimed, the limit already approaches the occurrence rate predicted when alpha-element abundance rather than iron sets planet formation, about 0.055%, leaving a narrow window to discriminate between the two hypotheses.","The single-transit observing strategy, using many short windows instead of continuous coverage, could be applied to other globular clusters or crowded fields where multi-transit searches are impractical.","The three newly cataloged detached eclipsing binaries are a byproduct of the search that may serve as independent tracers of the cluster's binary population and dynamics."],"forward_implications":["If the limit is right, hot Jupiters occur in 47 Tuc at least four times less often than in the Kepler field, making the cluster a genuinely different planet formation environment.","The result rules out, at 95% confidence, the occurrence rate expected if the cluster's stars hosted hot Jupiters at the same rate as Kepler stars of similar mass, before metallicity corrections are applied.","The quantified human vetting efficiency, near 90% and lower at longer periods, shows that visual inspection cannot be treated as perfect in future transit surveys and must be included in occurrence limits.","The $N_1$ framework gives a reusable way to combine independent null searches, as demonstrated by merging the outer-cluster survey with the inner-cluster Hubble search.","Extending the survey to the central chips and to fainter stars should push the combined sensitivity toward the predicted alpha-element-enhanced rate, which would require $N_1 \\approx 5450$ to rule out."],"supporting_citations":[{"why":"Supplies the independent 34,091-star inner-cluster Hubble search whose expected yield is added to this survey's to form the combined 0.11% limit.","marker":"R. L. Gilliland et al. (2000)"},{"why":"Re-analysis showing the Hubble null result is consistent with a Kepler-like occurrence rate; motivates the new survey and gives the 0.18% Kepler-field prediction for 47 Tuc-like stars.","marker":"K. Masuda & J. N. Winn (2017)"},{"why":"Previous ground-based 47 Tuc survey whose sensitivity is compared directly through $N_1$ in the same period and radius bins.","marker":"D. T. F. Weldrake et al. (2005)"},{"why":"Provides the Kepler-field hot Jupiter occurrence rates, 0.082% and 0.43%, used to measure the factor-of-four deficit.","marker":"F. Fressin et al. (2013)"},{"why":"Supplies the giant-planet occurrence dependence on host mass and metallicity used to predict 47 Tuc's rate under iron-based versus alpha-element-based scaling.","marker":"J. A. Johnson et al. (2010)"},{"why":"Provides 47 Tuc's [Fe/H] and [alpha/Fe] abundances used in the metallicity-scaled predictions.","marker":"M. J. Cordero et al. (2014)"}],"fun_headline_variants":["47 Tuc hot Jupiter rate capped at 0.11%","Strongest hot Jupiter limit yet in globular cluster","MISHAPS + Hubble: hot Jupiters rare in 47 Tuc","Globular cluster hot Jupiter occurrence below 0.11%","No hot Jupiters in 47 Tuc: new upper limit 0.11%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"That the measured recovery of injected, already-detrended synthetic transits, including the first human vetting pass, equals the real probability that a hot Jupiter transit would have been found, and that the later, unquantified vetting steps do not discard real planets.","fun_headline_variants_meta":{"raw":{"variants":["47 Tuc hot Jupiter rate capped at 0.11%","Strongest hot Jupiter limit yet in globular cluster","MISHAPS + Hubble: hot Jupiters rare in 47 Tuc","Globular cluster hot Jupiter occurrence below 0.11%","No hot Jupiters in 47 Tuc: new upper limit 0.11%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000951,"raw_usage":{"total_tokens":4121,"prompt_tokens":1073,"completion_tokens":3048,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":689,"completion_tokens_details":{"reasoning_tokens":2954}},"tokens_in":689,"tokens_out":3048,"duration_ms":17934,"temperature":1.0,"reasoning_tokens":2954,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T16:48:58.641307+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"An injection-recovery test that adds synthetic transits before detrending and independently audits the second vetting step; if the true recovery fraction falls materially below the paper's measured efficiency, the combined $N_1 = 2719$ and the 0.11% upper limit would be too optimistic.","supporting_citations":[],"review_version":1}