{"id":"002d60d5-2c4e-4730-8085-c078c6e1146f","arxiv_id":"2411.09096","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":14,"one_line_summary":"Three previously puzzling microlensing events are shown to be binary-lens, binary-source (2L2S) systems, with lens properties estimated via Bayesian analysis.","lead":"Three microlensing events from the KMTNet survey are best explained as binary lenses that each magnify a binary source star. The study adds three new examples of a rare four-body alignment and uses them to estimate the masses and distances of the lensing stars.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 2L2S classification rests on an unquantified dismissal of 3L1S alternatives; no chi-square or residual evidence is given for any of the three events.","rationale":"The reader's weakest_assumption is exactly the load-bearing issue. The paper's strongest claim is that all three events are genuine 2L2S systems, and the decisive competitor, the 3L1S model, is dismissed in one sentence per event with no fits, residuals, or chi-square comparisons. The text itself flags this gap: Section 3 states '3L1S models failed to account for the anomalies in all of the events analyzed' without showing any evidence, and Section 3.3 asserts the 3L1S fit is not equivalent without quantification. For KMT-2021-BLG-0284, the existence of a public 3L1S model by Hirao makes the missing Delta-chi-squared particularly acute, because the paper asks the reader to accept that reanalysis with re-reduced data overturned it. The paper does have independent support: the 2L2S models reproduce four caustic features with self-consistent source trajectories for the first two events, and the close-wide comparison for KMT-2024-BLG-0412 is quantified with Delta-chi-squared = 351.4. However, these strengths do not address the 3L1S degeneracy. The concern is not that the 2L2S interpretation is wrong, only that it is not yet demonstrated to be unique. This is exactly what a conditional verdict should hold open, and the requested test is straightforward and standard in the microlensing literature. I therefore find no reason to change the reader's CONDITIONAL verdict, and I agree with the reader's identification of the weakest assumption. My stress-test adds no new objection beyond confirming and sharpening that gap.","tokens_in":16622,"tokens_out":3518,"duration_ms":37788,"concrete_test":"For each of the three events, rerun the modeling with a full 3L1S parameter search using the same error-normalized KMTNet/OGLE/MOA data: a coarse grid over (s1, q1, s2, q2, alpha) followed by MCMC refinement, and for KMT-2021-BLG-0284 restart from the Hirao April 2021 model. Report the best-fit chi-squared, Delta-chi-squared = chi2(3L1S) - chi2(2L2S), and residual plots at the same scale as Figures 1, 3, and 5. If any 3L1S solution has Delta-chi-squared within about 10 or residuals visually comparable to the 2L2S residuals, the uniqueness of the 2L2S classification fails; if all three are worse by a large margin, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that all three events require a 2L2S model, but the main non-2L1S alternative, the 3L1S model, is dismissed without quantitative support. In the paragraph before Section 3.1, the authors state that '3L1S models failed to account for the anomalies in all of the events analyzed,' and Section 3.3 says a 3L1S interpretation 'does not yield a model with a fit equivalent to that of the 2L2S solution,' but no 3L1S fits, residuals, or chi-square values are shown. For KMT-2021-BLG-0284, a 3L1S model was in fact released by Yuki Hirao on April 16, 2021, and the text says reanalysis showed the 2L2S model 'significantly better explained' the data, again without giving Delta-chi-squared. Because 3L1S models routinely produce multiple caustic spikes and the paper's own Table 1 lists 14 established 3L1S events, this is not an exotic alternative; it is the principal degeneracy threat to the 2L2S classification. If a 3L1S fit with comparable chi-squared and clean residuals exists, then the reported lens masses, separations, distances, and even the binary-source interpretation are not uniquely determined by the data. The 2L2S solutions themselves are presented carefully with caustic geometries, but the missing quantitative model comparison is the load-bearing gap.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript reports the analysis of three KMTNet microlensing events (KMT-2021-BLG-0284, KMT-2022-BLG-2480, and KMT-2024-BLG-0412) whose light curves show complex anomalies. The authors argue that binary-lens single-source (2L1S) models cannot account for all features, while binary-lens binary-source (2L2S) models provide a comprehensive description. For each event they present best-fit lensing parameters, source color-magnitude information, angular Einstein radii where measurable, and Bayesian estimates of the lens masses and distances. They conclude that the lenses are binary stellar systems (M-dwarf pairs or a K dwarf with an M dwarf companion) and that the source binaries have comparable component brightnesses. The central claim is that all three events are genuine 2L2S systems, with the reported physical parameters describing the actual lens and source configurations.","tokens_in":17212,"tokens_out":3342,"duration_ms":37395,"significance":"If the 2L2S interpretation is correct, this paper adds three new events to the small sample of 2L2S microlensing systems and provides physical parameter constraints (mass ratios, separations, distances, source flux ratios) that are useful for studying the demographics of binary stars in the bulge and disk. The strengths of the paper include the use of multi-survey data, explicit caustic-geometry configurations for each event, and quantitative model comparison for the close/wide degeneracy in KMT-2024-BLG-0412, where the Δχ² = 351.4 preference for the wide solution is clearly reported. However, the central uniqueness claim—that the 2L2S model is required over plausible three-lens single-source (3L1S) alternatives—is supported only by qualitative statements; no χ² values or residuals for the rejected 3L1S fits are shown. Since 3L1S events are observationally established and can produce multiple caustic spikes, the missing quantitative comparison is a load-bearing gap that must be closed before the 2L2S classification can be considered secure.","major_comments":[{"comment":"The dismissal of 3L1S models is not quantified. The text states that '3L1S models failed to account for the anomalies in all of the events analyzed' and, for KMT-2024-BLG-0412, that 'a 3L1S interpretation does not yield a model with a fit equivalent to that of the 2L2S solution,' but no χ² values, Δχ², residual plots, or parameter tables for the best 3L1S fits are provided. For KMT-2021-BLG-0284, the paper mentions that Yuki Hirao released a 3L1S model that left noticeable residuals and that reanalysis found the 2L2S model 'significantly better explained' the data, again without a numerical Δχ². Because 3L1S models are a known and observationally established class (Table 1 lists 14 such events) and can produce multiple caustic spikes, this missing quantitative comparison is the principal degeneracy threat to the paper's central claim. The authors should present the best-fit 3L1S solutions, their residuals, and the Δχ² relative to the 2L2S solutions for all three events.","section":"Section 3 (paragraph before 3.1) and Section 3.3"},{"comment":"The claimed inadequacy of the 2L1S models is also stated qualitatively. For each event, the paper says the 2L1S model could not adequately describe the full light curve, and for KMT-2022-BLG-2480 and KMT-2024-BLG-0412 it shows 2L1S fits to restricted data segments, but no χ² values are given for the full 2L1S fits or for the segment-restricted fits used to justify adding a second source. The reader cannot assess whether the improvement from 2L1S to 2L2S is statistically decisive. Please report the best-fit χ² (or Δχ² with degrees of freedom) for the full 2L1S model, the segment-restricted 2L1S model, and the 2L2S model for each event.","section":"Sections 3.1–3.3"}],"minor_comments":[{"comment":"In the paragraph after Table 6, 'KMT-2024-BLG-0496' appears to be a typo for KMT-2024-BLG-0412; please correct it.","section":"Section 4, text near Eq. (1) and Table 6"},{"comment":"The last entry in the KMT-2024-BLG-0412 column, '(V − I, I)S2,0', ends with '0.0110' which appears to be a typographical artifact; it should likely be '0.011' consistent with the other uncertainties.","section":"Table 6"},{"comment":"The phrase 'the flux from the source generating the second set of caustic spikes is estimated to be 0.68 times less than the flux from the source producing the first set' is ambiguous. Since qF = 0.682 in Table 4 and Eq. (3) defines qF as the secondary-to-primary flux ratio, the text should state that the secondary source is fainter by a factor qF = 0.68.","section":"Section 3.2, text after Table 4"},{"comment":"For consistency with Figure 1, consider adding residual panels to Figures 3 and 5 so that the 2L2S fit quality can be directly compared across the three events.","section":"Figures 3 and 5"},{"comment":"The reference list entry 'Zang, W., Hwang, K.-H., Udalski, A., et al. 2021b, 162, 163' is missing the journal name; please verify and complete it.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"This is a well-structured modeling paper that fits the journal's scope. The main issue is the absence of quantitative model comparison for the 3L1S alternatives, which is a standard requirement in the microlensing literature and directly affects the uniqueness of the 2L2S classification. The revision is feasible within the paper's scope: the authors likely have the 3L1S fit results and can report Δχ² and residuals. If they do, the paper may become acceptable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a workmanlike addition to the small 2L2S catalog. Three events with complex caustic-crossing patterns, modeled with the team's standard pipeline. The light-curve fits look good, the caustic geometries are physically sensible, and they are upfront that two of the three events lack a theta_E measurement, so the masses and distances are Bayesian priors plus a timescale, i.e. weakly constrained. That honesty is worth noting.\n\nWhat's actually new: three new events, two with measured Einstein radii, and one close/wide degeneracy cleanly resolved by a large Delta chi^2 = 351.4. The detection-bias point about bright secondaries is a small but real contribution.\n\nThe soft spot is exactly where the stress-test put it. The 2L2S interpretation is the whole show, but the principal alternative—a 3L1S triple-lens model—is dismissed in a single sentence with no chi-square, no residuals, no parameter table. For KMT-2021-BLG-0284 they mention that Hirao's 3L1S model left 'noticeable residuals' and that reanalysis showed 2L2S 'significantly better explained' the data, but no numbers. For the other two events it's just 'failed to account' and 'does not yield a model with a fit equivalent.' Since the paper itself lists 14 established 3L1S events, this isn't an exotic alternative. The missing comparison is load-bearing, not cosmetic.\n\nThat said, I don't think the conclusion is wrong. The 2L1S subtractions shown in the figures do leave unmodeled structure, and the 2L2S geometry is specific enough—two sources crossing different parts of the caustic—that it's likely correct. But the paper doesn't give the reader the evidence to check the 3L1S branch, and for a classification claim that's a real gap.\n\nThe Bayesian physical parameters are standard products from their own Galactic model. That's fine as long as you read them as priors-plus-timescale, which they do state. The self-citation to Jung et al. 2021 is not a problem; it's the model they use everywhere.\n\nWho is this for? Microlensing specialists maintaining the four-body event census. General readers can skip. It deserves peer review—referee time is justified—but the referee should ask for the 3L1S comparison tables and residuals before acceptance.\n\nRecommendation: send it out, with a request for the missing model-comparison numbers.","headline":"Three new 2L2S events modeled carefully, but the paper never shows the numbers that would kill the main alternative, so the classification rests on assertion.","tokens_in":18006,"tokens_out":2158,"would_cite":false,"duration_ms":59339,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["95.30.Sf"],"model":"deepseek-v4-flash","headline":"Three microlensing events decoded as binary lens plus binary source","keywords":["gravitational microlensing","binary lens","binary source","2L2S events","light curve modeling","caustic crossings","M dwarf binaries","Galactic bulge"],"falsifier":"A concrete way to test the central claim is to take the three light curves and attempt a 3L1S model fit with a grid search over triple-lens parameters, then compare chi-square and residuals with the reported 2L2S solutions. The paper does not display this comparison, and for KMT-2021-BLG-0284 a 3L1S model was actually proposed before being rejected, so re-running that comparison is the decisive test. A second check would be high-cadence observation of a similar event with resolved caustic crossings to verify the two-source geometry that the model predicts.","tokens_in":16397,"feed_emoji":"🔭","tokens_out":7647,"duration_ms":63173,"temperature":0.7,"pith_summary":"This paper argues that three microlensing events with unusually intricate light curves are all best understood as 2L2S events, in which a binary lens magnifies a binary source. In each event, a standard binary-lens model could describe parts of the light curve but left one or more caustic features unexplained. Adding a second source to the model reproduces all of the observed anomalies. If the interpretation is correct, the events yield three new examples of 2L2S microlensing, with lenses that are likely M-dwarf binaries or, in one case, a K-type star with an M-dwarf companion, located mostly in the Galactic bulge. The paper also attributes the similar brightness of the two source stars in all three events to a detection bias favoring bright secondaries.","feed_headline":"Binary lens plus binary source explains three microlensing events","feed_subtitle":"Each light curve needed a second source star to fit all its caustic spikes; two events are M-dwarf binaries in the bulge.","key_machinery":"The key machinery is the 2L2S model, in which the observed magnification is the superposition of two source stars lensed by a common binary lens. The binary lens is characterized by its projected separation s and mass ratio q, and each source by its own closest-approach time t0, impact parameter u0, and normalized source radius; the relative brightness of the two sources is set by the flux ratio qF. The modeling strategy is to first fit a 2L1S model to segments of the light curve, then add the second source to account for the unexplained anomalies, and to check that a 3L1S model does not fit comparably. The physical lens parameters are obtained through a Bayesian analysis that uses the measured event timescale and, where available, the angular Einstein radius, along with a Galaxy model and mass function.","core_discovery":"The paper's central claim is that the light curves of KMT-2021-BLG-0284, KMT-2022-BLG-2480, and KMT-2024-BLG-0412 cannot be fully modeled as a binary lens acting on a single source, but are well fit by a 2L2S model in which the same binary lens magnifies two source stars. The reported best-fit parameters include lens mass ratios q from about 0.35 to 0.62 and projected separations s from about 0.81 to 2.79, with the close-wide degeneracy for one event resolved decisively in favor of the wide solution. For two events the lens is estimated to be a binary composed of M dwarfs located in the bulge; for the remaining event the primary is an early K-type main-sequence star and the companion an M dwarf, with the lens more likely in the disk. The paper further claims that the source binaries have similar magnitudes, a consequence of a detection bias that favors events with a brighter secondary source.","pith_inferences":["The paper's dismissal of 3L1S models is stated without quantitative support, so the claim of uniqueness rests on an invisible comparison; a sympathetic reading treats this as a gap to be filled rather than a proven result.","If the detection bias is real, then the fraction of 2L2S events with nearly equal-brightness secondaries should be much higher than the fraction with very faint secondaries; this could be tested by simulating 2L2S light curves with a range of flux ratios and measuring recovery efficiency on the same surveys.","The unresolved caustic crossings in these events mean that the Einstein radii for two of the events are unmeasured or only loosely constrained; future space-based parallax or high-cadence ground-based coverage could sharpen the mass and distance estimates."],"forward_implications":["The sample of known 2L2S microlensing events grows, providing more systems for statistical studies of binary lenses and binary sources.","The Bayesian masses and distances give two new candidate M-dwarf binaries in the bulge and one disk system, which can be compared with models of stellar populations and binary formation.","The detection-bias argument implies that many 2L2S events with faint secondaries are missed, so the true rate of binary-source events is higher than the observed rate.","For KMT-2024-BLG-0412, the wide solution (s ≈ 2.79, q ≈ 0.62) beats the close solution decisively, favoring a particular binary configuration that could be tested with future high-resolution observations."],"supporting_citations":[{"why":"Establishes the binary-lens caustic structure that produces the spike features in the light curves.","marker":"Mao & Paczyński 1991"},{"why":"Founds the binary-source interpretation as the superposition of two single-source lensing events.","marker":"Griest & Hu 1992"},{"why":"Provides the detection-efficiency argument for binary-source events with bright secondaries, used to interpret the similar source magnitudes.","marker":"Han & Jeong 1998"},{"why":"Supplies the Galaxy model and mass function priors used in the Bayesian estimation of lens masses and distances.","marker":"Jung et al. 2021"},{"why":"Describes the photometry pipeline used for reducing the survey data.","marker":"Albrow et al. 2009"},{"why":"Provides the method for recalibrating photometric error bars to normalize chi-square per degree of freedom.","marker":"Yee et al. 2012"}],"fun_headline_variants":["Three microlensing events need a second source star","Binary lens plus binary source solves 3 microlensing events","M-dwarf binaries and double sources fit complex light curves","Why these microlensing spikes required two source stars","Double-source models crack three tricky lensing events"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the three-lens (3L1S) interpretation has been definitively ruled out for all three events; the paper states this conclusion without showing the 3L1S fits, residuals, or chi-square comparisons, so if a 3L1S model fits comparably well the 2L2S conclusion loses its uniqueness.","fun_headline_variants_meta":{"raw":{"variants":["Three microlensing events need a second source star","Binary lens plus binary source solves 3 microlensing events","M-dwarf binaries and double sources fit complex light curves","Why these microlensing spikes required two source stars","Double-source models crack three tricky lensing events"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000349,"raw_usage":{"total_tokens":1985,"prompt_tokens":1104,"completion_tokens":881,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":720,"completion_tokens_details":{"reasoning_tokens":803}},"tokens_in":720,"tokens_out":881,"duration_ms":9306,"temperature":1.0,"reasoning_tokens":803,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T21:03:50.235414+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete way to test the central claim is to take the three light curves and attempt a 3L1S model fit with a grid search over triple-lens parameters, then compare chi-square and residuals with the reported 2L2S solutions. The paper does not display this comparison, and for KMT-2021-BLG-0284 a 3L1S model was actually proposed before being rejected, so re-running that comparison is the decisive test. A second check would be high-cadence observation of a similar event with resolved caustic crossings to verify the two-source geometry that the model predicts.","supporting_citations":[{"cited_title":"1992, ApJ, 397, 362","cited_arxiv_id":null,"evidence_quote":"Founds the binary-source interpretation as the superposition of two single-source lensing events."},{"cited_title":"1998, MNRAS, 301, 231","cited_arxiv_id":null,"evidence_quote":"Provides the detection-efficiency argument for binary-source events with bright secondaries, used to interpret the similar source magnitudes."}],"review_version":1}