{"id":"30b9be31-b230-4daf-b1e5-934cd81b4c8a","arxiv_id":"2601.10464","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"MitoFREQ bounds a mitogenome's population frequency from above as P(TLHG) × (frequency of the profile's rarest SNV within that TLHG), yielding conservative likelihood-ratio evidence weights.","lead":"MitoFREQ estimates how rare a person's complete mitochondrial genome is using only its major haplogroup and the frequency of its rarest single-letter DNA change, harvested from two large public mitogenome databases (HelixMTdb, gnomAD). This matters because forensic labs can now assign numerical weight to mitochondrial DNA evidence — hair, old bones, distant relatives — even when the full profile has never been seen in any database.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'guaranteed conservative' property holds only for true probabilities; MitoFREQ's plug-in estimates, especially the data-dependent rarest-SNV choice, are not shown to be conservative, and the geographic-transfer assumption is one major way the bound can fail.","rationale":"The reader's weakest_assumption (geographic transferability) is a real and important failure mode, and the paper itself flags it. My stress-test goes one level deeper: even if the databases were perfectly matched to the relevant population, the plug-in estimator and the data-dependent choice of the rarest SNV break the mathematical guarantee stated in Eq. (7). The inequality P(X) ≤ P(TLHG)P(X_k|TLHG) is about true probabilities; replacing them with observed frequencies, and then choosing k to minimize the observed frequency, invalidates the claim that the computed LR is conservative. This is not an internal mathematical error in the paper's proof, but it is a gap between the theorem and the method as implemented. The existing CONDITIONAL verdict is therefore appropriate: the mathematical core is sound, but the practical estimator needs uncertainty quantification and/or a conservative bias correction before it can be used in casework. I agree partially with the reader's identification: the geographic-transfer assumption is one specific and severe way the estimate can fail, but the more general concern is that the reported LRs are not proven to be conservative. The proposed concrete test — comparing MitoFREQ estimates to observed population frequencies in the paper's own high-quality datasets — would settle whether the anti-conservative failure is frequent enough to matter. If the test shows systematic underestimation of frequency, the paper would need a substantial revision (e.g., using upper confidence bounds or a penalized rarest-variant frequency); if it shows that the estimates are conservative in practice, the concern would be mitigated. I do not recommend rejecting the paper outright because the underlying bound is correct and the empirical validation may partly rescue the plug-in approach, but the current text overstates the guarantee.","tokens_in":17678,"tokens_out":10501,"duration_ms":115523,"concrete_test":"Compute MitoFREQ LRs for all unique haplotypes in the SWE (n=934) and US2015/US2020 datasets using the authors' protocol, then compare each estimated frequency 1/LR to the observed haplotype frequency in the same population sample (e.g., count/934 for SWE). If more than ~5% of haplotypes have 1/LR below the observed frequency, the plug-in estimator is anti-conservative and the 'guaranteed' bound fails even for a represented population. To isolate the geographic-transfer assumption, repeat the comparison for L0/L1/L2 (African-lineage) haplotypes using Helix SNV frequencies: if estimated frequencies are systematically below observed frequencies for these lineages, the §4 independence assumption is falsified.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that LR = 1/[P(TLHG)P(X_k|TLHG)] is a conservative lower bound on the true weight of evidence (Eqs. 6-7). This inequality is valid for true probabilities, but MitoFREQ computes plug-in estimates: P̂(TLHG) from one database and P̂(X_k|TLHG) from Helix/gnomAD/pooled, with k selected as the rarest observed SNV. For a singleton SNV, P̂(X_k|TLHG) = 1/n_TLHG, so the reported LR is capped near the database size (e.g., 195,983 in Table 4), while the true P(X_k|TLHG) — and hence the true P(X) — can be far larger. Thus the guarantee does not transfer to the reported LR. The data-dependent min-selection biases P̂ downward: choosing the smallest observed frequency among many candidates understates the true frequency of the chosen variant. The geographic-transfer assumption flagged in §4 ('potentially an oversimplification') is a frequent cause of such failure: if the relevant population is underrepresented in Helix, a variant common there can appear as a singleton in the database, making the LR anti-conservative relative to the case-relevant population. The 0.982 correlation in Fig. 2 is computed over all 4,762 shared SNVs and is dominated by common variants; MitoFREQ selects in the rare tail, where relative discrepancies are large. Without an explicit conservative adjustment (e.g., upper confidence bounds) and without validation against population-specific within-TLHG frequencies, the reported LRs can overstate evidence.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes MitoFREQ, a method for estimating the population frequency of a whole mitogenome using only a top-level haplogroup (TLHG) assignment and the frequency of the rarest SNV within that TLHG. The central derivation (§2.3, Eq. 6) shows that, under the assumption P(TLHG|X)=1, P(mitogenome) ≤ P(TLHG)·P(X_k|TLHG) for any SNV k, so the likelihood ratio LR = 1/[P(TLHG)·P(X_k|TLHG)] is a conservative lower bound on the true weight of evidence. The paper further proves (§2.4, Eq. 8) that a finer haplogroup yields a frequency estimate no larger than a coarser one, provided the same data source is used. The method is implemented in an open-source R package, TLHG prediction is reduced to 227 positions with 99.9% rank-1/2 concordance, and SNV frequencies from HelixMTdb and gnomAD are compared (log-correlation 0.982). LR distributions are reported on GenBank, Swedish, and U.S. datasets and compared with Brenner's κ and CGGT estimators.","tokens_in":18028,"tokens_out":8067,"duration_ms":86666,"significance":"The mathematical inequalities in §2.3–2.4 are correct, and the finer-vs-coarser property (Eq. 8) is exact even for plug-in counts from a single sample. The 227-position TLHG panel and the open-source implementation are practical strengths, and the use of large public mitogenome databases is a useful direction for forensic genetics. However, the paper's headline guarantee of a conservative LR is derived for true probabilities, not for the plug-in estimates actually reported. The data-dependent selection of the rarest SNV and the geographic portability of within-TLHG SNV frequencies are load-bearing risks that are not resolved. The paper is potentially valuable but requires substantial revision before its central claim is supportable.","major_comments":[{"comment":"The conservative bound is stated for true probabilities, but the implementation uses plug-in estimates P̂(TLHG) and P̂(X_k|TLHG) and selects the rarest SNV. The inequality does not automatically transfer to the reported LR. A concrete counterexample is in Table 4: for 'Large: Rare Africa L0' (count 1 in GenBank), the Helix plug-in LR is 195,983, i.e., p̂=5.1e-6, whereas the GenBank empirical frequency for that mitogenome is 1/61,295 ≈ 1.6e-5, so the reported LR exceeds the count-based LR by a factor of about 3.2. Thus the plug-in bound can be anti-conservative relative to an independent database. The paper should either apply a conservative adjustment (e.g., upper confidence bounds with a multiple-testing correction) or demonstrate empirically that the bound holds.","section":"§2.3, Eq. (6)–(7); §3.4, Table 4"},{"comment":"The assumption that P(X_k|TLHG) is independent of geographic origin, which the authors themselves flag as 'potentially an oversimplification,' is load-bearing. The 0.982 log-frequency correlation is computed over all shared SNVs and is dominated by common variants, whereas MitoFREQ operates in the rare tail where relative discrepancies are largest. If the case-relevant population is underrepresented in HelixMTdb or gnomAD, a variant common in that population can appear as a singleton in the database, making the LR anti-conservative. Population-specific within-TLHG validation is needed before the method can be considered robust for forensic use.","section":"§4 Discussion; §3.2 Fig. 2"},{"comment":"The comparison to Brenner's κ and CGGT estimators does not validate the conservative property. Agreement of average log10(LR) with these count estimators is not evidence that MitoFREQ's LRs are lower bounds on the true weight of evidence. The paper should directly validate the bound by checking, for the GenBank/US/SWE datasets, whether P̂(TLHG)P̂(X_k|TLHG) ≥ the observed frequency for each haplotype, rather than only reporting LR distributions.","section":"§3.3, Fig. 3"}],"minor_comments":[{"comment":"The assumption P(TLHG|X)=1 is not guaranteed in practice. The rank-1/2 concordance is 99.9%, so for about 0.1% of sequences the bound can fail. The paper should state that the method is intended for cases where TLHG assignment is unambiguous, or model the assignment uncertainty.","section":"§2.3, Eq. (5)"},{"comment":"The rarest-SNV selection is data-dependent and no multiple-testing adjustment is mentioned. A brief statement that this is the minimum over the profile, and that the computed LR is therefore a plug-in minimum not a fixed choice, would help readability.","section":"§3.4, Table 4"},{"comment":"The horizontal line of points at high gnomAD frequency but variable Helix frequency is noted as possible errors. This should be discussed more explicitly, as it affects the reliability of pooled estimates and the rare-tail comparison.","section":"§3.2, Fig. 2"}],"recommendation":"major_revision","confidential_remarks":"The Table 4 counterexample is concerning: it suggests the plug-in LR can be several times larger than the simple count-based LR in the paper's own data. If the authors cannot show that the bound holds empirically after appropriate adjustments, the central claim of conservative evidence evaluation would be unsupported. I would encourage a revision that adds uncertainty quantification and a direct validation of the inequality on the test datasets."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Do you know this MitoFREQ paper? It claims a simple way to turn top-level haplogroup and single-SNV frequencies from HelixMTdb/gnomAD into a conservative match probability for mitogenomes, and it mostly earns that claim at the level of probability theory. The genuinely new things are the estimator itself (TLHG frequency times the rarest within-TLHG SNV frequency), the 227-position panel that recovers the TLHG with ~99.9% accuracy, and the clean monotonicity result that a finer haplogroup gives a LR at least as high as a coarser one. The math is elementary the way good forensic math should be: the inequalities in Section 2 are correct and lead to a genuinely useful design principle. That part of the paper is solid.\n\nThe paper also does real service by packaging the method in an open-source R package with a Shiny app, and by testing on both curated forensic mitogenomes and a scrutinized GenBank set. The comparison to Brenner/CGGT singletons is useful context, though not a direct benchmark.\n\nNow the soft spots. The 'guaranteed conservative' bound in Eq. (6) holds for true probabilities. What MitoFREQ reports are plug-in estimates: P-hat(TLHG) from one database, P-hat(X_k | TLHG) from Helix, gnomAD, or pooled, with k selected as the rarest observed SNV. The data-dependent min-selection biases the estimate downward, and for a singleton variant the estimate is 1/n, which can overstate the LR. So the guarantee does not transfer to the reported LR as it stands. The authors do flag 'geographic independence' as potentially an oversimplification, and they note the African under-sampling in Helix — good, but the supporting correlation (0.982) is dominated by common variants, while the method operates in the rare tail. That is precisely where the bound can fail relative to the case-relevant population.\n\nThere is also a mixed-source issue in Table 4: GenBank haplotypes are evaluated with Helix TLHG and SNV frequencies, so the LRs can sit far above those haplotypes' own observable frequencies. That is not a flaw in the math, but it underscores that the reported numbers are not conservative in the operational sense without an explicit reference-population argument.\n\nIf the authors add either conservative confidence bounds (e.g., upper confidence limits for P-hat) or reframe the output as point estimates without a guarantee, and if they validate within-TLHG SNV frequencies on population-specific data, the paper would be ready for casework. As is, it reads as a promising method that needs one round of revision.\n\nMy own verdict: the central argument holds up; this is not a desk reject. The paper deserves a serious referee, and I would send it out. It is aimed at forensic geneticists, particularly those dealing with low-quality or partial mitogenome profiles. I would probably cite it for the monotonicity result, and I'd bring it to a reading group if we had one discussing forensic lineage marker statistics.","headline":"The probability inequalities are correct and the estimator is genuinely new, but the reported LRs are plug-in estimates whose 'guaranteed conservative' property does not automatically transfer to casework.","tokens_in":18614,"tokens_out":1973,"would_cite":true,"duration_ms":21352,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62P10","92D10"],"pacs":[],"model":"deepseek-v4-flash","headline":"A full mitogenome's frequency can be conservatively bounded using just its top-level haplogroup and one rare SNV.","keywords":["mitogenome","forensic genetics","likelihood ratio","population frequency","haplogroup","single nucleotide variant","mtDNA","match probability"],"falsifier":"Take a set of mitogenomes from an undersampled population (e.g., a specific African ethnic group), compute P(TLHG)P(X_k | TLHG) from HelixMTdb/gnomAD for their rarest SNVs, and compare the product against the observed frequency of those mitogenomes in a population-matched reference database such as EMPOP's population-specific sets. If, for a substantial fraction of profiles, the computed product is smaller than the observed frequency, the conservativeness claim fails for that population.","tokens_in":17515,"feed_emoji":"🧬","tokens_out":3789,"duration_ms":39054,"temperature":0.7,"pith_summary":"MitoFREQ is a method for estimating the population frequency of a whole mitochondrial genome when the exact sequence has no match in a reference database. It assigns the mitogenome to one of 30 top-level haplogroups and multiplies the haplogroup's frequency by the frequency of the profile's rarest SNV within that haplogroup, using publicly available data from HelixMTdb and gnomAD. The method is proven to give an upper bound on the true mitogenome frequency, hence a conservative likelihood ratio, and it is guaranteed to give a higher frequency estimate than would be obtained with a finer (more specific) haplogroup if such data existed. The authors also show that 227 selected positions are sufficient to infer the top-level haplogroup in 99.9% of tested mitogenomes, making the approach applicable to degraded or partial profiles. On high-quality forensic datasets and diverse GenBank sequences, the method yields likelihood ratios in the range of 100–100,000.","feed_headline":"One rare SNV gives a conservative bound on mitogenome frequency","feed_subtitle":"Forensic labs can now compute mtDNA match statistics from HelixMTdb and gnomAD without a full-sequence database match.","key_machinery":"The central identity is the nested inequality chain P(mitogenome) = P(X, TLHG) ≤ P(X_k, TLHG) = P(TLHG)P(X_k | TLHG), obtained by telescoping the joint probability and using the assumption P(TLHG | X) = 1. Combined with monotonicity of joint probabilities—P(H, X_k) ≤ P(G, X_k) whenever H ⊆ G—this guarantees that finer haplogrouping only increases the LR. The practical enabler is a panel of 227 mitogenome positions that reproduces top-level haplogroup calls from full-sequence data in 99.9% of tested cases, allowing the method to work on partial profiles.","core_discovery":"Under the assumption that the top-level haplogroup (TLHG) is known with certainty for the mitogenome X, the paper proves that for any SNV k in the profile, P(mitogenome) = P(X, TLHG) ≤ P(X_k, TLHG) = P(TLHG)P(X_k | TLHG). Therefore LR = 1/[P(TLHG)P(X_k | TLHG)] is a conservative lower bound on the true weight of evidence. The paper further proves a monotonicity property: if H is a finer haplogroup contained in G (e.g., M1 within M), then P(H)P(X_k | H) ≤ P(G)P(X_k | G), so using the coarser TLHG cannot underestimate the true frequency as long as the TLHG assignment is correct. This makes the method a valid, conservative source of match probabilities even when only top-level haplogroup and SN","pith_inferences":["The method's conservativeness relative to finer haplogroups suggests a path for incremental improvement: as fine-haplogroup SNV frequency data accumulate, LRs can only increase, so MitoFREQ provides a durable ceiling on frequency estimates.","The 0.982 log-frequency correlation between HelixMTdb and gnomAD is dominated by common variants; the method operates on the rare tail, where cross-database agreement is more uncertain and where a population-matched test would be most informative.","For populations undersampled in HelixMTdb (e.g., some African lineages), the assumption that SNV frequencies transfer across geographic origin may break down; a concrete test would compare MitoFREQ estimates against observed frequencies in a population-matched reference set.","The framework could be extended to multiple SNVs under a conditional-independence assumption within a TLHG, potentially tightening the upper bound without requiring full mitogenome counts—an extension the paper leaves open."],"forward_implications":["Forensic laboratories can compute mitogenome match statistics without requiring a full-sequence match in EMPOP or similar databases, drawing instead on the much larger HelixMTdb and gnomAD resources.","The likelihood ratio obtained from MitoFREQ is a conservative lower bound on the true weight of evidence, so it will not overstate the support for a match as long as the underlying assumptions hold.","Because only 227 positions are needed for top-level haplogroup inference, the method extends to degraded or partial mtDNA profiles that cannot be fully sequenced.","The TLHG frequency can be replaced by any population-specific distribution (e.g., country-specific), while SNV frequencies can be pooled across HelixMTdb and gnomAD, making the method adaptable to case-relevant reference populations.","Using a finer haplogroup, when data become available, can only increase the LR, so current estimates are a lower bound on future, more refined estimates."],"fun_headline_variants":["One rare SNV gives conservative mtDNA frequency","Conservative mitogenome frequency from a single SNV","Forensic mtDNA stats from top-level haplogroups","Proven bound: mitogenome frequency from rarest SNV","MitoFREQ: conservative mtDNA frequency estimation"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that the frequency of a SNV within a top-level haplogroup does not depend on geographic origin, so that frequencies from HelixMTdb and gnomAD can be applied to the case-relevant population; if this transfer fails for rare variants in an undersampled population, the claimed conservative upper bound can be violated and LRs could overstate evidence.","fun_headline_variants_meta":{"raw":{"variants":["One rare SNV gives conservative mtDNA frequency","Conservative mitogenome frequency from a single SNV","Forensic mtDNA stats from top-level haplogroups","Proven bound: mitogenome frequency from rarest SNV","MitoFREQ: conservative mtDNA frequency estimation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000438,"raw_usage":{"total_tokens":2167,"prompt_tokens":953,"completion_tokens":1214,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":697,"completion_tokens_details":{"reasoning_tokens":1134}},"tokens_in":697,"tokens_out":1214,"duration_ms":8733,"temperature":1.0,"reasoning_tokens":1134,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T10:18:32.059854+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a set of mitogenomes from an undersampled population (e.g., a specific African ethnic group), compute P(TLHG)P(X_k | TLHG) from HelixMTdb/gnomAD for their rarest SNVs, and compare the product against the observed frequency of those mitogenomes in a population-matched reference database such as EMPOP's population-specific sets. If, for a substantial fraction of profiles, the computed product is smaller than the observed frequency, the conservativeness claim fails for that population.","supporting_citations":[],"review_version":1}