{"id":"0590f3eb-8332-4b47-bac5-b29f3050036c","arxiv_id":"2607.14473","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Standard two-latent-class diagnostic models can mislabel latent classes when tests measure different measurands; modeling measurands explicitly via DAGs yields unbiased estimates and reinterprets two published analyses.","lead":"The paper shows that standard two-class statistical models for imperfect diagnostic tests can misidentify which condition they are measuring when tests target different biological quantities. It offers a graphical method to define the right latent classes and demonstrates the fix on simulated and real data for tuberculosis and leptospirosis.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The structural zero in Table 1 — x-ray abnormalities require active TB — is clinically questionable; if other respiratory disease can cause abnormalities, the 4LC TB estimates and the headline conclusion are not robust.","rationale":"The strongest claim — that a 2LC model can recover the dominant measurand rather than the target condition — is supported by the simulation in Section 4.2; that part is credible. The proposed remedy, however, requires the DAG and its implied restrictions to be correct. The reduced TB DAG in Figure 1b explicitly removes 'Other respiratory disease' in Section 2.1, and Table 1 then encodes the restriction that intrathoracic abnormalities cannot occur without active TB. This is a substantive clinical assumption: in pediatric respiratory cohorts, x-ray abnormalities are common in other respiratory infections. If that assumption fails, the 4LC model cannot represent the true data-generating process, and the reported accuracy and prevalence estimates for x-ray and active TB are biased. The issue is not nuisance misspecification; it directly affects the headline applied conclusion that the 2LC prevalence (0.27) overestimates active TB (0.22). The internal parameter-count inconsistency in Section 5.1 (13 vs 14 vs 19) and the unshown 'unbiased 4LC' simulation are secondary and addressable with additional details. The structural-zero concern is more fundamental, and the paper offers no sensitivity analysis to it. The reader's verdict (CONDITIONAL) already captures this uncertainty; our proposed alternative-model fit would determine whether the concern actually changes the TB estimates. If it does, the central applied claim would need to be weakened or substantially qualified.","tokens_in":11491,"tokens_out":5333,"duration_ms":52361,"concrete_test":"Re-analyze the Schumacher et al. pediatric TB data with a 5-latent-class version of the Table 1 model: add an 'other respiratory disease' latent variable that can cause x-ray abnormalities, relax the (0,0,1) structural zero, and assign a plausible prior to the prevalence of other respiratory disease (e.g., Beta(10,20) based on pediatric pneumonia literature). Compare the posterior median and 95% CrI for active TB prevalence and x-ray sensitivity to the reported 0.22 (0.18–0.28) and 0.95 (0.82–1). If active TB prevalence moves outside the reported CrI, or x-ray sensitivity changes by more than 0.1, the dropped 'other respiratory disease' node is load-bearing and the reported TB conclusions are conditional on an untestable restriction.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central applied claim — that the 2LC model overestimates active TB prevalence (0.27 vs 0.22) — rests on the structural-zero restrictions in Table 1, which are derived from the reduced DAG in Figure 1b. Section 2.1 drops 'Other respiratory disease' because it is 'not estimable' and relabels the x-ray measurand as 'Intrathoracic abnormalities due to TB.' Table 1 then declares the latent class (Active TB=0, TB Infection=0, Intrathoracic abnormalities=1) impossible. But in a pediatric population with respiratory symptoms, chest x-ray abnormalities frequently arise from pneumonia or other respiratory infections without active TB. Because no observed test measures 'other respiratory disease,' the model cannot distinguish such children from class 4 (0,0,0); it must absorb them into the no-abnormality class. This directly inflates the estimated prevalence of class 4 (0.52, 95% CrI 0.34–0.64) and biases x-ray sensitivity with respect to its measurand (reported 0.95, 0.82–1). The paper presents no sensitivity analysis to these restrictions. Moreover, the model is already only locally identifiable with rank 12 vs 14 parameters, so the estimates lean on informative priors; the combination of fixed structural zeros and informative priors makes the headline TB conclusion fragile. The leptospirosis example has an analogous issue: dropping 'other infection' assumes IgM/IgG arise only from leptospirosis, which is not clinically defensible.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper argues that conventional two-latent-class (2LC) latent class models for diagnostic test accuracy can misidentify the latent variable when tests do not all measure the same target condition. It introduces directed acyclic graphs (DAGs) to explicitly separate each test's measurand from the target condition, and proposes expanded latent class models with classes defined by combinations of target condition and distinct measurands. The authors derive the mixture likelihood in Eq. (1), discuss identifiability in terms of degrees of freedom and the Jacobian rank, and use simulations to show that when the majority of tests measure a non-target measurand, the 2LC model tends to label that measurand rather than the target condition. They re-analyze two published datasets (pediatric pulmonary tuberculosis and leptospirosis) and report that the 2LC model in the TB example overestimates active TB prevalence (0.27 vs 0.22 under the proposed 4LC model) and that in the leptospirosis example the 2LC latent class appears to correspond to IgM rather than Leptospirosis infection.","tokens_in":11852,"tokens_out":5680,"duration_ms":59110,"significance":"The methodological message is important and clearly demonstrated by the simulation: ignoring the measurand structure of diagnostic tests can lead to incorrectly labeled latent classes and biased accuracy/prevalence estimates. The DAG-based framework is a useful tool for making these assumptions transparent and for structuring the latent class model. The applied examples are engaging and the face validity checks (e.g., treatment patterns in the TB data) lend some support. However, the applied conclusions depend critically on structural-zero assumptions that are not clinically justified and are not subjected to sensitivity analysis, and the models are only locally identifiable with the help of informative priors. These limitations materially weaken the paper's central applied claims. If the structural assumptions were relaxed, the numerical findings could change substantially; the paper does not provide evidence of robustness.","major_comments":[{"comment":"The reduced DAG drops 'Other respiratory disease' and relabels the x-ray measurand as 'Intrathoracic abnormalities due to TB'. Table 1 then declares the latent class (Active TB=0, TB Infection=0, Intrathoracic abnormalities=1) as not possible (NP). In a pediatric cohort presenting with respiratory symptoms, this is a strong and clinically questionable restriction: chest x-ray abnormalities frequently occur without active TB (e.g., pneumonia). Because no observed test measures 'other respiratory disease', children with such abnormalities are forced into class 4 (0,0,0), directly inflating the estimated class-4 prevalence (0.52, 95% CrI 0.34–0.64) and biasing x-ray sensitivity with respect to its measurand (reported 0.95, 0.82–1). The manuscript provides no sensitivity analysis to this structural zero. The headline prevalence difference (0.22 vs 0.27) is therefore not robust to a clinicall","section":"§2.1, Table 1, Figure 1b"},{"comment":"The leptospirosis model suffers an analogous problem: the DAG in Figure 2b drops 'Other infection' because it is 'not estimable', and Table 3 restricts the model so that IgM/IgG arise only from Leptospirosis or from the target condition. The informative priors that the specificity of IgM exceeds 90% and that the specificity of IgG is 100% effectively impose the absence of other causes of these antibodies. In a hospital setting in Tanzania, this assumption is not clinically defensible. The resulting 6LC estimates—and the conclusion that the 2LC latent class corresponds to IgM rather than Leptospirosis—are conditional on this untestable exclusion. A sensitivity analysis that allows nonzero prevalence of other infections, even with weak prior information, is necessary before these applied conclusions can be considered reliable.","section":"§2.2, Table 3"},{"comment":"The authors report that the TB model's Jacobian rank is 12 with 14 unknown parameters, requiring at least 2 informative priors (Section 5.1). The structural-zero restrictions in Table 1 are additional, non-estimated constraints. This combination means the posterior estimates—including the active TB prevalence of 0.22 and the measurand-specific accuracies—depend on both the choice of which parameters receive informative priors and the validity of the structural restrictions. The manuscript does not report any sensitivity analysis varying either the prior specifications or the structural-zero set. Given that the local identifiability condition is not met, the applied estimates may be largely driven by prior and structural assumptions rather than by the data, and the paper should demonstrate stability of its conclusions under plausible perturbations.","section":"§5.1, Identifiability"}],"minor_comments":[{"comment":"The simulation uses the expected dataset rather than simulated random datasets, which is fine for illustrating mean behavior, but the sample size N is never stated. Please specify the sample size used to scale the expected frequencies, and consider reporting Monte Carlo standard errors or additional randomly simulated replicates to assess sampling variability.","section":"§4.1"},{"comment":"Table 2 is difficult to read: the column headers do not clearly indicate which columns are sensitivities versus specificities and with respect to which target/measurand. The reader has to reconstruct the structure from the prose. Please reorganize the table so that each test's accuracy with respect to each latent variable is presented unambiguously.","section":"Table 2"},{"comment":"There is a typo: 'thus illustrating they are are conditionally dependent' should read 'they are conditionally dependent'.","section":"§2.1"},{"comment":"In the leptospirosis results, the credible intervals for the prevalence of Leptospirosis under the 6LC model are extremely wide (0.28, 95% CrI 0.02–0.93; Table 4). Even the point estimates of sensitivity for several tests with respect to Leptospirosis have intervals spanning a large range. The conclusion that the 2LC model identifies IgM rather than Leptospirosis is partly based on point estimates with substantial uncertainty; the overlap of intervals should be acknowledged or handled more carefully in the interpretation.","section":"§5.2"},{"comment":"The Wasserman reference is incomplete ('2008-2010', 'Chapter 18') and should be updated to a proper citation. Also, 'Mycobaterium' is a typo in Section 2.1.","section":"References"},{"comment":"The text says '13 unknown parameters' and then later 'total number of unknown parameters was 19' after accounting for prevalence covariates and a continuous measurand. The derivation of the 19 total parameters (beyond the 13) is not shown and should be clarified, since the reader is referred to the supplementary material for details that are not available in the main text.","section":"§5.1"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a real and important problem in diagnostic test evaluation, and the simulation clearly supports the claim that a 2LC model can identify a dominant non-target measurand. The main issue is that the applied examples—which are central to demonstrating the value of the proposal—rely on strong structural-zero assumptions that are not clinically justified and for which no sensitivity analysis is provided. The identifiability limitations add to this concern. I would be willing to reconsider if the authors add (a) a sensitivity analysis for the TB and leptospirosis examples that relaxes the structural zeros (e.g., allowing a non-zero probability for the class with intrathoracic abnormalities but no active TB, or allowing other-sourced IgM/IgG), and (b) a more explicit discussion of how robust the numerical claims are to the choice of informative priors. The methodological core is sound and the paper would be a useful contribution to Biometrics if these concerns are addressed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague—\n\nThe thing to know: this paper has a real idea and demonstrates it cleanly in simulation, but the applied examples rest on structural assumptions that are more fragile than the text lets on. Read the simulation and the DAG framework; treat the TB and leptospirosis estimates as illustrative rather than definitive.\n\nWhat's actually new: framing the measurand of each test as a latent variable distinct from the target condition, using DAGs to do this transparently, and expanding the latent-class model accordingly. The simulation is the strongest part. When four of five tests measure a non-target measurand M, the 2LC model identifies M, not the target condition, and reports test sensitivities as if they were with respect to the target. That is a concrete demonstration of a real failure mode, and it should change how people specify and interpret LCMs. The likelihood in (1) is a correct mixture expression under the stated assumptions, and the notation makes the distinction between accuracy with respect to the measurand and with respect to the target condition usable.\n\nThe soft spots are real but not fatal. The stress-test note lands. In the TB application, the reduced DAG drops 'other respiratory disease' and then Table 1 declares the class (no active TB, no latent TB infection, x-ray abnormalities present) impossible. Clinically, x-ray abnormalities in symptomatic children can come from pneumonia or other respiratory infections. That structural zero is doing a lot of work: it inflates the no-abnormality class and shapes the x-ray sensitivity estimate. There is no sensitivity analysis to this assumption, and the same issue appears in the leptospirosis example where 'other infection' is dropped, leaving IgM and IgG arising only from leptospirosis. The leptospirosis estimates are also so wide that they don't support much beyond the qualitative point.\n\nThere is also an internal inconsistency in Section 5.1: first 13 unknown parameters, then 19 after adding a continuous measurand, then the Jacobian rank is 12 versus 14 parameters. That needs fixing before I'd trust the identifiability claims. And 'results not shown' for the unbiased 4LC simulation is an unkept promise; the code/data are absent. None of this sinks the central message, but it does cap confidence in the applied numbers.\n\nWho this is for: anyone doing latent class analysis of diagnostic tests, and reviewers of such work. It deserves a serious referee, if the authors clean up the identifiability reporting, add sensitivity analyses for the structural zeros, and share the simulation code.","headline":"A genuinely useful framing for latent class analysis—DAG-driven measurands—with a solid simulation, but the applied examples and identifiability reporting need cleanup before the estimates can be trusted.","tokens_in":12313,"tokens_out":2605,"would_cite":true,"duration_ms":23737,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62P10","62F15","62H30"],"pacs":[],"model":"deepseek-v4-flash","headline":"A standard two-latent-class model for diagnostic tests can identify the wrong latent class — the majority measurand — when most tests do not measure the target condition.","keywords":["latent class analysis","diagnostic test accuracy","measurand","directed acyclic graph","conditional dependence","identifiability","Bayesian inference","prevalence estimation"],"falsifier":"If, in a dataset where the majority of tests are known (by design or external gold standard) to measure something other than the target condition, the 2LC model's estimated latent class prevalence still equals the target-condition prevalence (and test sensitivities match the target-condition values) rather than the measurand's prevalence, the paper's central claim would be false. Concretely: simulate with five tests where four detect a proxy M and one detects the disease D, with D and M only weakly correlated; the 2LC model should return M prevalence as its class prevalence. Any simulation tha","tokens_in":11377,"feed_emoji":"🩺","tokens_out":6050,"duration_ms":48226,"temperature":0.7,"pith_summary":"The paper argues that in diagnostic accuracy studies without a gold standard, the conventional two-latent-class (2LC) model silently assumes every test measures the same target condition. In reality each test has its own measurand — the biological quantity it actually detects — which may be only related to the disease. The authors show that when most observed tests share a non-target measurand, the 2LC model's latent classes are labeled by that measurand rather than by the disease, biasing prevalence and sensitivity/specificity estimates. They propose drawing a DAG to expose each test's measurand, expanding the model to the resulting latent classes, and fitting that expanded model with Bayesian inference. Re-analysis of pediatric tuberculosis and leptospirosis data illustrates that the expanded model gives clinically more plausible class labels and different accuracy estimates than the original 2LC analyses.","feed_headline":"Latent classes track what most tests measure, not the disease","feed_subtitle":"If most tests detect a proxy marker, prevalence and accuracy estimates describe the marker, not the disease.","key_machinery":"The central device is the distinction between a test's measurand (the quantity it is designed to measure, such as IgM antibodies or M. tuberculosis bacteria) and the target condition of interest (the disease one wants to diagnose). The paper represents both, plus observed tests and covariates, in a directed acyclic graph (DAG). The DAG reveals shared and nonshared measurands, which implies conditional dependence among tests with the same measurand and determines how many latent classes exist (combinations of target condition and measurands). The expanded likelihood then factors each observed test's result through its measurand: P(T_j | M_p) times P(M_p | D), so test accuracy and measurand ac","core_discovery":"On the paper's own terms: a conventional two-latent-class model does not always identify the target condition. The latent classes it recovers are determined by the combination of observed tests; if the majority of tests measure a measurand distinct from the target condition, the classes are effectively M+ and M- (measurand positive/negative), not D+ and D- (target condition positive/negative). This happens because the tests' shared measurand induces conditional dependence that the 2LC model can only absorb by relabeling the latent variable. When the model is expanded so that latent classes are all combinations of the target condition and the measurands (as dictated by a DAG), the likelihood","pith_inferences":["The same mislabeling hazard applies to any finite mixture model applied to measurements that tap different constructs; the paper's DAG-plus-measurand recipe is a general diagnostic for mixture label validity, not just for medical tests.","A testable extension: when the majority of tests share a proxy measurand, the 2LC model's estimated prevalence should track the measurand's prevalence even if the target-condition prevalence is known from external sources; this can be checked in datasets with verified disease status.","The structural constraints (which latent class combinations are 'impossible') are assumptions; when those are wrong, the expanded model can be as mislabeled as the 2LC model. The paper does not sensitivity-analyze those constraints, but a robustness check that relaxes one constraint at a time would quantify the risk."],"forward_implications":["In any 2LC analysis, the latent class label is not guaranteed to be the target condition; it is an emergent property of which measurand dominates the test set.","When a majority of tests share a non-target measurand, prevalence and test-accuracy estimates from the 2LC model describe the measurand, not the disease.","Expanding the model to include measurand-based latent classes restores interpretable labels and separates test accuracy with respect to the measurand from accuracy with respect to the target condition.","The expanded models may fail local identifiability, but informative expert priors on one or two specificities can make them estimable.","Re-analysis of the leptospirosis data shows the previously 'low sensitivity' MAT estimate was an artifact: the 2LC classes were IgM+, not Leptospirosis+."],"fun_headline_variants":["Latent classes may track test measurands, not the disease","DAGs reveal when latent class models mislabel the target condition","Misaligned tests skew prevalence—DAGs correct the latent class view","When tests measure different things, latent classes can mislead"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The whole approach rests on the DAG being right: the chosen measurands, the dropped nodes, and the 'impossible' latent class combinations must genuinely reflect how the tests work; if any of these structural assumptions is wrong, the expanded model's class labels and accuracy estimates are just as unreliable as the 2LC model's.","fun_headline_variants_meta":{"raw":{"variants":["Latent classes may track test measurands, not the disease","DAGs reveal when latent class models mislabel the target condition","Misaligned tests skew prevalence—DAGs correct the latent class view","When tests measure different things, latent classes can mislead"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000327,"raw_usage":{"total_tokens":1685,"prompt_tokens":784,"completion_tokens":901,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":528,"completion_tokens_details":{"reasoning_tokens":826}},"tokens_in":528,"tokens_out":901,"duration_ms":8057,"temperature":1.0,"reasoning_tokens":826,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T01:59:10.691307+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"If, in a dataset where the majority of tests are known (by design or external gold standard) to measure something other than the target condition, the 2LC model's estimated latent class prevalence still equals the target-condition prevalence (and test sensitivities match the target-condition values) rather than the measurand's prevalence, the paper's central claim would be false. Concretely: simulate with five tests where four detect a proxy M and one detects the disease D, with D and M only weakly correlated; the 2LC model should return M prevalence as its class prevalence. Any simulation tha","supporting_citations":[],"review_version":1}