{"id":"52ed2a1f-5743-426d-b625-f40ce48970d0","arxiv_id":"2505.23447","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"New quality metrics quantify amount, joint, and conditional missingness, and guide visual exploration of non-random missing value structures in high-dimensional data.","lead":"This paper defines simple numerical scores that describe how and where data values are missing, such as how much is missing per column and whether missingness in one column is linked to values in another. It shows how plotting these scores can help analysts spot non-random gaps in large health datasets, demonstrated on a Parkinson's walking study.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"QCM_DiD/QCM_H are unstable when the conditioning subset S=DRk∩DMj is small; the QJM_dir<0.05 filter does not guarantee large S, so the case-study CM findings may reflect finite-sample noise rather than structure.","rationale":"The reader's weakest assumption is that the conditional-missingness metrics require DRk∩DMj to be large enough, and that Section 5.4's filter is an unvalidated mitigation. I agree, and would sharpen the analysis: the problem is not only high joint missingness, but any configuration in which the conditioning subset is small. In particular, low QJM_dir does not imply large S, because S ≈ P(dj missing)(1 − P(dk missing)) under near-independence. The synthetic data in Section 4 use N=106 and deliberately construct CM patterns with substantial S, so they do not stress this regime; the real case study does. The paper is transparent about the limitation, and the metric definitions are internally consistent, so this is not a fatal flaw. But the central claim that the QMs identify structural missingness remains conditional on establishing that observed QCM values exceed what sample-size variation alone would produce. A permutation null and sensitivity analysis of the thresholds would settle this directly and would not require abandoning the approach. Hence the existing CONDITIONAL verdict is appropriate; I do not move it.","tokens_in":16973,"tokens_out":4672,"duration_ms":49800,"concrete_test":"Run a permutation test under the independence null for every variable pair in the ICICLE analysis: fix the observed missingness masks, permute the missingness indicator of dj across rows, and recompute QCM_DiD and QCM_H 1,000 times, stratifying by |S|. Then compare the observed values for the pairs flagged by QJM_dir < 0.05 and QCM_DiD > 0.9 against this null distribution. If those pairs are not in the upper tail after FDR correction, or if the smallest |S| among flagged pairs is below 5–10% of N, the case-study CM findings are not distinguishable from small-sample artifacts; if they are significant and S is reasonably large, the concern is resolved. Reporting the null quantiles for a few representative |S| levels would also settle whether the proposed thresholds are appropriate.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that equations 5 and 6 identify meaningful conditional missingness structures. Both metrics compare the distribution of all recorded values in dk with the distribution in dk restricted to S = DRk∩DMj. The second histogram is estimated from |S| items only. When |S| is small, QCM_DiD is the total variation distance between a population distribution and a small-sample estimate; under MCAR its expected value is positive and grows as |S| shrinks, and QCM_H is inflated because sample entropy is biased downward in small samples. Section 5.4 acknowledges this for high joint missingness, but the proposed mitigation, filtering on QJM_dir < 0.05, does not control |S|. Writing P(both) for P(dj missing and dk missing), S has size P(dj missing) − P(both); under low QJM_dir, P(both) ≈ P(dj missing)P(dk missing), so S ≈ P(dj missing)(1 − P(dk missing)). Thus S is small whenever missingness in dj is low or missingness in dk is high, even with very low QJM_dir. The thresholds 0.05 and 0.9 are used once, with no sensitivity analysis or null distribution, and the case-study conclusions about AxLb, BMI, Height, and MoCA depend on the surviving edges. Without a sampling distribution or uncertainty quantification, the method cannot distinguish true conditional missingness from finite-sample fluctuation in this common regime.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper defines six quality metrics for missingness patterns—QAM, QJM_mag, QJM_dir, QJM_abs, QCM_DiD, and QCM_H—and demonstrates their use in variable ordering, filtering, and network, parallel-coordinates, and MissiG visualizations. Section 3 gives the definitions, Section 4 presents synthetically generated datasets with controlled missingness structures, and Section 5 applies the metrics to the ICICLE walking-monitoring dataset, reporting structures that are traced to data-collection procedures. The paper also acknowledges in Section 5.4 that the conditional-missingness metrics are problematic when the conditioning subset is small. The central claim is that these metrics can guide visual exploration of structured missingness in large high-dimensional data.","tokens_in":17234,"tokens_out":3767,"duration_ms":37687,"significance":"The contribution is potentially valuable: quality metrics tailored to missingness structures are a genuine gap in the visualization literature, and the proposed metrics are clearly specified and correctly normalized, including the division by 2 in Eq. (5) and by log(bk) in Eq. (6). The visual encodings are sensible, and the ICICLE case study is a rich real-world exemplar. The supplemental materials with synthetic datasets are a useful asset. However, the validation is not yet strong enough to support the effectiveness claim: the synthetic experiments largely confirm quantities that were planted using the same definitions, and the conditional-missingness metrics lack uncertainty quantification or a null-model baseline.","major_comments":[{"comment":"The synthetic evaluation plants missingness using exactly the quantities the metrics measure: JM patterns are generated by setting P(dj,dk) relative to E(dj,dk), which is what QJM_dir and QJM_abs compute by definition, and CM patterns are generated by placing missing values in predetermined value ranges, which is what QCM_DiD and QCM_H detect. The agreement reported in Sections 4.3 and 4.4 therefore confirms internal consistency, not detection power. Without a null model (e.g., MCAR replicates) or a comparison to standard missingness tests, the experiments do not establish that the metrics can distinguish structured missingness from chance, which is the central claim of Section 3. Please add a randomized-baseline evaluation reporting detection rates, precision/recall, or an equivalent performance measure.","section":"§4.1, Tables 1–2; §4.3–§4.4"},{"comment":"The CM metrics estimate a histogram and an entropy from the subset DRk∩DMj. When this subset is small, QCM_DiD is positively biased and QCM_H is inflated by the negative small-sample bias of sample entropy, so high values do not necessarily imply conditional missingness. Section 5.4 correctly acknowledges this for high joint missingness, but the proposed filter QJM_dir < 0.05 does not control the subset size: writing P(both) for P(DMj∩DMk), the conditioning subset has size P(DMj) − P(both), which under near-independence is approximately P(DMj)(1−P(DMk)) and can be tiny even when QJM_dir is near zero. The thresholds 0.05 and 0.9 are applied once, with no sensitivity analysis or null distribution, and the case-study conclusions about Height, BMI, and MoCA rest on the surviving edges. Please provide a sampling distribution or an explicit minimum-support criterion for Eqs. (5)–(6), and validate the filter on synthetic data with known small-S regimes.","section":"§5.4; Eqs. (5)–(6)"}],"minor_comments":[{"comment":"The dataset is described as containing 106 patients, but 64 cancer patients plus 52 healthy controls sums to 116; please correct this inconsistency.","section":"§4.1"},{"comment":"The last row reports '38.3.1%' where a valid percentage is intended; please fix this typographical error.","section":"Table 1"},{"comment":"The phrase 'aN aNstrings' appears throughout; it should read 'NaN strings'.","section":"§4.1"},{"comment":"There are several typographical slips, including 'furhtermore' and 'zoomed in in figure 14'; these should be cleaned up before publication.","section":"§5.4 and figure captions"},{"comment":"The text describes the filter QJM_dir < 0.05 as indicating low joint missingness, but QJM_dir is a signed deviation from expected joint missingness, not a magnitude; the wording should clarify that the filter is selecting a mix of negative and small positive deviations.","section":"§5.4"}],"recommendation":"major_revision","confidential_remarks":"The paper fits the journal's scope and the supplementary data are a plus. My main concern is that the evaluation validates the metrics against the same constructions used to define them, and the conditional-missingness findings in the case study depend on an unvalidated filter. I would be comfortable with a major revision that adds an MCAR null-model comparison and an uncertainty or support-size analysis for the CM metrics."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Honest take: this is a useful formalization paper, not a breakthrough. The six quality metrics (equations 1–6) are clearly defined and correctly normalized, and they give missing-data visualization a simple ordering/ranking layer that was missing. The authors build on Fernstad's earlier taxonomy and add the actual equations, which is a legitimate contribution. The synthetic data and the ICICLE Parkinson's case study are both included, and the supplemental materials are available, which I appreciate.\n\nThe weak spots are both in the evaluation. First, the synthetic validation is circular in a real sense: they plant missingness patterns using the same quantities the metrics measure, so the metrics trivially recover them. That's a sanity check, not proof of usefulness. Second, and more consequential for the case study, the conditional-missingness metrics QCM_DiD and QCM_H compare a full distribution against a small subset S = DRk∩DMj. Section 5.4 acknowledges the small-sample problem, but the proposed fix—filtering on QJM_dir < 0.05 and QCM_DiD > 0.9—doesn't control |S|. Your stress test is right: with low QJM_dir, |S| ≈ P(dj missing)(1 − P(dk missing)), which can be tiny. The thresholds look chosen once, with no sensitivity analysis or null model. So the case-study claims about BMI, Height, and MoCA are weaker than the authors present. They are honest about the limitation, and they call the findings speculative, but the method doesn't yet have a way to distinguish real conditional structure from finite-sample noise.\n\nThe metrics themselves are not deep statistics—marginal proportion, independence deviations, total variation distance, entropy difference—but that's not a flaw; packaging them for visual guidance is the contribution. People working on missing-data visualization or data-quality triage in high-dimensional health data will get value from this. I'd like to see the authors add a sampling distribution or permutation-based null for QCM, and test the filter more systematically. I'd send this to peer review rather than desk reject: the formalization is clean, the visualization examples are instructive, and the real dataset is a plus. With a sensitivity analysis and a more careful treatment of the small-S case, it could be a solid reference for missing-data visualization work.","headline":"Useful formalization of missingness metrics, but the conditional-missingness evaluation needs a sensitivity analysis before the case-study findings can be trusted.","tokens_in":17812,"tokens_out":3606,"would_cite":true,"duration_ms":33410,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Six quality metrics expose the hidden structure of missing data.","keywords":["missing data","quality metrics","missingness structures","high-dimensional visualization","joint missingness","conditional missingness","parallel coordinates","mobility monitoring"],"falsifier":"Take a dataset whose missing values are generated completely at random per variable (MCAR), with high per-variable missing rates so that many pairs have large joint missingness. If $QCM_{DiD}$ and $QCM_H$ frequently exceed their upper thresholds for these purely random pairs, the metrics cannot separate random from structured conditional missingness; the paper's own Section 5.4 observation predicts exactly this behavior, so an explicit receiver-operating-characteristic curve over synthetic MCAR data would settle how much of the case-study signal is artifact.","tokens_in":16740,"feed_emoji":"🔍","tokens_out":5781,"duration_ms":58825,"temperature":0.7,"pith_summary":"The paper proposes six quality metrics that reduce the messy, hard-to-see problem of missing values in high-dimensional data to a small set of scores an analyst can sort, filter, and map to visual channels. Three of the metrics quantify how much is missing per variable, how much missingness is jointly shared between pairs, and how far that joint missingness deviates from what chance would produce. Three more quantify conditional missingness: whether items that are missing in one variable have unusually low, high, or concentrated recorded values in another variable. The authors' central claim is that these scores guide visual exploration well enough to reveal non-random, structural missingness, which they demonstrate on controlled synthetic data and on a six-year Parkinson's walking-monitoring dataset with 56% missing values. If the claim holds, analysts can use the metrics to find data-collection artifacts and form hypotheses about why values are absent rather than treating missingness as noise to impute away.","feed_headline":"Six quality metrics expose the hidden structure of missing data","feed_subtitle":"Per-variable and pairwise scores turn scattered blanks into patterns analysts can see and act on.","key_machinery":"The load-bearing machinery is a set of six per-variable or per-pair scores, equations (1) through (6). $QAM$ is the fraction of missing entries in a variable. $QJM_{mag}$ is the fraction of items jointly missing in both variables; $QJM_{dir} = P(\\vec d_j,\\vec d_k) - E(\\vec d_j,\\vec d_k)$ is the signed deviation from the chance expectation $E = P(\\vec d_j) P(\\vec d_k)$; and $QJM_{abs}$ is the absolute value of that deviation, so high scores flag pairs whose co-missingness is unlikely to be accidental. The conditional metrics $QCM_{DiD}$ and $QCM_H$ compare, for each direction, the histogram of recorded values in $\\vec d_k$ for items missing in $\\vec d_j$ against the histogram for all recorded items, using the Shimazaki-Shinomoto rule to set bin counts, and normalize by the maximum possible difference or entropy. These scores make high dimensionality tractable: instead of inspecting all pairs of hundreds of variables, the analyst sorts, filters, and lays out variables by the scores and inspects the small set of outlier pairs.","core_discovery":"The paper's central claim is that structural missingness in large, high-dimensional data can be surfaced by six quality metrics defined over the three patterns of Amount Missing, Joint Missingness, and Conditional Missingness. $QAM$ gives each variable a score between 0 and 1 for the relative share of missing values. For each variable pair, $QJM_{mag}$ measures the observed share of jointly missing items, $QJM_{dir}$ measures the signed difference between observed and expected joint missingness, and $QJM_{abs}$ measures the absolute size of that deviation. For each directed pair, $QCM_{DiD}$ compares the distribution of recorded values for items that are missing in the other variable against the overall distribution, while $QCM_H$ compares their Shannon entropies. The paper demonstrates that mapping these values to variable ordering, subset selection, node size, edge width and colour, and histogram/glyph displays lets users spot clusters of attributes whose missingness is too structured to be random. In the ICICLE case study, the metrics expose blocks of missingness tied to study protocol, such as clinical measures not recorded for healthy controls and mutually exclusive activity-measurement sessions, and point to a testable hypothesis that declining health caused later dropouts.","pith_inferences":["The deviation-from-chance framing of $QJM_{dir}$ could be reused for any pairwise co-absence problem, such as co-occurring sensor failures or mutually exclusive survey response patterns; the paper does not state this application.","The known instability of $QCM_{DiD}$ and $QCM_H$ when the conditioning subset is tiny suggests a natural extension: report a minimum-support count or confidence interval alongside each conditional score, and calibrate thresholds per dataset rather than using fixed cutoffs.","A null-model test that permutes missingness within each variable while preserving marginal missing rates could turn these metrics into formal significance tests of non-random missingness; that test is not in the paper.","If applied at the item level rather than the variable level, the metrics could flag individual records whose missingness profile is anomalous, a direction the paper does not explore."],"forward_implications":["Ordering variables by $QAM$ immediately separates attributes whose missingness would make simple imputation unreliable from those that are nearly complete.","Pairs with high $QJM_{abs}$ and clearly positive or negative $QJM_{dir}$ can be prioritized for data-collection debugging, since their co-missingness cannot be explained by per-variable missing rates alone.","High $QCM$ values in a direction $\\vec d_j \\to \\vec d_k$ give an evidence-based reason to impute $\\vec d_j$ using recorded values in $\\vec d_k$, or to withhold imputation and model the missingness explicitly.","The same metrics can be applied to categorical variables by treating categories as histogram bins, extending the method beyond numerical sensor data.","In longitudinal studies, trends in $QAM$ and $QJM_{abs}$ across time points can track worsening cohort health or protocol drift as the study ages."],"supporting_citations":[{"why":"Defines the three missingness patterns (Amount Missing, Joint Missingness, Conditional Missingness) that the paper's metrics operationalize.","marker":"[16]"},{"why":"Provides the MissiG glyph representation and highlights high-dimensional missingness as a scalability challenge.","marker":"[11]"},{"why":"Supplies the pairwise-ordering algorithm used to arrange variables by QJM_abs and QCM_H in the visual encodings.","marker":"[4]"},{"why":"Gives the histogram bin-selection rule that fixes how conditional distributions are compared.","marker":"[32]"},{"why":"Documents the gap in scalable missing-data visualization that this paper addresses.","marker":"[2]"},{"why":"Supplies the no-missing-values dataset modified to create controlled AM, JM, and CM test cases.","marker":"[27]"},{"why":"Provides the ICICLE cohort data source used in the case study.","marker":"[26]"},{"why":"Provides the ICICLE-PD study cohort that generated the walking-monitoring dataset.","marker":"[42]"},{"why":"Defines the MCAR/MAR/MNAR missingness mechanism framing that motivates the deviation-from-expected metrics.","marker":"[28]"}],"fun_headline_variants":["Six metrics expose hidden missingness patterns","Missing data structures seen through six metrics","Quality metrics reveal missingness patterns visually","Map missingness with six quality metrics","Visualize structured missingness with six metrics"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the two conditional-missingness metrics give dependable signals even when only a handful of items are missing in one variable but recorded in the other; the paper itself notes in Section 5.4 that with high joint missingness these scores are forced high and uses an ad-hoc filter to suppress that effect.","fun_headline_variants_meta":{"raw":{"variants":["Six metrics expose hidden missingness patterns","Missing data structures seen through six metrics","Quality metrics reveal missingness patterns visually","Map missingness with six quality metrics","Visualize structured missingness with six metrics"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000162,"raw_usage":{"total_tokens":1282,"prompt_tokens":1029,"completion_tokens":253,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":645,"completion_tokens_details":{"reasoning_tokens":191}},"tokens_in":645,"tokens_out":253,"duration_ms":3530,"temperature":1.0,"reasoning_tokens":191,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T12:45:52.299289+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a dataset whose missing values are generated completely at random per variable (MCAR), with high per-variable missing rates so that many pairs have large joint missingness. If $QCM_{DiD}$ and $QCM_H$ frequently exceed their upper thresholds for these purely random pairs, the metrics cannot separate random from structured conditional missingness; the paper's own Section 5.4 observation predicts exactly this behavior, so an explicit receiver-operating-characteristic curve over synthetic MCAR data would settle how much of the case-study signal is artifact.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the ICICLE cohort data source used in the case study."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the ICICLE-PD study cohort that generated the walking-monitoring dataset."},{"cited_title":"Johansson Fernstad","cited_arxiv_id":null,"evidence_quote":"Defines the three missingness patterns (Amount Missing, Joint Missingness, Conditional Missingness) that the paper's metrics operationalize."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the MissiG glyph representation and highlights high-dimensional missingness as a scalability challenge."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the pairwise-ordering algorithm used to arrange variables by QJM_abs and QCM_H in the visual encodings."},{"cited_title":"Shimazaki and S","cited_arxiv_id":null,"evidence_quote":"Gives the histogram bin-selection rule that fixes how conditional distributions are compared."},{"cited_title":"Patr ´ ıcio, J","cited_arxiv_id":null,"evidence_quote":"Supplies the no-missing-values dataset modified to create controlled AM, JM, and CM test cases."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the MCAR/MAR/MNAR missingness mechanism framing that motivates the deviation-from-expected metrics."}],"review_version":1}