{"id":"4788a65e-4dd1-492b-8229-2df8a30f4d5f","arxiv_id":"2506.04429","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A ranking-based, human-in-the-loop monitoring system deployed in public health data review was associated with faster and more numerous event detection than prior baselines, but the headline 54x speedup is measured against a manual baseline, not an alerting system.","lead":"This paper describes a deployed public health data monitoring system that ranks anomalous data points by severity instead of raising alerts, and reports that reviewers found and documented events much faster and in greater numbers than with the previous setup. It is a real-world case study of human-AI collaboration in a high-volume, noisy data environment.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 54x speedup headline likely mischaracterizes the comparison baseline: Sec. 8 measures against the manual exploratory fallback (Baseline 1), not the alerting system named in the abstract, and the paper's own numbers are inconsistent (5x vs 6x).","rationale":"The paper reports a real deployment and a plausible design story, and the interaction modalities and design guidelines are useful contributions. The most load-bearing quantitative claim, however, is the 54x speedup. That claim is not supported as stated: the comparison baseline in Sec. 8 is the manual exploratory fallback of Baseline 1, not the alerting system named in the abstract. The paper even labels Baseline 1 as the 'Deployed Alerting System' but then describes reviewers abandoning it for manual inspection, so the reported speedup conflates the ranking paradigm with the fact that the baseline reviewers were not using an alerting system at all. The efficiency metric is also ambiguous: Fig. 6B plots recorded events per day, yet the text converts this to 'faster' without demonstrating that daily review time was constant; Fig. 6A shows time per row increasing across modifications, which complicates the speed interpretation. The 5x versus 6x discrepancy adds a concrete internal inconsistency. These issues warrant the reader's conditional verdict: accept only if the preregistration and raw logs are released, the baseline is described correctly, and the efficiency claim is tempered or verified with a proper denominator.","tokens_in":11666,"tokens_out":2940,"duration_ms":30217,"concrete_test":"Retrieve the OSF preregistration and artifacts, then recompute the efficiency multipliers from the raw reviewer logs with both numerator and denominator defined as events per minute of active review time. Next, rerun the speedup calculation substituting the original threshold-alerting interface — the system that generated tens of thousands of alerts — for Baseline 1. If the ratio drops materially below the claimed 54x, or cannot be computed from the logs, the abstract should be revised to state that the comparison is against manual exploratory review rather than traditional alert-based methods. Also resolve whether the Baseline 2 multiplier is 5x or 6x.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim depends on the efficiency ratio '54x faster' claimed in the abstract and Sec. 8. The paper defines Baseline 1 as 'Deployed Alerting System (120 weeks)' in Sec. 4.1, but then states that, without a reliable automated system, reviewers 'reverted to manual data inspection using exploratory tools.' Sec. 8 explicitly attributes the speedup to 'the exploratory system in Baseline 1.' Thus the denominator is a manual fallback, not the alert-based monitoring named in the abstract. The paper provides no raw event counts, session durations, or per-reviewer logs; the metric in Sec. 8 is 'recorded events per day' (Fig. 6B), which is not a speed measure unless total review time is held constant, and Fig. 6A shows time per row increased with modifications. The internal 5x (Sec. 8) vs 6x (Sec. 9) discrepancy for Baseline 2 further undermines confidence that the reported multipliers were computed from a stable, preregistered metric. If the actual alerting baseline produced 35,000+ alerts (Sec. 4.1), the efficiency comparison against that system would have a different denominator.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper describes a deployed public-health data monitoring system that replaces threshold-based alerting with an AI-based ranking paradigm, integrated into a human-in-the-loop interface for data reviewers at the Delphi group. The authors report a multi-year co-design process with reviewers, engineers, and computer scientists; a three-month longitudinal deployment; and a headline claim of a 54x increase in reviewer speed efficiency compared with traditional alert-based methods. They also report increased event and meta-event recording, positive reviewer feedback, and design guidelines for human-centered AI monitoring systems.","tokens_in":11870,"tokens_out":3467,"duration_ms":33633,"significance":"If the core claims hold, the paper would be a valuable real-world case study of a human-centered AI monitoring system deployed at national scale, with an unusual longitudinal evaluation and a concrete alternative to alert-fatigue-prone surveillance. The authors give appropriate emphasis to stakeholder-engaged design and to the practical failure modes of alerting systems. However, the quantitative headline claim is not currently supported by the evidence as presented, and the evaluation lacks the raw data and preregistration needed to verify the reported multipliers. The qualitative findings and design lessons are nonetheless likely to be useful to the HCI and public-health informatics communities.","major_comments":[{"comment":"The abstract claims \"54x increase in reviewer speed efficiency compared to traditional alert-based methods,\" but Sec. 8 measures the 54x against \"the exploratory system in Baseline 1.\" Sec. 4.1 defines Baseline 1 as the deployed alerting system, yet immediately states that \"reviewers reverted to manual data inspection using exploratory tools.\" Thus the comparison denominator is the manual fallback, not alert-based monitoring, and no quantitative alerting-phase data enter the ratio. The headline should be revised to compare like with like, or actual alerting-phase speed should be measured and reported.","section":"Abstract; Sec. 8; Sec. 4.1"},{"comment":"The efficiency claim is based on \"recorded events per day\" (Fig. 6B), which is an event count, not a speed, unless total review time is held constant. Fig. 6A shows that time spent per data row increased with each modification, so the reported 54x speedup cannot be separated from time investment without session-level duration and workload data. Please report raw event counts, session durations, per-reviewer logs, and the exact rate formula used to compute the multiplier.","section":"Sec. 8; Fig. 6"},{"comment":"The paper states that reviewers were 5x faster than Baseline 2 in Sec. 8 and 6x faster in Sec. 9, with no reconciliation. The paper also states that the evaluation strategy was preregistered on OSF, but no OSF link, preregistration document, or data release is provided. These issues, together with the small uncontrolled sequential design (no control condition; reviewer learning over three months is a plausible confound), limit confidence in the quantitative multipliers. Please reconcile the discrepancy and provide the preregistration and supporting data, or explicitly restrict claims to qualitative trends.","section":"Sec. 8; Sec. 9"}],"minor_comments":[{"comment":"The manuscript contains typos, including \"effeciency\" (Abstract and Sec. 8), \"appporpriate\" (Sec. 4.3), \"Reviwer Actions\" (Sec. 1), and \"Perscriptive Alerting Fallacy\" (Sec. 4.2). Please proofread the text carefully.","section":"Throughout"},{"comment":"The paper says \"We preregistered the Github commits and evaluation strategy on OSF before experiments began\" but provides no link or identifier. Please add the OSF record, or remove the preregistration claim if the record is not public.","section":"Sec. 1 and Sec. 8"},{"comment":"The caption says \"Baselines had no comparable metrics,\" which is contradicted by the numeric 54x and 5x claims in the text. Please clarify what baseline metrics were and were not available.","section":"Fig. 6 caption"},{"comment":"The label \"Baseline 1: Deployed Alerting System (120 weeks)\" is misleading because the description immediately shifts to a manual exploratory fallback. Consider renaming or restructuring this subsection to distinguish the alerting phase from the manual-inspection phase.","section":"Sec. 4.1"}],"recommendation":"major_revision","confidential_remarks":"The absence of the promised OSF record and raw data is a reproducibility concern; the editor may wish to ask the authors to provide these before any acceptance decision. The internal 5x vs 6x discrepancy and the abstract-vs-body mismatch for the 54x claim also suggest that the quantitative claims need careful re-analysis."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe real news here is that a deployed AI-based monitoring system at a national public health organization is running daily on up to 5M data points, and the team has documented how they moved from an alerting paradigm to ranking anomalies and what that took. That is genuinely scarce. The interaction modalities—simple filtering, situational awareness panels, the rolling mean/variance display—are concrete, and the design guidelines (meaningful interaction, situational awareness, engagement) are sensible. The qualitative account of reviewer fatigue with alerts and the shift to ranked lists is credible.\n\nThe quantitative headline, 54x faster, is the problem. The abstract says compared to 'traditional alert-based methods,' but Section 8 measures against Baseline 1, which is the manual exploratory fallback reviewers used when the alerting system failed. Those are not the same denominator. The paper also gives 5x in Section 8 and 6x in the Conclusion for the Baseline 2 comparison, which suggests the metric was not stable. The evaluation is small (three reviewers), sequential, and uncontrolled, so learning effects and motivation are not separated from interface effects. That is not disqualifying for a deployed system, but it means the paper should present the efficiency claim as 'faster than the manual fallback' and temper the absolute multiplier.\n\nThere is also a reproducibility gap: the paper says the evaluation was preregistered on OSF, but no link or artifact is given, and no raw logs are included. For a claim this specific, that is a fixable but real problem.\n\nWhat holds up: the core ranking algorithm is prior work, appropriately cited, and the new contribution is the deployment and interaction design. The failure-mode analysis of alerting is sound. The meta-event concept is useful.\n\nThis paper deserves a serious referee, but with a clear mandate for revision: release the preregistration and data (or explain why they cannot), correct the baseline description, reconcile the 5x/6x discrepancy, and soften the headline to match the actual comparison. If those are done, it is a solid contribution to human-centered AI and public health informatics. I would cite it for the deployment story, not for the efficiency multiplier.","headline":"A real deployed system and a good design story, but the 54x efficiency claim is measured against the wrong baseline and needs a major revision before the numbers can be used.","tokens_in":12424,"tokens_out":1926,"would_cite":true,"duration_ms":19392,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that replacing threshold alerts with a ranked anomaly list, plus context displays, lets a small team monitor millions of health data points daily and detect more events, including a 54x reviewer speedup in a deployed…","keywords":["public health monitoring","anomaly ranking","alert threshold","human-centered AI","outbreak detection","data quality","longitudinal evaluation","situational awareness"],"falsifier":"A crossover deployment in which the same reviewers alternate between the ranked-list interface and a threshold-alert interface that emits only a manageable number of top alerts per day would settle the claim: if the speed advantage over that alert interface is small or zero, the 54x result is specific to the exploratory baseline, not to alerting as such.","tokens_in":11443,"feed_emoji":"🩺","tokens_out":9823,"duration_ms":80837,"temperature":0.7,"pith_summary":"Public health data monitoring traditionally works by setting thresholds that fire alerts, but the authors argue this paradigm cannot scale to modern high-volume, noisy, nonstationary data: thresholds need constant manual retuning, alert volumes become unusable, and reviewers lose situational awareness. The paper's proposed fix is a ranking-based monitoring paradigm: an AI method scores each data point by its deviation from expectation, and reviewers work down a ranked list instead of wading through binary alerts. In a three-month pre-registered longitudinal deployment at a national organization, the system monitored up to 5,000,000 data points daily, and reviewers recorded events 54x faster on average than with the prior manual-exploratory workflow (described in the abstract as a 54x speedup over alert-based methods) while finding more events and higher-level 'meta-events.' If the result holds, it gives public health agencies and other high-volume monitoring domains a concrete model for combining AI anomaly scores with human triage.","feed_headline":"Ranked AI lists let reviewers spot public health events 54x faster","feed_subtitle":"A deployed system triages millions of daily data points, finding more events and meta-events than alert thresholds.","key_machinery":"The central object is the ranking-based monitoring paradigm: instead of emitting a binary alert when a threshold is crossed, an anomaly-detection method assigns every data point a score measuring deviation from expectation, and the system presents the top-scoring streams as a ranked list. The underlying method is an unsupervised multiple-univariate outlier-ranking approach validated for human-in-the-loop use, and the interface adds three interaction modalities that carry the argument: data-point filters for segmentation, situational-awareness panels (a county-level choropleth and indicator bar charts built from score aggregations), and rolling-mean heatmaps with variance tags that show how event scores change as data are revised. Because scores, unlike raw values, can be compared across geographic tiers and indicators, the ranked list makes it possible to aggregate context without the correlation errors that plague raw-value fusion.","core_discovery":"The paper's central claim is that replacing threshold-based alerts with a ranked list of anomalies, presented with contextual displays, transforms a failing public health monitoring workflow into one that works at national scale. The deployed system, built through an 18-month collaboration among data reviewers, engineers, and computer scientists, uses an unsupervised human-in-the-loop outlier-ranking method that scores data points by how far they deviate from reviewer-defined expectations, and the interface lets reviewers expand each ranked row to see sibling, parent, and child streams, toggle geographic context, and record structured events and meta-events. The longitudinal evaluation showed reviewers spent more time per row, recorded up to 49 events per session versus 1-2 in the prior baseline, and identified meta-events that were re-analyzed and matched notable public health events. The authors report that reviewers were 54x faster on average than with the previous exploratory manual system and faster than with the AI method presented alone; the abstract summarizes this as a 54x increase in reviewer speed efficiency compared with traditional alert-based methods.","pith_inferences":["Beyond the paper: the 54x figure is best read as a comparison to an exploratory manual-review workflow, not to a functioning threshold-alerting system; a fair comparison against a modest-volume alert list would likely show a much smaller speedup, though the ranking approach still removes the threshold-tuning burden.","Beyond the paper: the same ranking-plus-context design should transfer to other high-volume monitoring domains, such as environmental sensors, financial transactions, and agricultural reports, where the failure modes it targets appear in the same form.","Beyond the paper: a testable extension is whether the benefit persists with less expert reviewers or a different anomaly-scoring method, since the current evaluation involves three domain experts who helped design the system."],"forward_implications":["A small reviewer team can handle up to 5,000,000 data points per day, a workload that would overwhelm a manual or threshold-alert workflow.","Reviewers record more real events per session and can identify meta-events, combined anomalies that signal higher-level phenomena, which were absent or rare in the baselines.","The system cuts the need for per-stream threshold tuning; reviewers instead adjust data-expectation inputs roughly monthly as data dynamics change.","The pre-registered longitudinal evaluation strategy, with sequential interface changes and multi-week adaptation periods, provides a reusable pattern for assessing human-AI monitoring systems."],"supporting_citations":[{"why":"Documents the statistical challenges and distrust that plague threshold-based biosurveillance, motivating the shift to ranking.","marker":"Shmueli and Burkom [2010]"},{"why":"Survey of outlier-detection methods that grounds the choice of a multiple-univariate approach for the ranking system.","marker":"Blázquez-García et al. [2021]"},{"why":"Supplies the validated human-in-the-loop, linear-time outlier-ranking method at the core of the deployed system.","marker":"Joshi et al. [2024]"},{"why":"Reports the overwhelming alert volume, 35,000+ alerts, that motivates abandoning threshold-based alerting.","marker":"Coletta and Zhou [2019]"},{"why":"Characterizes the nonstationary hierarchical structure of public health data that prevents simple cross-dimension comparisons.","marker":"Reinhart et al. [2021]"},{"why":"Defines the drill-down fallacy that made the exploratory baseline inefficient and that the ranked interface avoids.","marker":"Lee et al. [2019]"},{"why":"Provides the online rolling mean/variance algorithm used to visualize how event scores evolve across data revisions.","marker":"Welford [1962]"},{"why":"Supports the design choice to communicate uncertainty over time in disease-data visualization.","marker":"Carroll et al. [2014]"}],"fun_headline_variants":["Ranked lists, not alerts, speed public health monitoring 54x","AI-ranked anomaly lists beat alert thresholds in deployed health system","54x faster outbreak spotting with AI-ranked data lists","From alerts to ranked AI triage: 54x speedup in health monitoring","National health monitoring finds events 54x faster via ranked AI"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The result assumes the measured speedup comes from the ranking paradigm and interface design rather than from reviewer learning, motivation, or the specific baseline used, since Baseline 1 was a manual exploratory tool, not a functioning alerting system.","fun_headline_variants_meta":{"raw":{"variants":["Ranked lists, not alerts, speed public health monitoring 54x","AI-ranked anomaly lists beat alert thresholds in deployed health system","54x faster outbreak spotting with AI-ranked data lists","From alerts to ranked AI triage: 54x speedup in health monitoring","National health monitoring finds events 54x faster via ranked AI"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000381,"raw_usage":{"total_tokens":1989,"prompt_tokens":882,"completion_tokens":1107,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":498,"completion_tokens_details":{"reasoning_tokens":1018}},"tokens_in":498,"tokens_out":1107,"duration_ms":7626,"temperature":1.0,"reasoning_tokens":1018,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T10:42:22.977006+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A crossover deployment in which the same reviewers alternate between the ranked-list interface and a threshold-alert interface that emits only a manageable number of top alerts per day would settle the claim: if the speed advantage over that alert interface is small or zero, the 54x result is specific to the exploratory baseline, not to alerting as such.","supporting_citations":[{"cited_title":"Statistical challenges facing early outbreak detection in biosurveillance","cited_arxiv_id":null,"evidence_quote":"Documents the statistical challenges and distrust that plague threshold-based biosurveillance, motivating the shift to ranking."},{"cited_title":"Outlier ranking in large-scale public health streams","cited_arxiv_id":null,"evidence_quote":"Supplies the validated human-in-the-loop, linear-time outlier-ranking method at the core of the deployed system."},{"cited_title":"What can you really do with 35,000 statistical alerts a week anyways? Online Journal of Public Health Informatics , 11(1), 2019","cited_arxiv_id":null,"evidence_quote":"Reports the overwhelming alert volume, 35,000+ alerts, that motivates abandoning threshold-based alerting."},{"cited_title":"An open repository of real-time covid-19 indicators","cited_arxiv_id":null,"evidence_quote":"Characterizes the nonstationary hierarchical structure of public health data that prevents simple cross-dimension comparisons."},{"cited_title":"Avoiding drill-down fallacies with vispilot: Assisted exploration of data subsets","cited_arxiv_id":null,"evidence_quote":"Defines the drill-down fallacy that made the exploratory baseline inefficient and that the ranked interface avoids."},{"cited_title":"Note on a method for calculating corrected sums of squares and products","cited_arxiv_id":null,"evidence_quote":"Provides the online rolling mean/variance algorithm used to visualize how event scores evolve across data revisions."},{"cited_title":"Visualization and analytics tools for infectious disease epidemiology: a systematic review","cited_arxiv_id":null,"evidence_quote":"Supports the design choice to communicate uncertainty over time in disease-data visualization."}],"review_version":1}