{"id":"aa5359f2-9c73-4b7d-9836-33256548d35e","arxiv_id":"2608.08892","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"low","formal_verification":"none","parameter_count":4,"one_line_summary":"In a fixed six-corpus panel, MyFixit and Doc2Dial both pass the individual content and container layer checks yet fail when the layers are unioned, and the panel cannot separate bridge-specific causes from simple graph density.","lead":"This paper audits six procedural corpora to see whether two relation layers each keep data in many small groups while their union collapses those groups into giant components. Two of six datasets show this pattern, and the paper carefully notes that the result is a descriptive panel finding, not a general claim.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central 2/6 qualifier count depends on a Human Know-How extraction path with a known mojibake risk; a corrected rerun could move a near-threshold nonqualifier into the qualifier set, so the frozen aggregate evidence cannot currently establish the count.","rationale":"The reader's stated weakest assumption was the author-chosen operational thresholds A1-A4. That is a real limitation, but the paper explicitly registers those thresholds as operational and makes no claim that the qualifier set is invariant to them; the central claim is scoped to the frozen rule. The more load-bearing uncertainty is the Human Know-How parser risk, which the paper itself discloses. The frozen aggregate artifact allows inspection and recomputation of the reported tables, but it cannot correct a known decoding flaw in the only parser path that was recovered. Human Know-How sits close to the qualification boundary, so the 2/6 count and the identity of the qualifier set could change under a corrected rerun. This does not make the paper dishonest or the reasoning circular; it means the exact central numerical claim is not fully verifiable from the public materials. The reader's conditional verdict already anticipates reproducibility concerns, and the recommended condition should explicitly include the Human Know-How corrected rerun. Because the reader's verdict is already CONDITIONAL, no verdict adjustment is needed; the condition should be sharpened to include this parser path. Agreement with the reader is partial: their weakest-assumption field identifies definition-dependence rather than the parser risk, although their rationale does mention the corrected Human Know-How rerun as a condition.","tokens_in":14415,"tokens_out":5151,"duration_ms":51076,"concrete_test":"Run a corrected Human Know-How loader on the original WikiHow-derived records: decode UTF-8 before any unicode-escape or normalization step, rebuild the C2/C3/C5 component tables for all six sources, and reapply the frozen source-level Gate. If Human Know-How remains a nonqualifier while MyFixit and Doc2Dial remain qualifiers, the 2/6 claim survives; if Human Know-How enters the qualifier set (union failure count rises from 1/4 to at least 2/4) or any stated qualifier drops out, the central count must be revised. Publish the corrected component metrics and the parser-fix diff as part of the artifact.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is the frozen source-level result '2 of 6 gate-eligible sources qualify.' That count is not a prevalence estimate, but it is still an exact assertion about the six-source panel. The weakest point is not the author-chosen A1-A4 thresholds: those are registered operational rules, and the paper explicitly declines to generalize beyond them. The load-bearing gap is the admitted extraction risk in the Human Know-How loader (Sections 10.3 and 11.2). The loader passes UTF-8 Turtle label bytes through unicode escape, which can introduce mojibake before text normalization, changing content-near-duplicate edges and hence C2/C5 component structure. Human Know-How is currently a nonqualifier, but by a narrow margin: under the union, its largest-component share is 0.1376 (A1 fails), while A2-A4 pass (share <=0.20, Neff=52.2>=30, share <=1/3). If a corrected parser shifts the union component structure, for instance pushing largest-share above 0.20 or Neff below 30, Human Know-How would fail at least two further criteria and would join the qualifier set, changing 2/6 to 3/6 and invalidating the stated qualifier set. The authors state that 'neither the direction nor the magnitude of possible drift in the frozen Human Know-How metrics is known,' so the aggregate artifact cannot rule this out. Because the central claim is exactly the count and identity of qualifiers, this unresolved parser path is the most load-bearing concern.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a symmetric, operationally defined audit of component structure in a fixed panel of six hierarchical procedural corpora (Human Know-How, MyFixit, OpenPI, WIQA, Doc2Dial, X-WLP). For each source, a content layer (C2, exact/near-duplicate normalized text), a container layer (C3, clique per document-like container), and their transitive-closure union (C5) are summarized by component count, largest-component share, and effective component count. A registered source-level rule (criteria A1–A4, with each layer required to pass and the union required to fail at least 3/4 of the criteria) yields two qualifiers, MyFixit and Doc2Dial, reported strictly as a descriptive panel fraction and explicitly not as a prevalence estimate. The paper further reports that a prespecified bridge-specific predictor (bridge edge density) is associated with the qualifier pattern but is not discriminated from a registered density control (mean union degree), so no bridge-specific mechanism is identified. Secondary analyses—a discrete τ-threshold sweep, a coverage-exposure near-identity on two corpora, and a lexical lower-bound diagnostic—are presented with explicitly narrow evidentiary scope. The paper's stated contribution is a bounded measurement and audit protocol with negative results, not a new splitting algorithm, a giant-component theorem, or a causal claim.","tokens_in":14763,"tokens_out":14110,"duration_ms":131915,"significance":"If the measurement is correct, the paper establishes that under one coherent operational rule the individual-layer-pass/union-fail pattern occurs in two of six procedurally distinct corpora (a repair-guide corpus and a government-service corpus) while four negative cases remain visible, including WIQA, where union formation does not cause collapse. The paper is exemplary in its self-scoping: it registers its thresholds, refuses to convert the panel fraction into a prevalence statement, and retains the mechanism negative result instead of rescuing a bridge-specific explanation with post hoc comparisons. The public artifact supports regeneration of the scientific tables and byte-for-byte verification of derived outputs, and the paper is explicit that this is not a raw-data-to-paper reproduction. The honest reporting of the density-control confound is a useful methodological caution for the leakage-control community. The contribution is modest—an audit rather than an algorithm—but it is falsifiable and carefully bounded, and the negative findings are themselves informative.","major_comments":[{"comment":"§10.3 and §11.2: the unresolved mojibake risk in the Human Know-How loader is load-bearing for the headline 2/6 count, and the manuscript's own disclosure of the risk does not close the gap. The loader passes UTF-8 Turtle label bytes through unicode escape before normalization, so corrected parsing could change C2 near-duplicate edges and, through transitive closure, the C5 component structure. Human Know-How is a near-threshold nonqualifier: its union largest-component share is 0.1376, which is 0.0624 below the A2 cap of 0.20 and well below the A4 cap of 1/3, and its union effective-component count is 52.2, only 22.2 above the A3 floor of 30; its content-layer largest share of 0.007897 is only 0.0021 below the A1 cap of 0.01. A corrected parser that raises union concentration (for example, pushing the largest share above 1/3, or above 0.20 while Neff falls below 30) would make the union fail at least 3/4 of A1–A4, so Human Know-How would join the qualifier set if its individual layers still pass, changing the count from 2/6 to 3/6. The paper states that neither the direction nor the magnitude of drift is known (§11.2), so the frozen aggregate evidence cannot currently establish the central count as a property of the six sources. The authors should either rerun the Human Know-How path with a corrected parser and report whether the frozen metrics and qualifier set survive, or re-scope the headline claim so that '2 of 6' is explicitly and prominently stated as an assertion about the frozen aggregate records, with the parser risk named at the point of assertion rather than only in the limitations.","section":"§10.3 and §11.2"},{"comment":"§1, §3, and §12 state the result as a property of the sources ('2 of 6 gate-eligible sources qualify'; 'MyFixit and Doc2Dial exhibit the individual-layer-pass/union-fail pattern'), whereas §10.3 and §11.2 concede that the evidence cannot establish that corrected parsing of Human Know-How leaves the qualifier set unchanged. This inconsistency matters because the exact count and identity of qualifiers is the paper's central claim. The abstract and conclusion should carry the same qualification that the limitations carry—either by describing the result as holding 'in the frozen aggregate records' or by stating that the Human Know-How path awaits a corrected-parser rerun. As written, a reader who only reads the abstract and conclusion would reasonably take the 2/6 count as a settled fact about the corpora, which §11.2 explicitly disclaims.","section":"§1, §3, §12 vs. §10.3, §11.2"}],"minor_comments":[{"comment":"The caption's visual encoding ('Thick filled-marker lines') does not match the marker glyphs ('c', 'u') shown in the plot; state explicitly that both line weight and marker form distinguish the configurations and qualifying status, and consider a small-multiples layout or direct source labels given the density of 18 markers on a log axis.","section":"Figure 1"},{"comment":"The term 'Gate-C exact endpoint' appears without prior definition; rename or define this object.","section":"§6"},{"comment":"The sentence 'The median column is included because all six frozen aggregate diagnostics contain a non-null value' is a confusing justification; state plainly that the median is reported for descriptive completeness.","section":"§2.1"},{"comment":"The expressions 'universe[0] tool' and 'universe[0] location family' are not defined anywhere; the reader cannot tell whether this is an array-index artifact of the frozen report or a domain term.","section":"§8"},{"comment":"The text reports that the bridge predictor's correlation with C5 largest-component share (0.943) equals its cross-predictor correlation with mean union degree (0.943); adding one sentence noting this equality is coincidental would prevent readers from inferring a structural relation.","section":"§5"},{"comment":"The column header layout ('n comp top1 share Neff' with repeated 'exact 0.50' labels) is hard to parse; align each statistic with its two grid-point columns.","section":"Table 6"}],"recommendation":"major_revision","confidential_remarks":"The paper is far more self-scoping than most submissions, and nearly every potential objection (threshold arbitrariness, outcome-enriched panel, mechanical one-source-deletion persistence, post hoc status of the A2–A4 rule-shape check) is already raised and addressed in the text. The single genuinely load-bearing unresolved point is the Human Know-How parser path, and the question for the editor is whether a re-scoped claim ('result holds for the frozen aggregate records; corrected-parser status open') is sufficient for this venue, or whether a corrected rerun should be required. The contribution is a measurement and audit rather than an algorithmic advance, and the cs.IR framing is appropriate given the leakage-control and near-duplicate literature. If the authors re-scope the abstract and conclusion as suggested and strengthen the artifact with a corrected-parser confirmation, I would regard the paper as publishable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. This is a careful, honest audit of relation-union component collapse in procedural corpora; it does not claim the phenomenon as new, and it goes out of its way to keep its claims scoped. The headline 2-of-6 qualifier count, however, is not fully reproducible: a disclosed decoding risk in the Human Know-How loader could plausibly change the count, and the authors have not rerun the pipeline.\n\nWhat is genuinely good: the symmetric layer-union protocol (C2 content-near-duplicate, C3 container-clique, C5 union) applied to a fixed six-source panel, with all six sources reported in Table 1, explicit negative cases, a prespecified Gate-2 mechanism test that fails and is reported as a null result, and post hoc analyses clearly labeled as presentation-only. The arithmetic is internally consistent; I spot-checked the component shares and effective counts, and nothing wobbles. The citation pattern is responsible: they attribute the merged-relation giant-component chain to Guvenilir and Dogan and do not claim priority.\n\nThe load-bearing soft spot is the Human Know-How parser. The loader passes UTF-8 Turtle labels through unicode escape, which can introduce mojibake before text normalization, and that can change content-near-duplicate edges and hence C2/C5 components. Human Know-How is currently a nonqualifier by one criterion: under the union it fails only A1 (largest share 0.138), while A2-A4 pass. A corrected parser that moves the union share above 0.20 or drops Neff below 30 would flip it to a qualifier, changing 2/6 to 3/6. The authors disclose this and explicitly say the direction and magnitude are unknown. That is honest, but it means the central empirical assertion is not yet established. They need a corrected rerun or a raw-data release before I'd trust the count.\n\nThe other limits are minor and mostly acknowledged: the A1-A4 thresholds are author-chosen operational rules, the panel is outcome-enriched and not a prevalence sample, and the mechanism is null. The paper says all of this up front.\n\nWho is it for? People building leakage-aware splits in procedural NLP or working on graph-based benchmark construction. It won't change broad practice. But it is a serious, well-scoped measurement study with a real flaw that is fixable. I would send it out; ask for the Human Know-How rerun (or at least a parser-sensitivity analysis) before final acceptance.","headline":"Honest, well-scoped audit of a known phenomenon; the 2/6 qualifier count remains unproven due to a disclosed but unquantified mojibake risk in the Human Know-How pipeline.","tokens_in":15253,"tokens_out":3338,"would_cite":false,"duration_ms":29830,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A symmetric audit shows two of six procedural corpora pass each split layer but fail their union.","keywords":["data leakage","component-disjoint splitting","near-duplicate detection","hierarchical procedural corpora","transitive closure","effective component count","train-test leakage","leakage-aware partitioning"],"falsifier":"Recompute the frozen component summaries from the raw corpora with a corrected Unicode loader and an independent implementation of the normalized near-duplicate rule; if either Doc2Dial or MyFixit fails to show both individual layers passing A1–A4 while the union fails at least three of them, the two-of-six result is wrong. A second check would be a full six-source threshold sweep: if the qualifier set changes across a plausible grid around $\\tau=0.85$, the reported outcome is an artifact of that cutoff.","tokens_in":14200,"feed_emoji":"🧩","tokens_out":7573,"duration_ms":72889,"temperature":0.7,"pith_summary":"The paper asks whether a data split that looks safe when edges of one type are considered alone remains safe when two edge families are combined. Across six hierarchical procedural corpora, it builds a content-near-duplicate layer and a common-container layer, closes each transitively, and then closes their union. It finds that two of the six sources, MyFixit and Doc2Dial, pass the individual-layer checks yet fail the same checks for the union, with Doc2Dial's largest component growing to 98.80% of units. The two-of-six count is explicitly a descriptive property of this deliberately assembled panel, not an estimate of how often such collapse happens. The paper also reports that a bridge-specific explanation could not be separated from an ordinary union-density control, so no causal mechanism is claimed.","feed_headline":"Two of six corpora pass split checks, fail when layers merge","feed_subtitle":"A layer-by-layer split audit can hide unions that concentrate up to 98.8% of units into one component.","key_machinery":"The argument runs on three typed graph layers over a common set of procedural units: C2 links units with near-duplicate content, C3 expands every source-defined container into a clique, and C5 is the transitive closure of their union. Each configuration is summarized by component count, largest-component share, and effective component count $N_{\\mathrm{eff}} = \\left(\\sum_j p_j^2\\right)^{-1}$, and judged by criteria A1–A4 (largest-component share at most $0.01$, at most $0.20$, at least 30 effective components, and largest-component share at most $1/3$). These summaries keep the two edge families visible and make the comparison symmetric, so a large union component can be checked against each contributing layer under the same rule. The key contrast is that a low share in C2 and C3 does not constrain the share in C5, because alternating C2–C3 paths can join units that neither layer connects alone.","core_discovery":"Under the paper's operational source-level rule, which requires both individual layers to remain dispersed under criteria A1–A4 while their transitive-closure union becomes concentrated, two of six gate-eligible sources qualify. MyFixit and Doc2Dial each keep their largest-component share low in the content and container layers, but the union share jumps to 41.64% for MyFixit and 98.80% for Doc2Dial. The other four sources are reported as explicit negative cases, including WIQA, where union formation leaves the largest-component share at 1.49%, showing that union alone does not imply collapse. The authors emphasize that this two-of-six fraction describes the frozen, outcome-enriched panel and is not a prevalence estimate, and that the near-alignment of a bridge-density predictor with mean union degree prevents any bridge-specific mechanism from being identified.","pith_inferences":["A natural extension would be to recruit result-blind sources where bridge density and mean union degree make discordant predictions, which is the design needed to test whether ordinary density or a bridge-specific mechanism explains the pattern.","The two-corpus coverage-exposure near-identity suggests that annotated-field incompleteness could serve as a cheap proxy for one exposure statistic in other hierarchical procedural corpora; checking that on a third corpus would be a direct test.","Because the criteria place repeated weight on largest-component share, an alternative audit built from statistically independent summary measures might classify sources differently, so a wider threshold-grid sensitivity analysis could show how much of the 2/6 result is definitional."],"forward_implications":["Layer-wise acceptability is not sufficient evidence of union-level capacity; the union is an object that deserves its own audit.","The 2/6 result is confined to the frozen panel and operational definitions: it is not a base rate, not a prevalence estimate, and not evidence about population frequency.","The absence of mechanism discrimination means claims that bridging specifically drives collapse are not supported; a density control fits the same pattern.","Negative cases such as WIQA show that merging two layers does not automatically create a giant component, so cross-container content repetition is the measured condition under which collapse appears.","Because A1, A2, and A4 are nested thresholds on the same statistic, the criteria repeat weight on largest-component share rather than providing four independent tests."],"supporting_citations":[{"why":"Establishes the direct precedent that merging relation types can create a giant component that obstructs component-disjoint splitting, which this paper explicitly does not claim as new.","marker":"[7]"},{"why":"Supplies the split-feasibility context by showing how a giant component in a similarity graph prevents a required split, grounding the A2 fit-bound rationale.","marker":"[12]"},{"why":"Provides the restricted single-linkage partitioning baseline against which the layer-union component structure is contrasted in the related-work discussion.","marker":"[13]"},{"why":"Documents multi-graph split configurations and component recording, the type of benchmark construction this paper extends with a symmetric layer-union audit.","marker":"[4]"},{"why":"Shows how transitive closure can propagate a false-positive link across matching passes, grounding the closure-risk assumption behind the union C5.","marker":"[6]"},{"why":"Establishes near-duplicate handling as an evaluation concern, supporting the definition and relevance of the content-near-duplicate layer.","marker":"[5]"}],"fun_headline_variants":["Merging layers collapses two of six audited corpora","Pass individually, fail merged: two of six in audit","Union collapse in 2 of 6 corpora; no bridge mechanism","Symmetric audit: individual layers pass, union fails 2/6","Two of six corpora fail only when layers unite"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire 2/6 result rests on the authors' registered operational definitions—near-duplicate threshold $\\tau=0.85$, containers expanded into cliques, largest-component share at most $0.01$, at least 30 effective components, and the four-criteria decision rule—so a different but equally defensible choice of thresholds could change which sources qualify.","fun_headline_variants_meta":{"raw":{"variants":["Merging layers collapses two of six audited corpora","Pass individually, fail merged: two of six in audit","Union collapse in 2 of 6 corpora; no bridge mechanism","Symmetric audit: individual layers pass, union fails 2/6","Two of six corpora fail only when layers unite"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00023,"raw_usage":{"total_tokens":1477,"prompt_tokens":933,"completion_tokens":544,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":549,"completion_tokens_details":{"reasoning_tokens":457}},"tokens_in":549,"tokens_out":544,"duration_ms":6288,"temperature":1.0,"reasoning_tokens":457,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:21:42.837390+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the frozen component summaries from the raw corpora with a corrected Unicode loader and an independent implementation of the normalized near-duplicate rule; if either Doc2Dial or MyFixit fails to show both individual layers passing A1–A4 while the union fails at least three of them, the two-of-six result is wrong. A second check would be a full six-source threshold sweep: if the qualifier set changes across a plausible grid around $\\tau=0.85$, the reported outcome is an artifact of that cutoff.","supporting_citations":[{"cited_title":"How to approach machine learning-based prediction of drug/compound–target interactions.Journal of Cheminformatics, 15:16, 2023","cited_arxiv_id":null,"evidence_quote":"Establishes the direct precedent that merging relation types can create a giant component that obstructs component-disjoint splitting, which this paper explicitly does not claim as new."},{"cited_title":"Lo-Hi: Practical ML Drug Discovery Benchmark","cited_arxiv_id":null,"evidence_quote":"Supplies the split-feasibility context by showing how a giant component in a similarity graph prevents a required split, grounding the A2 fit-bound rationale."},{"cited_title":"GraphPart: homology partitioning for biological sequence analysis.NAR Genomics and Bioinformatics, 5(4):lqad088, 2023","cited_arxiv_id":null,"evidence_quote":"Provides the restricted single-linkage partitioning baseline against which the layer-union component structure is contrasted in the related-work discussion."},{"cited_title":"PLINDER: The protein–ligand interactions dataset and evaluation resource.bioRxiv, 2024","cited_arxiv_id":null,"evidence_quote":"Documents multi-graph split configurations and component recording, the type of benchmark construction this paper extends with a symmetric layer-union audit."},{"cited_title":"Record linkage: Current practice and future directions","cited_arxiv_id":null,"evidence_quote":"Shows how transitive closure can propagate a false-positive link across matching passes, grounding the closure-risk assumption behind the union C5."},{"cited_title":"The effect of content- equivalent near-duplicates on the evaluation of search engines","cited_arxiv_id":null,"evidence_quote":"Establishes near-duplicate handling as an evaluation concern, supporting the definition and relevance of the content-near-duplicate layer."}],"review_version":1}