{"id":"1705d01c-af33-45d8-93b0-3e751b6a1a57","arxiv_id":"2505.11925","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"This paper provides six publicly available screw driving datasets with 34,182 operations covering thread degradation, surface conditions, assembly faults, and injection molding variations.","lead":"This paper releases six industrial screw driving datasets with over 34,000 operations, along with a Python library for loading and preprocessing them. The data is meant to give researchers a standardized benchmark for quality control, anomaly detection, and fault classification in assembly automation.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Headline count (34,182) contradicts Table 1 sum (34,082), and Section 5.1 gives two different DOIs/versions for the same collection; the paper's central access/metadata claims are not self-consistent.","rationale":"This is a dataset-description paper, so the central claim is not an analytical derivation but an availability/metadata claim: six datasets, >34,000 operations, controlled collection, persistent DOI, and a Python library. The most load-bearing weakness is not the acknowledged limitation about laboratory-to-production transferability, which the authors disclose in Section 6 and which does not falsify the record as a description of these specific experiments. The more concrete problem is that the manuscript's own headline numbers and access metadata are inconsistent: the stated total of 34,182 does not match the sum of the dataset table (34,082), and Section 5.1 lists two different DOIs/versions for the same collection. These inconsistencies sit directly on the central claim that the resource is accessible and citable, and they are checkable by any reader. The reader's verdict was CONDITIONAL, and this finding reinforces that condition; it does not move the verdict to a different category, so UNCHANGED is appropriate. Agreement is partial because the reader's stated weakest assumption was external validity, whereas the sharper, immediately verifiable concern is internal metadata integrity; the reader's rationale did mention the DOI inconsistency, which is the part I agree with.","tokens_in":14472,"tokens_out":8092,"duration_ms":69105,"concrete_test":"Download both Zenodo records (10.5281/zenodo.14729547 and 10.5281/zenodo.15273503), inspect their version metadata, and for each scenario ZIP recompute the operation count from labels.csv. Then compare the recomputed total to 34,182 and to the sum of Table 1 (34,082), and check whether both DOIs resolve to the same record/version. If the totals differ by 100 operations or the DOIs point to different versions, the manuscript's count and citation metadata are wrong as printed and must be corrected before the resource is widely cited.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central deliverable is a citable, accessible dataset, so the integrity of its counts and access metadata is load-bearing. Section 4 states 'the collection contains 34,182 individual screw driving operations' immediately before Table 1. Summing the Table 1 entries (5,000 + 12,500 + 1,700 + 5,000 + 2,400 + 7,482) gives 34,082, a 100-operation discrepancy that is not explained anywhere in the manuscript. Separately, Section 5.1 identifies the persistent DOI as https://doi.org/10.5281/zenodo.14729547 and says the current version is v1.2.2, but the recommended citation in the same subsection cites v1.2.1 with a different DOI (https://doi.org/10.5281/zenodo.15273503), while reference [4] cites v1.2.2 with the persistent DOI. A reader who follows the printed citation will reference a different version/DOI than the one the text designates as persistent. These are precisely the fields a dataset paper must get right: the total number of operations and the canonical way to cite and retrieve the data. Neither can be trusted from the manuscript alone, so the central claim 'accessible via a persistent Zenodo DOI' and the headline scale figure require external verification. The underlying dataset may be perfectly fine; the manuscript as written is internally inconsistent on its own key numbers.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper describes PyScrew, a collection of six industrial screw driving datasets (s01-s06) with a claimed total of 34,182 operations. It documents the automated screwing station, the four process phases, the acquisition system (torque, angle, gradient at 833.33 Hz), the OK/NOK classification criteria, and each scenario's experimental design: thread degradation, surface friction, two assembly-fault sets, and upper/lower workpiece injection-molding variations. Access is via a Zenodo DOI and a PyPI/GitHub Python library, with code examples for loading and filtering data. The paper closes with limitations (controlled laboratory setting, single screw type and plastic housing) and future work.","tokens_in":14753,"tokens_out":9982,"duration_ms":87514,"significance":"PyScrew's intended contribution is a large, standardized public resource for screw driving process data. The experimental design is documented in unusual detail: a single automatic station with Bosch Rexroth components, four torque- and angle-controlled phases, 833.33 Hz acquisition, and explicit OK/NOK criteria. The six scenarios cover thread degradation, surface conditions, assembly faults, and injection-molding parameter variations, at a scale an order of magnitude beyond AURSAD. The companion Python library and persistent repository are concrete reproducibility artifacts. If the stated counts and access metadata are corrected, this is a strong contribution to reproducible manufacturing research; the paper is also honest about the controlled-laboratory limitation and its implications for transfer to real production environments.","major_comments":[{"comment":"Section 4 states that 'the collection contains 34,182 individual screw driving operations' immediately before Table 1, but the Table 1 entries (5,000 + 12,500 + 1,700 + 5,000 + 2,400 + 7,482) sum to 34,082. The 100-operation discrepancy is unexplained anywhere in the manuscript. I verified that the per-scenario class tables for s02, s03, s04, s05, and s06 sum to their claimed scenario totals, so the arithmetic error is localized to the collection-level total. Because the scale of the collection is a central claim, please correct the count in the abstract, introduction, Section 4, and conclusion, or reconcile the table entries. I did not download the repository, so I cannot determine which number is correct; the manuscript as written is not self-consistent on its headline figure.","section":"Section 4, Table 1"},{"comment":"The data access section identifies https://doi.org/10.5281/zenodo.14729547 as the persistent DOI and states that the current version is v1.2.2, but the recommended citation in the same subsection cites v1.2.1 with DOI https://doi.org/10.5281/zenodo.15273503, while reference [4] cites v1.2.2 under the persistent DOI. A reader following the printed citation will reference a different version and DOI than the one the text calls persistent. Please make the version number and DOI consistent among Section 5.1, the recommended citation, and the reference list, and confirm which DOI is the concept DOI that tracks the latest version.","section":"Section 5.1, References [4]"},{"comment":"The text says that the 833.33 Hz sampling rate results in approximately 3-4 data points per degree of rotation. At the stated spindle speeds (finding 150 rpm, screw-in 600 rpm, pre-tightening 200 rpm, final tightening 40 rpm), the corresponding densities are about 0.93, 0.23, 0.69, and 3.47 samples per degree, so the '3-4 per degree' statement holds only for the final tightening phase. The claim of 5,000-7,000 data points per operation also appears hard to reconcile with the phase angles and speeds. Please clarify whether the controller records at fixed angular increments (with 0.25 degree resolution) or fixed time intervals, and correct the derived estimates accordingly.","section":"Section 3.3"}],"minor_comments":[{"comment":"Reference [9] contains a typo in the title, 'Machien Leanring'; it should read 'Machine Learning'.","section":"References [9]"},{"comment":"The phrase 'data completeness above 95%' is used repeatedly but never defined; please state how completeness is measured (for example, the fraction of expected measurement points present after validation).","section":"Section 3.3"},{"comment":"The first code example reads `data[\"torque values\"]`, while the configuration example passes `return_measurements=[\"torque\", \"angle\"]`; please clarify the actual key names returned by `get_data` so the examples are internally consistent.","section":"Section 5.2"},{"comment":"The row for `003_control-group-from-s01` says it uses 'first cycles from s01 control group' without specifying whether the 200 samples are the first cycle of each of 100 workpieces at two locations; please state the selection rule to make the row unambiguous.","section":"Table 4"},{"comment":"The description of West et al. [2] mentions an imbalanced dataset with 50,000 normal and 96 anomalous samples; since no PyScrew scenario has 50,000 samples, please state explicitly that this was a separate industrial dataset so that readers do not confuse it with the collection presented here.","section":"Section 2.2"}],"recommendation":"major_revision","confidential_remarks":"For the editor: the related-work section is dominated by the authors' own prior publications, but the paper acknowledges this explicitly and those works are genuine applications of the same data, so I do not view this as a citation-ethics problem. The blocking issues are internal consistency of the headline count and the citation/DOI metadata; both are fixable with a careful pass against the repository. I did not download the data, so I cannot adjudicate which count is correct, but the manuscript itself must be unambiguous before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: PyScrew is exactly the kind of resource this area needs — a large, openly available, multi-scenario screw-driving time series collection — but the paper's central metadata is not internally consistent, and that needs to be fixed before the dataset is widely cited.\n\nWhat's new: compared with AURSAD (~2,045 samples, five operation types), PyScrew offers six datasets, nominally 34k operations, with broader fault taxonomies (s03 has 26 classes, s04 25, s05 42, s06 44), repeated-use degradation, surface-condition variations, and injection-molding parameter sweeps. The experimental setup is documented in unusual detail: Bosch Rexroth station, four process phases, torque/angle/gradient at 833.33 Hz, OK/NOK rules, and a hierarchical data model. The dual access path (Zenodo + Python library) is practical and reproducible, assuming the data is actually there. The authors also acknowledge obvious limitations — one station, one screw type, one plastic housing, lab conditions — so they are not overselling transferability.\n\nSoft spots, in order of importance:\n\n1. The headline count does not match the table. Section 4 says 34,182; Table 1 sums to 34,082 (5,000+12,500+1,700+5,000+2,400+7,482). A 100-operation unexplained difference is not a fatal scientific flaw, but for a dataset paper the total count is a load-bearing number. It must be reconciled.\n\n2. The DOI/version situation is tangled. Section 5.1 names https://doi.org/10.5281/zenodo.14729547 as persistent and current v1.2.2, but the recommended citation in the same subsection gives v1.2.1 with a different DOI (https://doi.org/10.5281/zenodo.15273503), while reference [4] gives v1.2.2 with the persistent DOI. A reader who follows the printed citation will cite a different artifact than the one the text designates as canonical. Also fixable, but it is exactly the field a dataset paper must get right.\n\n3. Verification requires downloading the data. That is normal for a dataset paper, but it means the paper's claims are checkable only after the fact. I would not weigh this heavily — the internal statistics, the per-class tables, and the setup description are coherent enough that I expect the data to exist as described.\n\nThe related-work section leans heavily on the authors' own prior papers, but those are applications of the same data, not evidence for its validity, so I do not see a circularity problem.\n\nVerdict: the central contribution is real and needed. Fix the count and the DOI/version mismatch, verify the repository matches the paper, and this deserves publication as a resource paper. I would send it to peer review. If I were working on industrial time series or anomaly detection, I would cite it once I had done a quick download check.","headline":"PyScrew is a genuinely useful open dataset for screw-driving research, but its headline count and DOI citations do not add up and need fixing before anyone cites it.","tokens_in":15263,"tokens_out":2531,"would_cite":true,"duration_ms":24260,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper introduces PyScrew, a public collection of six screw-driving datasets with more than 34,000 operations, aiming to give industrial process-monitoring research a standardized, reproducible resource.","keywords":["screw driving","industrial time series dataset","process monitoring","anomaly detection","quality control","manufacturing","machine learning benchmark","assembly automation"],"falsifier":"Train a fault classifier on the PyScrew s04 assembly-fault data and evaluate it on screw-driving traces recorded from a different automatic station with the same screw and housing; if performance on held-out industrial data falls to near chance while within-collection cross-validation stays high, the claim that the collection generalizes as a general benchmark is falsified. A lighter check: recompute the s03 control-group separability from the March 2023 and February 2024 recordings and see whether method rankings shift across recording times.","tokens_in":14271,"feed_emoji":"🔩","tokens_out":5169,"duration_ms":49621,"temperature":0.7,"pith_summary":"The paper introduces PyScrew, a public collection of six datasets containing 34,182 screw driving operations logged on one automatic screw station under controlled conditions. The datasets systematically vary thread wear, surface friction, assembly faults, and injection-molding parameters of plastic housings, each run recorded as torque, angle, and gradient time series. The claim is that this collection gives researchers a standardized, reproducible resource for developing and comparing anomaly detection, classification, and process-monitoring methods in industrial assembly. If that holds, work that previously required proprietary factory data can be done openly, and published methods can be compared on the same data.","feed_headline":"34,000 screw-driving runs released as open benchmark data","feed_subtitle":"Six controlled datasets cover thread wear, surface faults, and molding variations so industrial ML results become comparable.","key_machinery":"The load-bearing artifact is the standardized experimental setup combined with a hierarchical data model. Each screw drive is a multivariate time series of torque, angle, and gradient sampled at 833.33 Hz over four phases—finding, driving in, pre-tightening, and final tightening—organized by scenario, class, and operation ID, with metadata including usage count and outcome labels. This structure is what lets the same measurements serve anomaly detection, classification, feature-extraction evaluation, and process-monitoring research.","core_discovery":"The central claim is that PyScrew is a comprehensive open collection of screw driving data that addresses the scarcity of standardized industrial datasets. It comprises six scenarios: natural thread degradation from repeated use (s01, 5,000 runs), surface friction variations including contamination and treatments (s02, 12,500 runs), two assembly-fault collections with up to 26 and 25 error classes (s03 with 1,700 runs and s04 with 5,000 runs), and injection-molding parameter variations for the upper and lower workpieces (s05 with 2,400 runs and s06 with 7,482 runs). Every operation is captured as time series of torque, angle, and gradient at 833.33 Hz across four process phases, with metadata such as workpiece usage count and OK/NOK outcome labels. The authors provide both raw data through a persistent DOI and a purpose-built Python library that handles loading, validation, and preprocessing, so that researchers can reproduce analyses and compare methodologies fairly.","pith_inferences":["An implied next step the authors do not spell out: the usage-count metadata in s01 could support remaining-useful-life models for plastic threads, turning the degradation dataset into a prognosis benchmark rather than just a classification one.","Because everything was recorded on one station with one screw type and one plastic housing material, the collection is best read as a controlled testbed; whether insights transfer to other geometries, materials, or screw types is an open empirical question not settled by this paper.","The paired error conditions in s02 (water versus oil lubricant, coarse versus fine sanding) invite studies of error-severity or signature similarity, a comparison the authors mention but do not perform.","A testable extension would be to measure how stable feature-extraction rankings remain across the repeated control groups recorded months apart, which would quantify the collection's temporal consistency."],"forward_implications":["The same 34,182 operations can serve as a common benchmark, so anomaly-detection and classification results from different research groups become directly comparable.","Each scenario isolates one family of influences—thread wear (s01), surface friction (s02), assembly faults (s03-s04), fabrication parameters (s05-s06)—letting researchers attribute performance differences to a specific physical factor.","The s05 and s06 fabrication datasets, with mostly OK outcomes, allow studies of how process parameters alter torque and angle traces before quality limits are exceeded.","The hierarchical data model and filtering options let users focus on specific phases, measurements, or classes, supporting both exploratory analysis and machine-learning pipelines.","Open access reduces the need for proprietary factory data, which the authors argue is a barrier to reproducible research in manufacturing quality control."],"supporting_citations":[{"why":"Supplies the predecessor open screwdriving dataset (about 2,000 samples, five operation types) whose scale and scope PyScrew is positioned to expand.","marker":"[9]"},{"why":"The persistent dataset record on the open repository that makes the six collections accessible and citable.","marker":"[4]"},{"why":"Establishes the anomaly-detection baseline on s01 (thread degradation) that the collection is meant to support.","marker":"[5]"},{"why":"Shows supervised detection of surface-based anomalies on s02, one of the collection's benchmark scenarios.","marker":"[6]"},{"why":"Compares feature-extraction methods on s02, providing a reference evaluation the collection enables.","marker":"[7]"},{"why":"Demonstrates multi-class error detection across 25 error types on s04, a central use case for the assembly-conditions dataset.","marker":"[8]"}],"fun_headline_variants":["34k screw-driving runs open for ML benchmark","Six screw-driving datasets: 34k runs released","Open screw-driving data covers faults and wear","PyScrew: 34k industrial screw-driving operations","Screw-driving benchmark: 34k runs, 6 scenarios"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The collection's usefulness as a general benchmark depends on one laboratory station, one screw type, and one plastic housing material capturing enough of real production variability that methods validated on it transfer to industrial practice—a limitation the authors state in Section 6.","fun_headline_variants_meta":{"raw":{"variants":["34k screw-driving runs open for ML benchmark","Six screw-driving datasets: 34k runs released","Open screw-driving data covers faults and wear","PyScrew: 34k industrial screw-driving operations","Screw-driving benchmark: 34k runs, 6 scenarios"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000177,"raw_usage":{"total_tokens":1321,"prompt_tokens":1003,"completion_tokens":318,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":619,"completion_tokens_details":{"reasoning_tokens":241}},"tokens_in":619,"tokens_out":318,"duration_ms":3801,"temperature":1.0,"reasoning_tokens":241,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:44:13.039092+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train a fault classifier on the PyScrew s04 assembly-fault data and evaluate it on screw-driving traces recorded from a different automatic station with the same screw and housing; if performance on held-out industrial data falls to near chance while within-collection cross-validation stays high, the claim that the collection generalizes as a general benchmark is falsified. A lighter check: recompute the s03 control-group separability from the March 2023 and February 2024 recordings and see whether method rankings shift across recording times.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The persistent dataset record on the open repository that makes the six collections accessible and citable."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the anomaly-detection baseline on s01 (thread degradation) that the collection is meant to support."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows supervised detection of surface-based anomalies on s02, one of the collection's benchmark scenarios."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Compares feature-extraction methods on s02, providing a reference evaluation the collection enables."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Demonstrates multi-class error detection across 25 error types on s04, a central use case for the assembly-conditions dataset."}],"review_version":1}