{"id":"09ba905a-eed5-4b21-b484-57169763c4be","arxiv_id":"2501.14841","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey of 11 participants from one industry-academia project and an analysis of 280,732 Zenodo datasets show that data planning is hard, few datasets ship code, and licences are often misused.","lead":"The authors surveyed participants in a European industry-academia project that published 13 open datasets and analyzed metadata from nearly 281,000 Zenodo datasets. They report that planning, not technology, is the main challenge in publishing data, and that only about 2.4% of Zenodo datasets include scripts to help others reuse them.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 2.4% script rate is based on file extensions in Zenodo metadata and likely misses code packed inside archives, so the headline gap is not actually established.","rationale":"The reader's weakest-assumption analysis correctly identifies the file-extension proxy in Section 3.5 as the load-bearing point for the quantitative claim. My stress-test sharpens that concern: the paper's own observation that most Zenodo files are archives means the proxy is not merely an upper bound but has a systematic blind spot. Code shipped inside archives is a normal way to distribute parsing and reproduction scripts, so the reported 2.4% could be a substantial undercount. This does not invalidate the paper's qualitative lessons from the InSecTT survey, which are modestly framed and supported by the deposited secondary data; it does mean the headline numerical claim should be presented as an upper-bound estimate conditional on top-level file metadata, not as a measured fact. Since the reader already assigned CONDITIONAL for essentially this reason, no change in verdict is needed, but the concern is real and can be settled by the archive-sampling check described above.","tokens_in":16762,"tokens_out":2610,"duration_ms":25808,"concrete_test":"Take a random sample of 200 Zenodo datasets from the same 2000-2023 scrape whose top-level files are archives (zip, tar, gz, etc.), download each archive, and check whether it contains any .py, .m, .R, .ipynb, or other executable source file. Recompute the percentage of sampled datasets accompanied by scripts including these archive-contained scripts, and also compute a separate rate counting only .ipynb files. If the adjusted percentage is materially higher than 2.4% (for example, more than double), then the headline claim must be qualified as depending on top-level metadata only and as likely undercounting supporting code.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim, stated in the Abstract and repeated in Section 5, is that 'less than 2.4%' of Zenodo datasets are accompanied by scripts that improve reuse. The computation in Section 3.5 counts a dataset as having supporting software only if a file with extension .py, .m, or .R appears in the Zenodo file-level metadata. Two publication patterns break this proxy. First, the paper itself reports that the majority of Zenodo files are archives (Figure 2), and their contents are undisclosed in the metadata; code for parsing or reproducing datasets is routinely bundled inside .zip or .tar.gz files and would be invisible to the scraper. Second, Jupyter notebooks (.ipynb), shell scripts, and code in other languages (C/C++, Java, etc.) are not counted at all. The paper labels the 2.4% figure an 'upper bound' because some identified scripts may not be functional supporting software, but false negatives from archive contents push in the opposite direction and are arguably more prevalent. If even a small fraction of archived datasets contain scripts, the true rate could be several times 2.4%. Therefore the Abstract's claim that 'only few datasets had accompanying scripts' is not supported by the scrape as described.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports on open data publication practices in the InSecTT industry-academia collaboration, combining an open-ended survey of 11 project participants (covering 13 datasets) with a metadata analysis of 280,732 Zenodo datasets up to 2023. The authors identify qualitative lessons about planning, documentation, licensing, and synthetic data, and they quantify how often Zenodo datasets are accompanied by script-like files, reporting a rate of about 2.4%. The paper closes with a set of recommendations for publishing open datasets in industry-academia projects.","tokens_in":16939,"tokens_out":4954,"duration_ms":41560,"significance":"The topic is timely and the paper brings together a rare inside view of an industry-academia project with a large public-repository metadata sample. The authors deposit their secondary data at Zenodo [14], which supports reproducibility of the metadata analysis and the survey responses. The qualitative findings about planning, licences, and synthetic data are plausible and useful even with the small sample. However, the central quantitative claim—that only 2.4% of Zenodo datasets have accompanying scripts for improved reuse—rests on a file-extension proxy that the paper itself acknowledges to be incomplete, and the Abstract and Conclusion state the claim without the required caveats.","major_comments":[{"comment":"The 2.4% figure counts a dataset as accompanied by a script only if the Zenodo file-level metadata contains a file with extension .py, .m, or .R. Since §3.3 reports that the majority of published files are archives whose contents remain undisclosed, scripts bundled inside .zip or .tar.gz files are invisible to this count, as are Jupyter notebooks and scripts in other languages. Consequently the paper's own 'upper bound' characterization in §3.5 applies only to whether identified scripts truly function as supporting software; it does not address false negatives from archive contents. The Abstract's statement that 'only few datasets (2.4%) had accompanying scripts for improved reuse' is therefore a factual claim that the described method cannot support.","section":"§3.5, Figure 2, Abstract"},{"comment":"The Conclusion repeats the unsupported interpretation as 'very few Zenodo datasets (less than 2.4%) were accompanied with scripts that would improve reuse.' The phrase 'less than 2.4%' is justified neither by the measurement nor by the authors' own caveats; the measurement is a lower bound on datasets with directly visible script-like files and an uncertain proxy for datasets with supporting software. Because this claim is one of the paper's main contributions, it needs to be either re-analyzed (e.g., by sampling archive contents) or reworded to describe exactly what was measured: the proportion of datasets with at least one directly visible .py, .m, or .R file.","section":"§5 and §1 contribution 2"},{"comment":"The Limitations section acknowledges that the authors 'did not explore the contents of the datasets in detail, relying instead on metadata analysis.' This caveat is directly load-bearing for the headline quantitative claim, yet it appears only in the limitations and is not reflected in the Abstract or Conclusion. The paper should state prominently that the script-provision rate is based on metadata alone and may substantially undercount accompanying software.","section":"§4.3"}],"minor_comments":[{"comment":"'Point Cloud Data (pdc)' should be 'Point Cloud Data (PCD)'.","section":"§3.3"},{"comment":"The respondent quotation 'easily overcomed' should be 'easily overcome' and should be marked as a direct quotation.","section":"§3.8"},{"comment":"Reference [3] truncates the RFC URL to 'rfc70'; it should read 'rfc7012'.","section":"References"},{"comment":"'Zenodos offers' should be 'Zenodo offers'.","section":"§5.1"},{"comment":"The manuscript alternates between 'data set' and 'dataset'; a single spelling would improve readability.","section":"General"},{"comment":"Table 3 would benefit from a note clarifying that cell entries are counts of research questions, and the 'Conditionality' column header appears to correspond to 'Third order' in the text.","section":"Table 3"}],"recommendation":"major_revision","confidential_remarks":"The paper is likely to be acceptable after a revision that properly qualifies or reworks the 2.4% claim. The qualitative contribution and the recommendations are within scope for a software-engineering or open-science venue. The main risk is that the headline number, as currently phrased, overstates the evidence and could mislead readers about open data practices. I would suggest the editor ask the authors to either supplement the metadata analysis with a small archive-content inspection or reframe the claim as a visibility-based statistic rather than a statement about actual supporting software."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper is an experience report on publishing open data in a European industry-academia project, plus a metadata scrape of 280,732 Zenodo records. The genuinely new bit is the angle: nobody has looked specifically at open data publishing inside industry-academia collaborations, and the 280k snapshot is a fresh empirical view. The qualitative findings—planning matters more than technical work, licensing awareness is low, synthetic data has real value—are plausible and well supported by the eleven survey responses, which cover all thirteen InSecTT datasets. The authors also deposited the secondary data, which is a genuine reproducibility plus.\n\nThe soft spot is the headline 2.4% figure. The abstract and conclusion state that 'only few datasets had accompanying scripts' as a measured fact. The body (Section 3.5) does better: it calls this an upper bound based on counting files with .py, .m, or .R extensions. The stress-test is right that this proxy undercounts code packed inside archives—Figure 2 shows most Zenodo files are archives with undisclosed contents—and it ignores Jupyter notebooks, shell scripts, and other languages. So the true rate could be several times 2.4%. The claim is directionally plausible, but the specific number is not established by the scrape as described.\n\nOther weaknesses are minor: the survey is small and self-selected, though the authors say so; the scraping code is not released, which is a lost opportunity given the data is; and the Gartner synthetic-data prediction is used loosely, but that is not load-bearing.\n\nOverall, this is a modest but honest experience report. The qualitative contribution stands on its own, and the Zenodo analysis is a useful preliminary snapshot. It needs a serious referee, but with a requirement to fix the abstract-to-body inconsistency and either inspect a sample of archives or explicitly reframe the 2.4% as a file-extension-level floor on the upper bound. I'd bring it to a reading group as an example of how to do an honest limitations section, and I'd cite the qualitative recommendations if I were writing about data-sharing practices.\n\nRecommendation: send to peer review; expect revision on the quantitative claim.","headline":"The 2.4% script figure is a headline that outruns the evidence; the paper's real worth is the honest qualitative lessons from a small industry-academia collaboration.","tokens_in":17488,"tokens_out":4125,"would_cite":false,"duration_ms":32417,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Only 2.4% of datasets in a major open repository come with reuse scripts.","keywords":["open data","industry-academia collaboration","dataset publication","data reuse","Zenodo metadata","data licences","synthetic data","survey analysis"],"falsifier":"Unpack the archives and inspect the script-like files in a random sample of, say, one thousand of the same repository records, then check whether the scripts actually parse or operate on their dataset; if many .py, .m, or .R files turn out to be unrelated or broken while many archives hide usable scripts, the 2.4% figure and the significant-gap interpretation would not survive.","tokens_in":16536,"feed_emoji":"📊","tokens_out":9318,"duration_ms":82161,"temperature":0.7,"pith_summary":"This paper aims to establish what actually happens when research partnerships between companies and universities publish open datasets, and what can go wrong. Drawing on a survey of eleven contributors covering thirteen datasets in a European industry–academia project, plus metadata from 280,732 records in the generalist repository Zenodo, it argues that planning the data-collection workflow is more challenging than the technical work, that only about 2.4% of repository datasets include scripts that would help others parse or reuse them, and that many authors do not choose licences carefully. The paper also argues that synthetic data, generated by simulations and sometimes mixed with real measurements, can be a legitimate and valuable alternative when real-world data is scarce. If these findings hold, projects that release datasets could improve reuse and citation by planning early, attaching example code, choosing permissive licences, and considering simulated data.","feed_headline":"Just 2.4% of Zenodo datasets include reuse scripts","feed_subtitle":"A survey of an industry-academia project plus 280k repository records shows what helps and what hinders dataset reuse.","key_machinery":"The load-bearing object of the quantitative analysis is a simple proxy: a dataset counts as having accompanying scripts for improved reuse if its metadata lists at least one file whose extension is .py, .m, or .R. Applied to the 280,732 Zenodo records collected through the repository's API, this proxy yields the headline 2.4% figure; the paper treats it as an upper bound because it cannot verify what the files do and cannot see inside archive files. The qualitative side is carried by an inductive analysis of eleven open-ended survey responses, structured additionally by two published taxonomies of research questions, which produced the lessons about planning, licensing, and synthetic data.","core_discovery":"The paper's central claim is that open-data publication in industry–academia collaboration is valuable but often poorly supported. On the qualitative side, eleven survey respondents covering all thirteen datasets of the InSecTT project reported that planning the data-collection workflow—including cleaning, privacy, stakeholder involvement, and documentation—was the dominant challenge, greater than any technical hurdle. On the quantitative side, a scrape of metadata for 280,732 Zenodo datasets from 2000 to 2023 found that only 2.4% had at least one file with a script-like extension (.py, .m, or .R); the authors explicitly call this an upper bound because they cannot confirm that such files are functional supporting software. The paper further reports that most Zenodo datasets use the repository's default CC-BY licence while several InSecTT datasets were published with no licence or with copyleft licences, and argues that synthetic or semi-synthetic data can be highly meaningful, as illustrated by simulated pedestrian-tracking and simulated factory-sensor datasets.","pith_inferences":["The 2.4% figure is probably a lower bound on the true availability of supporting code, because many datasets store code inside archive files whose contents the metadata-only analysis cannot see; checking archive contents would likely raise the measured share.","The paper's advice to pair GitHub with Zenodo suggests that evaluating open data quality from repository metadata alone undercounts actual reuse support, since code often lives in a linked version-control repository rather than beside the data files.","The licence findings might reflect default settings in upload interfaces as much as author ignorance; if so, interface changes that force an explicit licence choice could improve licensing practice faster than guidance documents.","Because the survey covers one project with eleven respondents, the qualitative lessons are best read as candidate practices to test in other industry–academia collaborations, not as proven general laws."],"forward_implications":["A dataset released with even one parsing or example script would belong to a small minority of the repository records the paper examined, so the quantified gap is also an opportunity for differentiation.","Following the recommendation to use both GitHub and Zenodo would give each dataset both a DOI-backed permanent archive and an integration-friendly development home.","Adopting the paper's workflow advice—planning cleaning, privacy, documentation, and stakeholder involvement before collection—would reduce the restructuring effort that respondents reported.","Accepting synthetic and semi-synthetic data would let more industry partners publish useful datasets without privacy violations or costly manual labelling.","Choosing permissive licences such as CC-BY, as the paper recommends, would make downstream reuse legally simpler than copyleft licences do."],"supporting_citations":[{"why":"Establishes the general inductive approach used to code the open-ended survey responses.","marker":"[15]"},{"why":"Identifies patterns and anti-patterns of industry-academia collaboration that this paper extends to open-data publishing.","marker":"[10]"},{"why":"Supplies prior experiences and challenges from industry-academia cyber-physical systems projects that motivate the lessons-learned focus.","marker":"[4]"},{"why":"Provides the first research-question taxonomy used to classify the survey respondents' research questions.","marker":"[12]"},{"why":"Provides the second research-question taxonomy used to classify the survey respondents' research questions.","marker":"[5]"},{"why":"Is the published secondary dataset from the study, making the survey and Zenodo analysis reproducible.","marker":"[14]"},{"why":"Supplies the industry forecast that synthetic data would make up 60% of AI training data by 2024, which the paper's synthetic-data finding engages.","marker":"[8]"},{"why":"Provides the generalist repository comparison chart that underlies the platform analysis and recommendations.","marker":"[13]"}],"fun_headline_variants":["Script-less open data: only 2.4% of Zenodo sets include code","Open data's missing scripts: why planning beats tech hurdles","Synthetic data shines in industry-academia open data study","Licensing blind spot in open data: survey of 280k datasets","Data sharing lessons: plan early, script rarely, license blind"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The quantitative headline depends on treating a file with a script-like extension (.py, .m, or .R) in repository metadata as evidence of a script that genuinely improves reuse; the paper itself notes it cannot verify the files' function and cannot inspect archives, so the true share of reuse-ready datasets could differ.","fun_headline_variants_meta":{"raw":{"variants":["Script-less open data: only 2.4% of Zenodo sets include code","Open data's missing scripts: why planning beats tech hurdles","Synthetic data shines in industry-academia open data study","Licensing blind spot in open data: survey of 280k datasets","Data sharing lessons: plan early, script rarely, license blind"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000168,"raw_usage":{"total_tokens":1240,"prompt_tokens":902,"completion_tokens":338,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":518,"completion_tokens_details":{"reasoning_tokens":247}},"tokens_in":518,"tokens_out":338,"duration_ms":3820,"temperature":1.0,"reasoning_tokens":247,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T15:14:12.004849+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Unpack the archives and inspect the script-like files in a random sample of, say, one thousand of the same repository records, then check whether the scripts actually parse or operate on their dataset; if many .py, .m, or .R files turn out to be unrelated or broken while many archives hide usable scripts, the 2.4% figure and the significant-gap interpretation would not survive.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Identifies patterns and anti-patterns of industry-academia collaboration that this paper extends to open-data publishing."},{"cited_title":"Page 15 of 15","cited_arxiv_id":null,"evidence_quote":"Provides the first research-question taxonomy used to classify the survey respondents' research questions."},{"cited_title":"Where has the data been published, and what was the process for publishing it? 5","cited_arxiv_id":null,"evidence_quote":"Provides the second research-question taxonomy used to classify the survey respondents' research questions."}],"review_version":1}