REVIEW 3 major objections 6 minor 8 references
Insights from Publishing Open Data in Industry-Academia Collaboration
T0 review · 3 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Only 2.4% of datasets in a major open repository come with reuse scripts.
desk verdict The 2.4% script figure is a headline that outruns the evidence; the paper's real worth is the honest qualitative lessons from a small industry-academia collaboration. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object of the quantitative analysis is a simple proxy: a dataset counts as having accompanying scripts for improved reuse if its metadata lists at least one file whose extension is .py, .m, or .R. Applied to the 280,732 Zenodo records collected through the repository's API, this proxy yields the headline 2.4% figure; the paper treats it as an upper bound because it cannot verify what the files do and cannot see inside archive files. The qualitative side is carried by an inductive analysis of eleven open-ended survey responses, structured additionally by two published taxonomies of research questions, which produced the lessons about planning, licensing, and synthetic data.
What would settle it
Unpack the archives and inspect the script-like files in a random sample of, say, one thousand of the same repository records, then check whether the scripts actually parse or operate on their dataset; if many .py, .m, or .R files turn out to be unrelated or broken while many archives hide usable scripts, the 2.4% figure and the significant-gap interpretation would not survive.
Extended reading notes
Core claim
The paper's central claim is that open-data publication in industry–academia collaboration is valuable but often poorly supported. On the qualitative side, eleven survey respondents covering all thirteen datasets of the InSecTT project reported that planning the data-collection workflow—including cleaning, privacy, stakeholder involvement, and documentation—was the dominant challenge, greater than any technical hurdle. On the quantitative side, a scrape of metadata for 280,732 Zenodo datasets from 2000 to 2023 found that only 2.4% had at least one file with a script-like extension (.py, .m, or .R); the authors explicitly call this an upper bound because they cannot confirm that such files are functional supporting software. The paper further reports that most Zenodo datasets use the repository's default CC-BY licence while several InSecTT datasets were published with no licence or with copyleft licences, and argues that synthetic or semi-synthetic data can be highly meaningful, as illustrated by simulated pedestrian-tracking and simulated factory-sensor datasets.
Load-bearing premise
The quantitative headline depends on treating a file with a script-like extension (.py, .m, or .R) in repository metadata as evidence of a script that genuinely improves reuse; the paper itself notes it cannot verify the files' function and cannot inspect archives, so the true share of reuse-ready datasets could differ.
Editorial extensions
If this is right
- A dataset released with even one parsing or example script would belong to a small minority of the repository records the paper examined, so the quantified gap is also an opportunity for differentiation.
- Following the recommendation to use both GitHub and Zenodo would give each dataset both a DOI-backed permanent archive and an integration-friendly development home.
- Adopting the paper's workflow advice—planning cleaning, privacy, documentation, and stakeholder involvement before collection—would reduce the restructuring effort that respondents reported.
- Accepting synthetic and semi-synthetic data would let more industry partners publish useful datasets without privacy violations or costly manual labelling.
- Choosing permissive licences such as CC-BY, as the paper recommends, would make downstream reuse legally simpler than copyleft licences do.
Reading between the lines
- The 2.4% figure is probably a lower bound on the true availability of supporting code, because many datasets store code inside archive files whose contents the metadata-only analysis cannot see; checking archive contents would likely raise the measured share.
- The paper's advice to pair GitHub with Zenodo suggests that evaluating open data quality from repository metadata alone undercounts actual reuse support, since code often lives in a linked version-control repository rather than beside the data files.
- The licence findings might reflect default settings in upload interfaces as much as author ignorance; if so, interface changes that force an explicit licence choice could improve licensing practice faster than guidance documents.
- Because the survey covers one project with eleven respondents, the qualitative lessons are best read as candidate practices to test in other industry–academia collaborations, not as proven general laws.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reports on open data publication practices in the InSecTT industry-academia collaboration, combining an open-ended survey of 11 project participants (covering 13 datasets) with a metadata analysis of 280,732 Zenodo datasets up to 2023. The authors identify qualitative lessons about planning, documentation, licensing, and synthetic data, and they quantify how often Zenodo datasets are accompanied by script-like files, reporting a rate of about 2.4%. The paper closes with a set of recommendations for publishing open datasets in industry-academia projects.
Significance. The topic is timely and the paper brings together a rare inside view of an industry-academia project with a large public-repository metadata sample. The authors deposit their secondary data at Zenodo [14], which supports reproducibility of the metadata analysis and the survey responses. The qualitative findings about planning, licences, and synthetic data are plausible and useful even with the small sample. However, the central quantitative claim—that only 2.4% of Zenodo datasets have accompanying scripts for improved reuse—rests on a file-extension proxy that the paper itself acknowledges to be incomplete, and the Abstract and Conclusion state the claim without the required caveats.
major comments (3)
- [§3.5, Figure 2, Abstract] The 2.4% figure counts a dataset as accompanied by a script only if the Zenodo file-level metadata contains a file with extension .py, .m, or .R. Since §3.3 reports that the majority of published files are archives whose contents remain undisclosed, scripts bundled inside .zip or .tar.gz files are invisible to this count, as are Jupyter notebooks and scripts in other languages. Consequently the paper's own 'upper bound' characterization in §3.5 applies only to whether identified scripts truly function as supporting software; it does not address false negatives from archive contents. The Abstract's statement that 'only few datasets (2.4%) had accompanying scripts for improved reuse' is therefore a factual claim that the described method cannot support.
- [§5 and §1 contribution 2] The Conclusion repeats the unsupported interpretation as 'very few Zenodo datasets (less than 2.4%) were accompanied with scripts that would improve reuse.' The phrase 'less than 2.4%' is justified neither by the measurement nor by the authors' own caveats; the measurement is a lower bound on datasets with directly visible script-like files and an uncertain proxy for datasets with supporting software. Because this claim is one of the paper's main contributions, it needs to be either re-analyzed (e.g., by sampling archive contents) or reworded to describe exactly what was measured: the proportion of datasets with at least one directly visible .py, .m, or .R file.
- [§4.3] The Limitations section acknowledges that the authors 'did not explore the contents of the datasets in detail, relying instead on metadata analysis.' This caveat is directly load-bearing for the headline quantitative claim, yet it appears only in the limitations and is not reflected in the Abstract or Conclusion. The paper should state prominently that the script-provision rate is based on metadata alone and may substantially undercount accompanying software.
minor comments (6)
- [§3.3] 'Point Cloud Data (pdc)' should be 'Point Cloud Data (PCD)'.
- [§3.8] The respondent quotation 'easily overcomed' should be 'easily overcome' and should be marked as a direct quotation.
- [References] Reference [3] truncates the RFC URL to 'rfc70'; it should read 'rfc7012'.
- [§5.1] 'Zenodos offers' should be 'Zenodo offers'.
- [General] The manuscript alternates between 'data set' and 'dataset'; a single spelling would improve readability.
- [Table 3] Table 3 would benefit from a note clarifying that cell entries are counts of research questions, and the 'Conditionality' column header appears to correspond to 'Third order' in the text.
Circularity Check
No circularity found; the paper's claims are inductive empirical findings rather than derived results, and the only self-citations are not load-bearing.
full rationale
The paper does not present a mathematical derivation, fitted model, or uniqueness argument that could reduce to its own inputs. Its central quantitative claim, that approximately 2.4% of Zenodo datasets contain files identified as scripts, is an inductive count from scraped metadata with an explicitly acknowledged proxy limitation: the authors state that they can only establish an upper bound because they cannot verify what the scripts do. This is a measurement-validity concern, not circularity. The survey conclusions are qualitative interpretations of open-ended responses, and the Gartner citation about synthetic data is an external prediction that the paper uses for comparison, not a premise fitted to its own data. The only self-citations are the deposited secondary data at Zenodo [14] and a prior co-authored industry-academia collaboration study [4]; neither carries the load of the paper's conclusions, and the prior study is cited merely as related work. No circular step can be exhibited by quoting equations or showing that an output is equivalent to an input by construction. The paper is self-contained as an empirical study, so the appropriate circularity score is 0.
Assumptions & free parameters
assumptions (4)
- domain assumption Survey respondents' self-reported descriptions of data collection, publication, and lessons are accurate.
- domain assumption File extensions (.py, .m, .R) are a valid proxy for identifying accompanying scripts in Zenodo datasets.
- domain assumption Zenodo metadata is representative of open dataset publication practices generally.
- domain assumption Gartner's prediction that 60% of AI data will be synthetic by 2024 is a reliable external benchmark.
Cite this review
Pith. "Pith review of Insights from Publishing Open Data in Industry-Academia Collaboration." pith.science (2026). https://pith.science/paper/HO2KI7DQ
@misc{pith2026250114841,
author = {Pith},
title = {Pith review of: Insights from Publishing Open Data in Industry-Academia Collaboration},
year = {2026},
howpublished = {\url{https://pith.science/paper/HO2KI7DQ}},
note = {Machine review of arXiv:2501.14841}
}
read the original abstract
Effective data management and sharing are critical success factors in industry-academia collaboration. This paper explores the motivations and lessons learned from publishing open data sets in such collaborations. Through a survey of participants in a European research project that published 13 data sets, and an analysis of metadata from almost 281 thousand datasets in Zenodo, we collected qualitative and quantitative results on motivations, achievements, research questions, licences and file types. Through inductive reasoning and statistical analysis we found that planning the data collection is essential, and that only few datasets (2.4%) had accompanying scripts for improved reuse. We also found that authors are not well aware of the importance of licences or which licence to choose. Finally, we found that data with a synthetic origin, collected with simulations and potentially mixed with real measurements, can be very meaningful, as predicted by Gartner and illustrated by many datasets collected in our research project.
Reference graph
Works this paper leans on
-
[1]
Information Model for IP Flow Information Export (IPFIX),
Arthur, C. (2013). Tech giants may be huge, but nothing matches big data. The Guardian 23 Aug 2013. Online: https://www.theguardian.com/technology/2013/aug/23/tech-giants-data [2] Bansal, M. A., Sharma, D. R., & Kathuria, D. M. (2022). A systematic review on data scarcity problem in deep learning: solution and applications. ACM Computing Surveys (CSUR), 5...
-
[2]
How long is your working experience? (1-5 years, 6-10 years, 11-15 years, 16-20 years, 21-25 years, 26-30 years, or Over 30 years)
-
[3]
Which type of organization are you working in? (Small or medium enterprise, Large enterprise, University, Research institute, Public sector or Other)
-
[5]
Where has the data been published, and what was the process for publishing it? 5
Which InSecTT dataset did you contribute to? 4. Where has the data been published, and what was the process for publishing it? 5. What has been the original purpose of the dataset? 6. What has been the main motivation to publish the dataset? 7. Which have been the research questions or other challenge targets studied with the dataset? 8. What was the type...
-
[6]
Economist (Anonymous author). (2017). The world’s most valuable resource is no longer oil, but data. The Economist. 6 May 2017. Online: https://www.economist.com/leaders/2017/05/06/the-worlds-most-valuable-resource-is-no-longer-oil-but-data [7] Garousi, V., Pfahl, D., Fernandes, J. M., Felderer, M., Mäntylä, M. V., Shepherd, D., ... & Tekinerdogan, B. (20...
arXiv 2017
-
[10]
Have you published supporting software or tools with your dataset? If yes, please describe in short
-
[11]
Which are the main achievements and results sprung up from the dataset so far? These include, but are not limited to, publications, commercialization, and exploitation plans
-
[12]
Have you learnt some specific lessons during the data collection process? For example, share your experiences (in brief) about encouraging the stakeholder(s) to get involved , ways/workflows in solving encountered technical challenges, discovered best practices on data management/cleaning/processing/anonymization, viable solutions applied on ethical chall...
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.