Pith. sign in

REVIEW 3 major objections 4 minor 77 references

TEDI: Trustworthy and Ethical Dataset Indicators to Analyze and Compare Dataset Documentation

T0 review · 3 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Scraped AI data skips consent and privacy, audit of 114 shows

desk verdict A genuinely useful rubric for dataset documentation with a substantial audit of voice datasets, though the hand-coded labels need reliability evidence before the cross-method pattern is taken as established. read the letter →

arxiv 2505.17841 v1 pith:V6HUZSWP submitted 2025-05-23 cs.CY cs.AIeess.AS

classification cs.CYcs.AIeess.AS
keywords TEDIdatasetdocumentationtransparencymultimodaldatasetsdatacollectionmethodsconsentprivacyaudit
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper introduces TEDI, a rubric of 143 questionnaire-style indicators that extract verifiable information about trustworthy and ethical attributes from dataset documentation. Using it, the authors manually annotated 114 multimodal datasets that include human voices and found that only a few document consent, privacy, or harmful content. The documentation that does exist is patterned by collection method: crowdsourced and direct collection are more likely to mention ethical indicators, while scraped and derived datasets dominate scale but almost never document consent or privacy. The paper claims that data collection methods shape the ethical attributes of datasets, and offers TEDI plus a seven-way taxonomy of collection methods as tools for comparing datasets and eventually automating documentation analysis.

What carries the argument

The carrying objects are TEDI and the data sourcing taxonomy. TEDI is a three-level hierarchy of 143 indicators, each rephrased as a verifiable question answerable with not applicable, no, yes, or yes with evidence or justification, with top-level categories drawn from the Belmont Report and the EU Trustworthy AI Guidelines. The taxonomy sorts collection into seven methods — sampling, direct collection, proprietary, crowdsourced, scraped, derived, synthetic — with 25 subcategories, applied separately to primary modalities, secondary modalities, and annotations. Together they convert free-text dataset documentation into structured, comparable data, and the paper uses them to show how collection method, dataset size, and modality pairings correlate with documented ethical indicators.

What would settle it

Take a random sample of 30 of the 114 datasets and have multiple independent annotators re-apply TEDI at the fine-grained indicator level, reporting agreement per category; if the scraped-versus-crowdsourced gap in documented consent and privacy does not survive stricter coding (e.g., counting only yes with evidence as documented), the central pattern is an artifact of the optimistic any-indicator-yes rule.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central discovery is that data collection methods impact the ethical attributes of datasets, and that this impact is visible in documentation: privacy and consent are largely ignored and remain conspicuously absent from scraped and derived datasets, whereas crowdsourced and direct collections are markedly more likely to document consent, control, compensation, and representation. The finding relies on a deliberately optimistic coding rule — a category counts as documented if any one of its indicators is answered yes — so the true state of documentation is likely worse than reported. The paper also documents that trustworthiness indicators, especially utility and provenance, are far more frequently reported than ethical indicators, and that scraping is the dominant route to datasets larger than 1,000 hours but is not the only viable one.

Load-bearing premise

The annotation is consistent enough across datasets that the differences between collection methods reflect the datasets themselves rather than the coders' judgment.

Editorial extensions

If this is right

  • If TEDI is adopted, dataset documentation can be compared across datasets on a common set of verifiable indicators rather than case-by-case narrative.
  • Dataset creators designing new collections know which indicators their documentation will be judged against, so ethics can be addressed at collection design time.
  • Regulators and auditors can use the indicator set as a checklist for training-data transparency obligations.
  • The empirical pattern implies that scaling by scraping predictably trades away documented consent and privacy, a trade-off that should be made explicit rather than incidental.
  • Automating TEDI-style extraction from documentation becomes a tractable next step because indicators are phrased as answerable questions.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the collection-method pattern generalizes beyond voice datasets, then a dataset's ethical profile can be partially predicted from its sourcing method alone, which would let users triage documentation review effort before reading a single page.
  • A testable extension would be to run the same 143-indicator audit on image-text and text-only corpora; the framework predicts scraped subsets like web-crawled image-text pairs show the same consent and privacy absence.
  • The optimistic coding rule implies that even the low reported rates are upper bounds; a fine-grained re-analysis using the 143 indicators directly would likely find lower coverage, especially for consent revocation and data-worker conditions.
  • TEDI could be turned into a prospective design tool: dataset creators could answer the 143 questions before collection and use negative answers as a checklist of missing safeguards, not just a documentation score.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper introduces TEDI, a three-level hierarchy of 143 fine-grained indicators for assessing trustworthy and ethical aspects of dataset documentation, and applies it by manually annotating 114 human-voice multimodal datasets. It also proposes a two-tier data sourcing taxonomy and labels each dataset's primary, secondary, and annotation modalities according to that taxonomy. The main empirical findings are that trustworthiness indicators such as utility and provenance are documented far more often than ethical indicators; that consent, privacy, and harmful-content indicators are rarely documented; and that documentation patterns vary by collection method, with scraped datasets less likely to mention ethical indicators and crowdsourced/direct collections more likely to do so. The paper argues that these patterns reflect a tension between dataset scale and ethical documentation.

Significance. If the empirical patterns hold, this is a useful contribution: TEDI offers a systematic, verifiable rubric for comparing dataset documentation, and the data sourcing taxonomy provides a vocabulary for discussing collection practices that is currently missing for multimodal speech datasets. The focus on underexplored modalities and the manual annotation of 114 datasets is valuable, and the authors are transparent about the 'most optimistic' category-level coding rule and about the documentation-vs-dataset distinction. The main limitation is methodological: the central cross-method comparison in Figure 3 depends entirely on hand-assigned labels with no reported inter-annotator reliability, and the analysis lacks uncertainty quantification and control for confounds. These issues currently prevent the empirical claim from being fully established.

major comments (3)
  1. [Section 4.3, Figure 3] The central cross-method comparison rests entirely on hand-assigned labels, but no inter-annotator reliability is reported for either the TEDI category-level decisions (aggregated from 143 indicators under the any-indicator-yes rule) or the primary collection-method labels. Without a second coder or an agreement statistic, the reported differences between scraped and crowdsourced/direct datasets could be an artifact of the coding rubric rather than a property of the datasets. Please report inter-annotator agreement (e.g., Cohen's kappa on a subset), provide the annotation protocol or codebook, and run a sensitivity analysis using stricter category-level aggregation rules.
  2. [Table 3, Figure 3] The collection-method categories are not mutually exclusive: the percentages in Table 3 sum to 106% for primary and 130% for annotation modalities, so a single dataset can appear in multiple method columns, yet Figure 3 compares methods as if they were independent groups. In addition, no confidence intervals and no per-cell sample sizes are given, and several categories are very small (e.g., synthetic 2%, proprietary 7%). The paper should report exact counts, handle overlapping membership (e.g., by analyzing only datasets with a single primary method), and add a statistical test or confidence intervals for the cross-method differences.
  3. [Section 3.1, Section 4.2] The selection procedure (citation threshold, snowball sampling, purposive sampling) and the strong association between collection method, dataset size, and release year mean that the observed 'impact' of collection method on ethical indicators may be confounded. For example, large recent scraped datasets and older small direct-collection datasets differ in many ways beyond collection method. The conclusion in Section 4.3 that 'data collection methods impact ethical attributes' uses causal language for what is currently a descriptive association. Please add a robustness check (e.g., stratifying by size or release-year groups) or soften the causal wording so that the headline claim is supported by the evidence presented.
minor comments (4)
  1. [Section 4.3, Section 5] The 'most optimistic' any-indicator-yes rule is acknowledged, but several conclusion sentences (e.g., 'privacy and consent were largely ignored') should consistently read 'not documented' or 'not mentioned' to match TEDI's stated scope of assessing documentation rather than the underlying dataset practices.
  2. [Figure 7 caption] Assuming a 10-second clip duration for datasets where duration is not stated is a strong approximation; please report a sensitivity analysis for Figure 2 or at least state explicitly how sensitive the size-based conclusions are to this assumption.
  3. [Table 3] The table title says 'Proportional', but the percentages do not sum to 100 for any modality; please clarify the denominator, note that a dataset can be assigned multiple collection methods, and explain how the proportions were computed.
  4. [Figure 2 caption] The caption states that data collection methods 'influence' dataset size; this is an observed association and should be phrased as such to avoid implying a causal direction that the data do not establish.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the empirical findings are measured from external dataset documentation, not derived from the framework's own definitions.

full rationale

The paper's central empirical claim is that data collection methods are associated with the extent to which trustworthy and ethical indicators are documented. This is an empirical measurement result: the TEDI indicators are defined from external sources such as the Belmont Report, the EU Trustworthy AI Guidelines, and existing datasheets, and they are applied to dataset documentation that is independent of the framework's construction. The collection-method labels are assigned using a separate data sourcing taxonomy, and no indicator value is computed from a collection-method label, no fitted parameter is used to predict the reported proportions, and no equation in the paper makes the outcome equal to the input by construction. The disclosed 'any indicator yes means category yes' aggregation rule affects sensitivity but does not make the cross-method comparison definitional, and the paper explicitly states that a positive TEDI response reflects documentation availability rather than a judgment about the dataset itself. The only self-citation found is a background citation to author Alice Xiang's prior work on privacy and fairness, which is not load-bearing for the paper's central claims. The lack of reported inter-annotator reliability is a validity and reproducibility concern, not evidence of circularity, because it does not reduce the measured proportions to the authors' prior definitions. The paper is therefore self-contained as an empirical audit: its conclusions could be contradicted by applying the rubric to a different corpus or by independent re-coding of the same corpus.

Assumptions & free parameters 3 free parameters · 4 assumptions · 3 invented entities

The central claims rest on two constructed instruments (the TEDI indicator set and the sourcing taxonomy) plus a manual annotation process. No parametric model is fit, but several hand-set choices shape the results: the citation threshold, the assumed clip duration for size estimates, and the optimistic category-level coding rule. The framework's validity depends on the assumption that documentation mentions are a usable proxy, that the ethical ontology from Belmont and EU guidelines is complete, and that annotators agree.

free parameters (3)
  • citation threshold = >10 Google Scholar citations
    Hand-set selection criterion for datasets from PapersWithCode, Section 3.1. It shapes the corpus composition and excludes many recent datasets.
  • assumed clip duration = 10 seconds
    Used to approximate dataset size in hours for video datasets where only clip counts were reported, as noted in the Figure 7 caption.
  • optimistic category coding = yes if any indicator in the third-level category is yes
    Coding rule for collapsing 143 indicators to category-level analysis, described in Section 4.3. It maximizes positive signals, making reported proportions upper bounds.
assumptions (4)
  • domain assumption Documentation mentions are a usable proxy for dataset attributes
    TEDI converts dataset documentation into an information source for analysis. Section 2 explicitly distinguishes documentation evidence from actual dataset practices, but the empirical claims still rely on documentation as the evidence base.
  • domain assumption Belmont principles and EU Trustworthy AI Guidelines provide a complete ethical ontology
    The 3-level TEDI hierarchy is grounded in these frameworks (Section 2). If these frameworks omit important ethical dimensions, TEDI would systematically miss them.
  • domain assumption Manual annotation across 114 datasets is consistent
    Section 3.1 describes the annotation process but reports no inter-annotator reliability or validation, so consistency is assumed.
  • domain assumption The corpus is representative of multimodal datasets with human voices
    Selection used PapersWithCode keywords, a citation threshold, and snowball and purposive sampling (Section 3.1). The authors acknowledge in Section 5 that the corpus may miss recent datasets and skews toward speech processing.
invented entities (3)
  • TEDI indicator hierarchy
    purpose: Rubric to quantify documentation coverage across 143 indicators in a 3-level hierarchy
    Introduced by the authors as a measurement framework. It has no external falsifiable handle beyond the rubric itself and its application to documentation.
  • Data sourcing taxonomy
    purpose: Classify dataset collection methods into 7 top-level and 25 sub-categories
    Constructed by the authors from literature and corpus observation. Its inter-rater reliability is not measured.
  • Primary/secondary/annotation modality convention
    purpose: Disambiguate how modalities are collected in multimodal datasets
    A definitional convention introduced to support the analysis, not an independently validated construct.

how reviews work

0 comments
Cite this review

Pith. "Pith review of TEDI: Trustworthy and Ethical Dataset Indicators to Analyze and Compare Dataset Documentation." pith.science (2026). https://pith.science/paper/V6HUZSWP

@misc{pith2026250517841,
  author       = {Pith},
  title        = {Pith review of: TEDI: Trustworthy and Ethical Dataset Indicators to Analyze and Compare Dataset Documentation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/V6HUZSWP}},
  note         = {Machine review of arXiv:2505.17841}
}
read the original abstract

Dataset transparency is a key enabler of responsible AI, but insights into multimodal dataset attributes that impact trustworthy and ethical aspects of AI applications remain scarce and are difficult to compare across datasets. To address this challenge, we introduce Trustworthy and Ethical Dataset Indicators (TEDI) that facilitate the systematic, empirical analysis of dataset documentation. TEDI encompasses 143 fine-grained indicators that characterize trustworthy and ethical attributes of multimodal datasets and their collection processes. The indicators are framed to extract verifiable information from dataset documentation. Using TEDI, we manually annotated and analyzed over 100 multimodal datasets that include human voices. We further annotated data sourcing, size, and modality details to gain insights into the factors that shape trustworthy and ethical dimensions across datasets. We find that only a select few datasets have documented attributes and practices pertaining to consent, privacy, and harmful content indicators. The extent to which these and other ethical indicators are addressed varies based on the data collection method, with documentation of datasets collected via crowdsourced and direct collection approaches being more likely to mention them. Scraping dominates scale at the cost of ethical indicators, but is not the only viable collection method. Our approach and empirical insights contribute to increasing dataset transparency along trustworthy and ethical dimensions and pave the way for automating the tedious task of extracting information from dataset documentation in future.

Figures

Figures reproduced from arXiv: 2505.17841 by the authors.

Figure 1
Figure 1. Proportional dataset use per task for datasets from PapersWithCode [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Data collection methods influence the size (i.e. recorded hours) of datasets. [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Proportion of datasets that have considered trustworthy (bottom) and ethical (top) dataset [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figures from the paper (7 more)
Figure 4
Figure 4. Figure 4: Research approach of this study, including dataset selection, annotation, and analysis. [PITH_FULL_IMAGE:figures/full_fig_p017_4.png]
Figure 5
Figure 5. Figure 5: Histograms of detailed data types for video, text and speech for primary, secondary and [PITH_FULL_IMAGE:figures/full_fig_p018_5.png]
Figure 6
Figure 6. Figure 6: Count of papers per speech processing task on PapersWithCode [PITH_FULL_IMAGE:figures/full_fig_p018_6.png]
Figure 7
Figure 7. Figure 7: New dataset releases (bar chart) overlaid by mean dataset size (hours) over time, with error [PITH_FULL_IMAGE:figures/full_fig_p019_7.png]
Figure 8
Figure 8. Figure 8: Primary data collection methods (highlighted in colour) and sources [PITH_FULL_IMAGE:figures/full_fig_p021_8.png]
Figure 9
Figure 9. Figure 9: Secondary data collection methods (highlighted in colour) and sources [PITH_FULL_IMAGE:figures/full_fig_p022_9.png]
Figure 10
Figure 10. Figure 10: Annotation data collection methods (highlighted in colour) and sources [PITH_FULL_IMAGE:figures/full_fig_p023_10.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

77 extracted references · 34 canonical work pages

  1. [1]

    AB-2013 Generative artificial intelligence: training data transparency.en. 2024. (Visited on 05/12/2025)

  2. [2]

    Croissant: A Metadata Format for ML-Ready Datasets

    Mubashara Akhtar et al. “Croissant: A Metadata Format for ML-Ready Datasets”. en. In: Proceedings of the Eighth Workshop on Data Management for End-to-End Machine Learning. Santiago AA Chile: ACM, June 2024, pp. 1–6. ISBN : 9798400706110. DOI: 10 . 1145 / 3650203 . 3663326. URL: https : / / dl . acm . org / doi / 10 . 1145 / 3650203 . 3663326 (visited on ...

  3. [3]

    Ethical Considerations for Responsible Data Curation

    Jerone T A Andrews et al. “Ethical Considerations for Responsible Data Curation”. en. In: 2023

  4. [4]

    Common Voice: A Massively-Multilingual Speech Corpus

    Rosana Ardila et al. Common Voice: A Massively-Multilingual Speech Corpus . arXiv:1912.06670 [cs]. Mar. 2020. DOI: 10 . 48550 / arXiv . 1912 . 06670. URL: http : //arxiv.org/abs/1912.06670 (visited on 04/25/2025)

  5. [5]

    Training Data for the Price of a Sandwich

    Stefan Baack. Training Data for the Price of a Sandwich. en. Tech. rep. Mozilla Foundation,

  6. [6]

    Addressing

    Jack Bandy and Nicholas Vincent. “Addressing "Documentation Debt" in Machine Learning: A Retrospective Datasheet for BookCorpus”. en. In: Sydney, Australia, 2021

  7. [7]

    Principles of biomedical ethics

    Tom Beauchamp and James Childress. Principles of biomedical ethics . Oxford University Press, 2013

  8. [8]

    Multimodal datasets: misog- yny, pornography, and malignant stereotypes

    Abeba Birhane, Vinay Uday Prabhu, and Emmanuel Kahembwe. Multimodal datasets: misog- yny, pornography, and malignant stereotypes. en. arXiv:2110.01963 [cs]. Oct. 2021. URL: http://arxiv.org/abs/2110.01963 (visited on 04/26/2024)

Show all 77 references
  1. [9]

    The Foundation Model Transparency Index v1.1 May 2024

    Rishi Bommasani et al. The Foundation Model Transparency Index v1.1 May 2024. en. 2023

  2. [10]

    Bias and Fairness in Multimodal Machine Learning: A Case Study of Automated Video Interviews

    Brandon M. Booth et al. “Bias and Fairness in Multimodal Machine Learning: A Case Study of Automated Video Interviews”. en. In: Proceedings of the 2021 International Conference on Multimodal Interaction. Montréal QC Canada: ACM, Oct. 2021, pp. 268–277. ISBN : 978- 1-4503-8481-...

  3. [11]

    Scripted dialogs versus improvisation: lessons learned about emotional elicitation techniques from the IEMOCAP database

    Carlos Busso and Shrikanth S. Narayanan. “Scripted dialogs versus improvisation: lessons learned about emotional elicitation techniques from the IEMOCAP database”. en. In: Inter- speech 2008. ISCA, Sept. 2008, pp. 1670–1673. DOI: 10.21437/Interspeech.2008-463. URL: https://www...

  4. [12]

    arXiv:2106.06909 [cs]

    Guoguo Chen et al.GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio. arXiv:2106.06909 [cs]. June 2021. DOI: 10.48550/arXiv.2106.06909. URL: http://arxiv.org/abs/2106.06909 (visited on 04/25/2025)

  5. [13]

    Chmielinski et al

    Kasia S. Chmielinski et al. The Dataset Nutrition Label (2nd Gen): Leveraging Context to Mitigate Harms in Artificial Intelligence . en. arXiv:2201.03954 [cs]. Mar. 2022. DOI: 10.48550/arXiv.2201.03954 . URL: http://arxiv.org/abs/2201.03954 (visited on 04/06/2025)

  6. [14]

    V oxCeleb2: Deep speaker recogni- tion

    Joon Son Chung, Arsha Nagrani, and Andrew Zisserman. “V oxCeleb2: Deep speaker recogni- tion”. In: arXiv ii (2018). ISSN : 23318422

  7. [15]

    FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech

    Alexis Conneau et al. FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech. arXiv:2205.12446 [cs]. May 2022. DOI: 10.48550/arXiv.2205.12446 . URL: http://arxiv.org/abs/2205.12446 (visited on 04/28/2025)

  8. [16]

    Data Equity: Foundational Concepts for Generative AI. Tech. rep. World Economic Forum, Oct. 2023. URL: https://www3.weforum.org/docs/WEF_Data_Equity_Concepts_ Generative_AI_2023.pdf (visited on 04/07/2025)

  9. [17]

    ImageNet: A large-scale hierarchical image database

    Jia Deng et al. “ImageNet: A large-scale hierarchical image database”. In: 2009 IEEE Confer- ence on Computer Vision and Pattern Recognition. ISSN: 1063-6919. June 2009, pp. 248–255. DOI: 10.1109/CVPR.2009.5206848. URL: https://ieeexplore.ieee.org/document/ 5206848 (visited on...

  10. [18]

    CrowdWorkSheets: Accounting for Individual and Collective Identities Underlying Crowdsourced Dataset Annotation

    Mark Diaz et al. “CrowdWorkSheets: Accounting for Individual and Collective Identities Underlying Crowdsourced Dataset Annotation”. en. In: 2022 ACM Conference on Fairness, Accountability, and Transparency. arXiv:2206.08931 [cs]. June 2022, pp. 2342–2351. DOI: 10.1145/3531146....

  11. [19]

    Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus

    Jesse Dodge et al. Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus. arXiv:2104.08758 [cs]. Sept. 2021. DOI: 10.48550/arXiv.2104.08758. URL: http://arxiv.org/abs/2104.08758 (visited on 05/12/2025)

  12. [20]

    What’s In My Big Data? en

    Yanai Elazar et al. What’s In My Big Data? en. arXiv:2310.20707 [cs]. Mar. 2024. DOI: 10.48550/arXiv.2310.20707 . URL: http://arxiv.org/abs/2310.20707 (visited on 05/12/2025)

  13. [21]

    Pages: 1-39 Publication Title: European Commission

    Ethics Guidelines for Trustworthy AI. Pages: 1-39 Publication Title: European Commission

  14. [22]

    Uncurated Image-Text Datasets: Shedding Light on Demographic Bias

    Noa Garcia et al. Uncurated Image-Text Datasets: Shedding Light on Demographic Bias . en. arXiv:2304.02828 [cs]. Apr. 2023. DOI: 10.48550/arXiv.2304.02828 . URL: http: //arxiv.org/abs/2304.02828 (visited on 05/12/2025)

  15. [23]

    Datasheets for datasets

    Timnit Gebru et al. “Datasheets for datasets”. In: Communications of the ACM 64.12 (2021). arXiv: 1803.09010, pp. 86–92. ISSN : 15577317. DOI: 10.1145/3458723

  16. [24]

    Audio Set: An ontology and human-labeled dataset for audio events

    Jort F. Gemmeke et al. “Audio Set: An ontology and human-labeled dataset for audio events”. In: 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). ISSN: 2379-190X. Mar. 2017, pp. 776–780. DOI: 10.1109/ICASSP.2017.7952261. URL: https://ieeex...

  17. [25]

    Ego4D: Around the World in 3,000 Hours of Egocentric Video

    Kristen Grauman et al. Ego4D: Around the World in 3,000 Hours of Egocentric Video . arXiv:2110.07058 [cs]. Mar. 2022. DOI: 10 . 48550 / arXiv . 2110 . 07058. URL: http : //arxiv.org/abs/2110.07058 (visited on 04/25/2025)

  18. [26]

    Speaker recognition by machines and humans: A tutorial review

    John H. L. Hansen and Taufiq Hasan. “Speaker recognition by machines and humans: A tutorial review”. In: IEEE Signal Processing Magazine 32.6 (2015), pp. 74–99. ISSN : 10535888. DOI: 10.1109/MSP.2015.2462851

  19. [27]

    ActivityNet: A large-scale video benchmark for human activity understanding

    Fabian Caba Heilbron et al. “ActivityNet: A large-scale video benchmark for human activity understanding”. In: 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). ISSN: 1063-6919. June 2015, pp. 961–970. DOI: 10.1109/CVPR.2015.7298698. URL: https://ieeexplo...

  20. [28]

    Design Research in Information Systems

    Alan Hevner and Samir Chatterjee. Design Research in Information Systems. Ed. by Ramesh Sharda and Stefan V oß. Springer, 2010.ISBN : 978-1-4419-5652-1. DOI: 10.1007/978-1- 4419-5653-8. URL: http://www.springer.com/series/6157

  21. [29]

    URL: https://partnershiponai

    Improving Conditions for Data Enrichment Workers. URL: https://partnershiponai. org/responsible-sourcing-library/ (visited on 04/07/2025)

  22. [30]

    A Standardized Machine-readable Dataset Documentation Format for Responsible AI

    Nitisha Jain et al. A Standardized Machine-readable Dataset Documentation Format for Responsible AI. en. arXiv:2407.16883 [cs]. June 2024. URL: http://arxiv.org/abs/2407. 16883 (visited on 08/21/2024)

  23. [31]

    Libri-Light: A Benchmark for ASR with Limited or No Supervision

    Jacob Kahn et al. “Libri-Light: A Benchmark for ASR with Limited or No Supervision”. In: ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). arXiv:1912.07875 [cs]. May 2020, pp. 7669–7673. DOI: 10.1109/ ICASSP40776.2020.9052942...

  24. [32]

    Libriheavy: a 50,000 hours ASR corpus with punctuation casing and context

    Wei Kang et al. Libriheavy: a 50,000 hours ASR corpus with punctuation casing and context. arXiv:2309.08105 [eess]. Jan. 2024. DOI: 10 . 48550 / arXiv . 2309 . 08105. URL: http : //arxiv.org/abs/2309.08105 (visited on 04/28/2025)

  25. [33]

    Golos: Russian Dataset for Speech Research

    Nikolay Karpov, Alexander Denisenko, and Fedor Minkin. Golos: Russian Dataset for Speech Research. arXiv:2106.10161 [eess]. June 2021. DOI: 10.48550/arXiv.2106.10161. URL: http://arxiv.org/abs/2106.10161 (visited on 04/28/2025)

  26. [34]

    A Data Perspective on Ethical Challenges in V oice Biometrics Research

    Anna Leschanowsky et al. “A Data Perspective on Ethical Challenges in V oice Biometrics Research”. In: IEEE Transactions on Biometrics, Behavior, and Identity Science 7.1 (Jan. 2025), pp. 118–131. ISSN : 2637-6407. DOI: 10.1109/TBIOM.2024.3446846. URL: https: //ieeexplore.ieee...

  27. [35]

    YODAS: Youtube-Oriented Dataset for Audio and Speech

    Xinjian Li et al. YODAS: Youtube-Oriented Dataset for Audio and Speech . en. arXiv:2406.00899 [cs]. June 2024. DOI: 10 . 48550 / arXiv . 2406 . 00899. URL: http : //arxiv.org/abs/2406.00899 (visited on 05/13/2025)

  28. [36]

    Foundations & Trends in Multi- modal Machine Learning: Principles, Challenges, and Open Questions

    Paul Pu Liang, Amir Zadeh, and Louis-Philippe Morency. “Foundations & Trends in Multi- modal Machine Learning: Principles, Challenges, and Open Questions”. en. In: ACM Comput- ing Surveys (Apr. 2024), p. 3656580. ISSN : 0360-0300, 1557-7341. DOI: 10.1145/3656580. URL: https://...

  29. [37]

    Bridging the Data Provenance Gap Across Text, Speech and Video

    Shayne Longpre et al. Bridging the Data Provenance Gap Across Text, Speech and Video . en. arXiv:2412.17847 [cs]. Feb. 2025. DOI: 10.48550/arXiv.2412.17847 . URL: http: //arxiv.org/abs/2412.17847 (visited on 04/16/2025)

  30. [38]

    Data Authenticity, Consent, & Provenance for AI are all broken: what will it take to fix them? en

    Shayne Longpre et al. Data Authenticity, Consent, & Provenance for AI are all broken: what will it take to fix them? en. arXiv:2404.12691 [cs]. Apr. 2024. URL: http://arxiv.org/ abs/2404.12691 (visited on 05/16/2024)

  31. [39]

    The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI

    Shayne Longpre et al. The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI. en. arXiv:2310.16787 [cs]. Nov. 2023. URL: http://arxiv. org/abs/2310.16787 (visited on 05/07/2024)

  32. [40]

    Data Statements: From Technical Concept to Community Practice

    Angelina McMillan-Major, Emily M. Bender, and Batya Friedman. “Data Statements: From Technical Concept to Community Practice”. en. In: ACM Journal on Responsible Computing 1.1 (Mar. 2024), pp. 1–17. ISSN : 2832-0565. DOI: 10.1145/3594737 . URL: https://dl. acm.org/doi/10.1145/...

  33. [41]

    Gender Artifacts in Visual Datasets

    Nicole Meister et al. Gender Artifacts in Visual Datasets. en. arXiv:2206.09191 [cs]. Sept

  34. [42]

    Documenting Computer Vision Datasets: An Invitation to Reflexive Data Practices

    Milagros Miceli et al. “Documenting Computer Vision Datasets: An Invitation to Reflexive Data Practices”. In: ACM Fairness Accountability and Transparency (FAccT) ’21 . ACM, 2021, p. 12. ISBN : 978-1-4503-8309-7. DOI: 10.1145/3442188.3445880 . URL: https: //doi.org/10.1145/344...

  35. [43]

    Measuring Data

    Margaret Mitchell et al. Measuring Data. en. arXiv:2212.05129 [cs]. Feb. 2023. URL: http: //arxiv.org/abs/2212.05129 (visited on 03/12/2024)

  36. [44]

    On responsible machine learning datasets emphasizing fairness, privacy and regulatory norms with examples in biometrics and healthcare

    Surbhi Mittal et al. “On responsible machine learning datasets emphasizing fairness, privacy and regulatory norms with examples in biometrics and healthcare”. en. In: Nature Machine Intelligence (Aug. 2024). Publisher: Nature Publishing Group, pp. 1–14.ISSN : 2522-5839. DOI: 1...

  37. [45]

    Data Collection in Music Generation Training Sets: A Critical Analysis

    Fabio Morreale, Megha Sharma, and I-Chieh Wei. “Data Collection in Music Generation Training Sets: A Critical Analysis”. en. In: Milan, Italy, 2023

  38. [46]

    V oxceleb: A large-scale speaker identification dataset

    Arsha Nagrani, Joon Son Chung, and Andrew Zisserman. “V oxceleb: A large-scale speaker identification dataset”. In: arXiv (2017), pp. 2616–2620. ISSN : 23318422

  39. [47]

    DMLR: Data-centric Machine Learning Research – Past, Present and Future

    Luis Oala et al. DMLR: Data-centric Machine Learning Research – Past, Present and Future. en. arXiv:2311.13028 [cs, eess]. Nov. 2023. URL: http://arxiv.org/abs/2311.13028 (visited on 05/06/2024)

  40. [48]

    Librispeech: An ASR corpus based on public domain audio books

    Vassil Panayotov et al. “Librispeech: An ASR corpus based on public domain audio books”. In: ICASSP , IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings 2015-Augus (2015). Publisher: IEEE ISBN: 9781467369978, pp. 5206–5210. ISSN : 15206149. ...

  41. [49]

    MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations

    Soujanya Poria et al. MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations. en. arXiv:1810.02508 [cs]. June 2019. DOI: 10.48550/arXiv.1810.02508. URL: http://arxiv.org/abs/1810.02508 (visited on 04/25/2025)

  42. [50]

    Large image datasets: A pyrrhic win for computer vision? en

    Vinay Uday Prabhu and Abeba Birhane. Large image datasets: A pyrrhic win for computer vision? en. arXiv:2006.16923 [cs]. July 2020. DOI: 10.48550/arXiv.2006.16923 . URL: http://arxiv.org/abs/2006.16923 (visited on 05/12/2025)

  43. [51]

    MLS: A Large-Scale Multilingual Dataset for Speech Research

    Vineel Pratap et al. “MLS: A Large-Scale Multilingual Dataset for Speech Research”. en. In: Interspeech 2020. arXiv:2012.03411 [eess]. Oct. 2020, pp. 2757–2761. DOI: 10.21437/ Interspeech . 2020 - 2826. URL: http : / / arxiv . org / abs / 2012 . 03411(visited on 04/25/2025). 12

  44. [52]

    Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI

    Mahima Pushkarna, Andrew Zaldivar, and Oddur Kjartansson. “Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI”. en. In: 2022 ACM Conference on Fairness, Accountability, and Transparency . Seoul Republic of Korea: ACM, June 2022, pp. 1776–1826. ISBN...

  45. [53]

    Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 laying down harmonised rules on artificial intelligence and amending Regulations (EC) No 300/2008, (EU) No 167/2013, (EU) No 168/2013, (EU) 2018/858, (EU) 2018/1139 and (EU) 2019/2144 and D...

  46. [54]

    Open Datasheets: Machine-readable Documentation for Open Datasets and Responsible AI Assessments

    Anthony Cintron Roman et al. Open Datasheets: Machine-readable Documentation for Open Datasets and Responsible AI Assessments. en. arXiv:2312.06153 [cs]. Mar. 2024. DOI: 10. 48550/arXiv.2312.06153 . URL: http://arxiv.org/abs/2312.06153 (visited on 04/06/2025)

  47. [55]

    Measuring Social Biases in Grounded Vision and Language Embeddings

    Candace Ross, Boris Katz, and Andrei Barbu. “Measuring Social Biases in Grounded Vision and Language Embeddings”. en. In:Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Online: Asso...

  48. [56]

    About V oice: A Longitudinal Study of Speaker Recognition Dataset Dynamics

    Casandra Rusti et al. “About V oice: A Longitudinal Study of Speaker Recognition Dataset Dynamics”. In: (Apr. 2023). arXiv: 2304.03858. URL: http://arxiv.org/abs/2304. 03858

  49. [57]

    Second Draft of the General-Purpose AI Code of Practice. en. 2025. URL: https://digital- strategy.ec.europa.eu/en/library/second-draft-general-purpose-ai-code- practice-published-written-independent-experts (visited on 05/12/2025)

  50. [58]

    Representation Bias in Data: A Survey on Identification and Resolution Techniques

    Nima Shahbazi et al. “Representation Bias in Data: A Survey on Identification and Resolution Techniques”. en. In: ACM Computing Surveys 55.13s (Dec. 2023), pp. 1–39. ISSN : 0360-0300, 1557-7341. DOI: 10.1145/3588433. URL: https://dl.acm.org/doi/10.1145/3588433 (visited on 03/19/2024)

  51. [59]

    The Great Scrape: The Clash Between Scraping and Privacy

    Daniel J. Solove and Woodrow Hartzog. “The Great Scrape: The Clash Between Scraping and Privacy”. en. In: SSRN Electronic Journal (2024). ISSN : 1556-5068. DOI: 10.2139/ssrn. 4884485. URL: https://www.ssrn.com/abstract=4884485 (visited on 07/19/2024)

  52. [60]

    Mitigating Gender Bias in Captioning Systems

    Ruixiang Tang et al. “Mitigating Gender Bias in Captioning Systems”. en. In: Proceedings of the Web Conference 2021. Ljubljana Slovenia: ACM, Apr. 2021, pp. 633–645. ISBN : 978- 1-4503-8312-7. DOI: 10.1145/3442381.3449950. URL: https://dl.acm.org/doi/10. 1145/3442381.3449950 (...

  53. [61]

    The Belmont Report. en. Tech. rep. US Department of Health, Education, and Welfare, 1979. URL: https://www.hhs.gov/ohrp/sites/default/files/the- belmont- report- 508c_FINAL.pdf

  54. [62]

    URL: https://keithito.com/LJ-Speech-Dataset (visited on 04/25/2025)

    The LJ Speech Dataset. URL: https://keithito.com/LJ-Speech-Dataset (visited on 04/25/2025)

  55. [63]

    Third Draft of the General-Purpose AI Code of Practice. en. 2025. URL: https://digital- strategy.ec.europa.eu/en/library/third-draft-general-purpose-ai-code- practice-published-written-independent-experts (visited on 05/12/2025)

  56. [64]

    YFCC100M: The New Data in Multimedia Research

    Bart Thomee et al. “YFCC100M: The New Data in Multimedia Research”. In:Communications of the ACM 59.2 (Jan. 2016). arXiv:1503.01817 [cs], pp. 64–73. ISSN : 0001-0782, 1557-7317. DOI: 10 . 1145 / 2812802. URL: http : / / arxiv . org / abs / 1503 . 01817(visited on 04/25/2025)

  57. [65]

    Interim Measures for the Management of Generative Artificial In- telligence Services

    China Law Translate. Interim Measures for the Management of Generative Artificial In- telligence Services . en. July 2023. URL: https : / / www . chinalawtranslate . com / generative-ai-interim/ (visited on 05/12/2025). 13

  58. [66]

    V oxPopuli: A Large-Scale Multilingual Speech Corpus for Represen- tation Learning, Semi-Supervised Learning and Interpretation

    Changhan Wang et al. “V oxPopuli: A Large-Scale Multilingual Speech Corpus for Represen- tation Learning, Semi-Supervised Learning and Interpretation”. In: Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint ...

  59. [67]

    Contrastive Language-Vision AI Models Pretrained on Web-Scraped Multimodal Data Exhibit Sexual Objectification Bias

    Robert Wolfe et al. “Contrastive Language-Vision AI Models Pretrained on Web-Scraped Multimodal Data Exhibit Sexual Objectification Bias”. en. In: 2023 ACM Conference on Fairness, Accountability, and Transparency. Chicago IL USA: ACM, June 2023, pp. 1174–

  60. [68]

    Sierra Wyllie, Ilia Shumailov, and Nicolas Papernot.Fairness Feedback Loops: Training on Synthetic Data Amplifies Bias. en. arXiv:2403.07857 [cs]. Mar. 2024. URL: http://arxiv. org/abs/2403.07857 (visited on 04/30/2024)

  61. [69]

    Being ’Seen’ vs

    Alice Xiang. Being ’Seen’ vs. ’Mis-Seen’: Tensions between Privacy and Fairness in Computer Vision. en. SSRN Scholarly Paper. Rochester, NY, Feb. 2022. DOI: 10.2139/ssrn.4068921. URL: https://papers.ssrn.com/abstract=4068921 (visited on 05/12/2025)

  62. [70]

    CSTR VCTK Corpus: English Multi-speaker Corpus for CSTR V oice Cloning Toolkit (version 0.92)

    Junichi Yamagishi, Christophe Veaux, and Kirsten MacDonald. “CSTR VCTK Corpus: English Multi-speaker Corpus for CSTR V oice Cloning Toolkit (version 0.92)”. eng. In:The Rainbow Passage which the speakers read out can be found in the International Dialects of English Archive: (...

  63. [71]

    Navigating Dataset Documentations in AI: A Large-Scale Analysis of Dataset Cards on Hugging Face

    Xinyu Yang, Weixin Liang, and James Zou. Navigating Dataset Documentations in AI: A Large-Scale Analysis of Dataset Cards on Hugging Face . en. arXiv:2401.13822 [cs]. Jan

  64. [72]

    Position: Measure Dataset Diversity, Don’t Just Claim It

    Dora Zhao et al. “Position: Measure Dataset Diversity, Don’t Just Claim It”. en. In:Proceedings of the 41 st International Conference on Machine Learning. Vienna, Austria, 2024. 14 A Appendix A.1 Supporting Material: Trustworthy and Ethical Dataset Indicators Table 4: Detailed...

  65. [76]

    URL: http://arxiv.org/abs/2401.13822 (visited on 04/06/2025)

    DOI: 10.48550/arXiv.2401.13822. URL: http://arxiv.org/abs/2401.13822 (visited on 04/06/2025)

  66. [1185]

    DOI: 10.1145/3593013.3594072

    ISBN : 9798400701924. DOI: 10.1145/3593013.3594072 . URL: https://dl.acm. org/doi/10.1145/3593013.3594072 (visited on 04/17/2024)

  67. [2019]

    URL: https : / / digital - strategy . ec . europa . eu / en / library / ethics - guidelines-trustworthy-ai

  68. [2023]

    URL: http://arxiv.org/abs/2206.09191 (visited on 05/12/2025)

    DOI: 10.48550/arXiv.2206.09191. URL: http://arxiv.org/abs/2206.09191 (visited on 05/12/2025)

  69. [2024]

    URL: https://foundation.mozilla.org/en/research/library/generative- ai-training-data/common-crawl/

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.