REVIEW 3 major objections 4 minor 77 references
TEDI: Trustworthy and Ethical Dataset Indicators to Analyze and Compare Dataset Documentation
T0 review · 3 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Scraped AI data skips consent and privacy, audit of 114 shows
desk verdict A genuinely useful rubric for dataset documentation with a substantial audit of voice datasets, though the hand-coded labels need reliability evidence before the cross-method pattern is taken as established. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying objects are TEDI and the data sourcing taxonomy. TEDI is a three-level hierarchy of 143 indicators, each rephrased as a verifiable question answerable with not applicable, no, yes, or yes with evidence or justification, with top-level categories drawn from the Belmont Report and the EU Trustworthy AI Guidelines. The taxonomy sorts collection into seven methods — sampling, direct collection, proprietary, crowdsourced, scraped, derived, synthetic — with 25 subcategories, applied separately to primary modalities, secondary modalities, and annotations. Together they convert free-text dataset documentation into structured, comparable data, and the paper uses them to show how collection method, dataset size, and modality pairings correlate with documented ethical indicators.
What would settle it
Take a random sample of 30 of the 114 datasets and have multiple independent annotators re-apply TEDI at the fine-grained indicator level, reporting agreement per category; if the scraped-versus-crowdsourced gap in documented consent and privacy does not survive stricter coding (e.g., counting only yes with evidence as documented), the central pattern is an artifact of the optimistic any-indicator-yes rule.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that data collection methods impact the ethical attributes of datasets, and that this impact is visible in documentation: privacy and consent are largely ignored and remain conspicuously absent from scraped and derived datasets, whereas crowdsourced and direct collections are markedly more likely to document consent, control, compensation, and representation. The finding relies on a deliberately optimistic coding rule — a category counts as documented if any one of its indicators is answered yes — so the true state of documentation is likely worse than reported. The paper also documents that trustworthiness indicators, especially utility and provenance, are far more frequently reported than ethical indicators, and that scraping is the dominant route to datasets larger than 1,000 hours but is not the only viable one.
Load-bearing premise
The annotation is consistent enough across datasets that the differences between collection methods reflect the datasets themselves rather than the coders' judgment.
Editorial extensions
If this is right
- If TEDI is adopted, dataset documentation can be compared across datasets on a common set of verifiable indicators rather than case-by-case narrative.
- Dataset creators designing new collections know which indicators their documentation will be judged against, so ethics can be addressed at collection design time.
- Regulators and auditors can use the indicator set as a checklist for training-data transparency obligations.
- The empirical pattern implies that scaling by scraping predictably trades away documented consent and privacy, a trade-off that should be made explicit rather than incidental.
- Automating TEDI-style extraction from documentation becomes a tractable next step because indicators are phrased as answerable questions.
Reading between the lines
- If the collection-method pattern generalizes beyond voice datasets, then a dataset's ethical profile can be partially predicted from its sourcing method alone, which would let users triage documentation review effort before reading a single page.
- A testable extension would be to run the same 143-indicator audit on image-text and text-only corpora; the framework predicts scraped subsets like web-crawled image-text pairs show the same consent and privacy absence.
- The optimistic coding rule implies that even the low reported rates are upper bounds; a fine-grained re-analysis using the 143 indicators directly would likely find lower coverage, especially for consent revocation and data-worker conditions.
- TEDI could be turned into a prospective design tool: dataset creators could answer the 143 questions before collection and use negative answers as a checklist of missing safeguards, not just a documentation score.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces TEDI, a three-level hierarchy of 143 fine-grained indicators for assessing trustworthy and ethical aspects of dataset documentation, and applies it by manually annotating 114 human-voice multimodal datasets. It also proposes a two-tier data sourcing taxonomy and labels each dataset's primary, secondary, and annotation modalities according to that taxonomy. The main empirical findings are that trustworthiness indicators such as utility and provenance are documented far more often than ethical indicators; that consent, privacy, and harmful-content indicators are rarely documented; and that documentation patterns vary by collection method, with scraped datasets less likely to mention ethical indicators and crowdsourced/direct collections more likely to do so. The paper argues that these patterns reflect a tension between dataset scale and ethical documentation.
Significance. If the empirical patterns hold, this is a useful contribution: TEDI offers a systematic, verifiable rubric for comparing dataset documentation, and the data sourcing taxonomy provides a vocabulary for discussing collection practices that is currently missing for multimodal speech datasets. The focus on underexplored modalities and the manual annotation of 114 datasets is valuable, and the authors are transparent about the 'most optimistic' category-level coding rule and about the documentation-vs-dataset distinction. The main limitation is methodological: the central cross-method comparison in Figure 3 depends entirely on hand-assigned labels with no reported inter-annotator reliability, and the analysis lacks uncertainty quantification and control for confounds. These issues currently prevent the empirical claim from being fully established.
major comments (3)
- [Section 4.3, Figure 3] The central cross-method comparison rests entirely on hand-assigned labels, but no inter-annotator reliability is reported for either the TEDI category-level decisions (aggregated from 143 indicators under the any-indicator-yes rule) or the primary collection-method labels. Without a second coder or an agreement statistic, the reported differences between scraped and crowdsourced/direct datasets could be an artifact of the coding rubric rather than a property of the datasets. Please report inter-annotator agreement (e.g., Cohen's kappa on a subset), provide the annotation protocol or codebook, and run a sensitivity analysis using stricter category-level aggregation rules.
- [Table 3, Figure 3] The collection-method categories are not mutually exclusive: the percentages in Table 3 sum to 106% for primary and 130% for annotation modalities, so a single dataset can appear in multiple method columns, yet Figure 3 compares methods as if they were independent groups. In addition, no confidence intervals and no per-cell sample sizes are given, and several categories are very small (e.g., synthetic 2%, proprietary 7%). The paper should report exact counts, handle overlapping membership (e.g., by analyzing only datasets with a single primary method), and add a statistical test or confidence intervals for the cross-method differences.
- [Section 3.1, Section 4.2] The selection procedure (citation threshold, snowball sampling, purposive sampling) and the strong association between collection method, dataset size, and release year mean that the observed 'impact' of collection method on ethical indicators may be confounded. For example, large recent scraped datasets and older small direct-collection datasets differ in many ways beyond collection method. The conclusion in Section 4.3 that 'data collection methods impact ethical attributes' uses causal language for what is currently a descriptive association. Please add a robustness check (e.g., stratifying by size or release-year groups) or soften the causal wording so that the headline claim is supported by the evidence presented.
minor comments (4)
- [Section 4.3, Section 5] The 'most optimistic' any-indicator-yes rule is acknowledged, but several conclusion sentences (e.g., 'privacy and consent were largely ignored') should consistently read 'not documented' or 'not mentioned' to match TEDI's stated scope of assessing documentation rather than the underlying dataset practices.
- [Figure 7 caption] Assuming a 10-second clip duration for datasets where duration is not stated is a strong approximation; please report a sensitivity analysis for Figure 2 or at least state explicitly how sensitive the size-based conclusions are to this assumption.
- [Table 3] The table title says 'Proportional', but the percentages do not sum to 100 for any modality; please clarify the denominator, note that a dataset can be assigned multiple collection methods, and explain how the proportions were computed.
- [Figure 2 caption] The caption states that data collection methods 'influence' dataset size; this is an observed association and should be phrased as such to avoid implying a causal direction that the data do not establish.
Circularity Check
No significant circularity: the empirical findings are measured from external dataset documentation, not derived from the framework's own definitions.
full rationale
The paper's central empirical claim is that data collection methods are associated with the extent to which trustworthy and ethical indicators are documented. This is an empirical measurement result: the TEDI indicators are defined from external sources such as the Belmont Report, the EU Trustworthy AI Guidelines, and existing datasheets, and they are applied to dataset documentation that is independent of the framework's construction. The collection-method labels are assigned using a separate data sourcing taxonomy, and no indicator value is computed from a collection-method label, no fitted parameter is used to predict the reported proportions, and no equation in the paper makes the outcome equal to the input by construction. The disclosed 'any indicator yes means category yes' aggregation rule affects sensitivity but does not make the cross-method comparison definitional, and the paper explicitly states that a positive TEDI response reflects documentation availability rather than a judgment about the dataset itself. The only self-citation found is a background citation to author Alice Xiang's prior work on privacy and fairness, which is not load-bearing for the paper's central claims. The lack of reported inter-annotator reliability is a validity and reproducibility concern, not evidence of circularity, because it does not reduce the measured proportions to the authors' prior definitions. The paper is therefore self-contained as an empirical audit: its conclusions could be contradicted by applying the rubric to a different corpus or by independent re-coding of the same corpus.
Assumptions & free parameters
free parameters (3)
- citation threshold =
>10 Google Scholar citations
- assumed clip duration =
10 seconds
- optimistic category coding =
yes if any indicator in the third-level category is yes
assumptions (4)
- domain assumption Documentation mentions are a usable proxy for dataset attributes
- domain assumption Belmont principles and EU Trustworthy AI Guidelines provide a complete ethical ontology
- domain assumption Manual annotation across 114 datasets is consistent
- domain assumption The corpus is representative of multimodal datasets with human voices
invented entities (3)
-
TEDI indicator hierarchy
-
Data sourcing taxonomy
-
Primary/secondary/annotation modality convention
Cite this review
Pith. "Pith review of TEDI: Trustworthy and Ethical Dataset Indicators to Analyze and Compare Dataset Documentation." pith.science (2026). https://pith.science/paper/V6HUZSWP
@misc{pith2026250517841,
author = {Pith},
title = {Pith review of: TEDI: Trustworthy and Ethical Dataset Indicators to Analyze and Compare Dataset Documentation},
year = {2026},
howpublished = {\url{https://pith.science/paper/V6HUZSWP}},
note = {Machine review of arXiv:2505.17841}
}
read the original abstract
Dataset transparency is a key enabler of responsible AI, but insights into multimodal dataset attributes that impact trustworthy and ethical aspects of AI applications remain scarce and are difficult to compare across datasets. To address this challenge, we introduce Trustworthy and Ethical Dataset Indicators (TEDI) that facilitate the systematic, empirical analysis of dataset documentation. TEDI encompasses 143 fine-grained indicators that characterize trustworthy and ethical attributes of multimodal datasets and their collection processes. The indicators are framed to extract verifiable information from dataset documentation. Using TEDI, we manually annotated and analyzed over 100 multimodal datasets that include human voices. We further annotated data sourcing, size, and modality details to gain insights into the factors that shape trustworthy and ethical dimensions across datasets. We find that only a select few datasets have documented attributes and practices pertaining to consent, privacy, and harmful content indicators. The extent to which these and other ethical indicators are addressed varies based on the data collection method, with documentation of datasets collected via crowdsourced and direct collection approaches being more likely to mention them. Scraping dominates scale at the cost of ethical indicators, but is not the only viable collection method. Our approach and empirical insights contribute to increasing dataset transparency along trustworthy and ethical dimensions and pave the way for automating the tedious task of extracting information from dataset documentation in future.
Figures
Figures from the paper (7 more)
Reference graph
Works this paper leans on
-
[1]
AB-2013 Generative artificial intelligence: training data transparency.en. 2024. (Visited on 05/12/2025)
work page 2013
-
[2]
Croissant: A Metadata Format for ML-Ready Datasets
Mubashara Akhtar et al. “Croissant: A Metadata Format for ML-Ready Datasets”. en. In: Proceedings of the Eighth Workshop on Data Management for End-to-End Machine Learning. Santiago AA Chile: ACM, June 2024, pp. 1–6. ISBN : 9798400706110. DOI: 10 . 1145 / 3650203 . 3663326. URL: https : / / dl . acm . org / doi / 10 . 1145 / 3650203 . 3663326 (visited on ...
work page 2024
-
[3]
Ethical Considerations for Responsible Data Curation
Jerone T A Andrews et al. “Ethical Considerations for Responsible Data Curation”. en. In: 2023
work page 2023
-
[4]
Common Voice: A Massively-Multilingual Speech Corpus
Rosana Ardila et al. Common Voice: A Massively-Multilingual Speech Corpus . arXiv:1912.06670 [cs]. Mar. 2020. DOI: 10 . 48550 / arXiv . 1912 . 06670. URL: http : //arxiv.org/abs/1912.06670 (visited on 04/25/2025)
-
[5]
Training Data for the Price of a Sandwich
Stefan Baack. Training Data for the Price of a Sandwich. en. Tech. rep. Mozilla Foundation,
-
[6]
Jack Bandy and Nicholas Vincent. “Addressing "Documentation Debt" in Machine Learning: A Retrospective Datasheet for BookCorpus”. en. In: Sydney, Australia, 2021
work page 2021
-
[7]
Principles of biomedical ethics
Tom Beauchamp and James Childress. Principles of biomedical ethics . Oxford University Press, 2013
work page 2013
-
[8]
Multimodal datasets: misog- yny, pornography, and malignant stereotypes
Abeba Birhane, Vinay Uday Prabhu, and Emmanuel Kahembwe. Multimodal datasets: misog- yny, pornography, and malignant stereotypes. en. arXiv:2110.01963 [cs]. Oct. 2021. URL: http://arxiv.org/abs/2110.01963 (visited on 04/26/2024)
arXiv 2021
Show all 77 references
-
[9]
The Foundation Model Transparency Index v1.1 May 2024
Rishi Bommasani et al. The Foundation Model Transparency Index v1.1 May 2024. en. 2023
2024
-
[10]
Bias and Fairness in Multimodal Machine Learning: A Case Study of Automated Video Interviews
Brandon M. Booth et al. “Bias and Fairness in Multimodal Machine Learning: A Case Study of Automated Video Interviews”. en. In: Proceedings of the 2021 International Conference on Multimodal Interaction. Montréal QC Canada: ACM, Oct. 2021, pp. 268–277. ISBN : 978- 1-4503-8481-...
2021
-
[11]
Scripted dialogs versus improvisation: lessons learned about emotional elicitation techniques from the IEMOCAP database
Carlos Busso and Shrikanth S. Narayanan. “Scripted dialogs versus improvisation: lessons learned about emotional elicitation techniques from the IEMOCAP database”. en. In: Inter- speech 2008. ISCA, Sept. 2008, pp. 1670–1673. DOI: 10.21437/Interspeech.2008-463. URL: https://www...
2008 doi
- [12]
- [13]
-
[14]
V oxCeleb2: Deep speaker recogni- tion
Joon Son Chung, Arsha Nagrani, and Andrew Zisserman. “V oxCeleb2: Deep speaker recogni- tion”. In: arXiv ii (2018). ISSN : 23318422
2018
-
[15]
FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech
Alexis Conneau et al. FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech. arXiv:2205.12446 [cs]. May 2022. DOI: 10.48550/arXiv.2205.12446 . URL: http://arxiv.org/abs/2205.12446 (visited on 04/28/2025)
-
[16]
Data Equity: Foundational Concepts for Generative AI. Tech. rep. World Economic Forum, Oct. 2023. URL: https://www3.weforum.org/docs/WEF_Data_Equity_Concepts_ Generative_AI_2023.pdf (visited on 04/07/2025)
2023
-
[17]
ImageNet: A large-scale hierarchical image database
Jia Deng et al. “ImageNet: A large-scale hierarchical image database”. In: 2009 IEEE Confer- ence on Computer Vision and Pattern Recognition. ISSN: 1063-6919. June 2009, pp. 248–255. DOI: 10.1109/CVPR.2009.5206848. URL: https://ieeexplore.ieee.org/document/ 5206848 (visited on...
2009
-
[18]
CrowdWorkSheets: Accounting for Individual and Collective Identities Underlying Crowdsourced Dataset Annotation
Mark Diaz et al. “CrowdWorkSheets: Accounting for Individual and Collective Identities Underlying Crowdsourced Dataset Annotation”. en. In: 2022 ACM Conference on Fairness, Accountability, and Transparency. arXiv:2206.08931 [cs]. June 2022, pp. 2342–2351. DOI: 10.1145/3531146....
2022 arXiv
-
[19]
Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus
Jesse Dodge et al. Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus. arXiv:2104.08758 [cs]. Sept. 2021. DOI: 10.48550/arXiv.2104.08758. URL: http://arxiv.org/abs/2104.08758 (visited on 05/12/2025)
- [20]
-
[21]
Pages: 1-39 Publication Title: European Commission
Ethics Guidelines for Trustworthy AI. Pages: 1-39 Publication Title: European Commission
- [22]
-
[23]
Datasheets for datasets
Timnit Gebru et al. “Datasheets for datasets”. In: Communications of the ACM 64.12 (2021). arXiv: 1803.09010, pp. 86–92. ISSN : 15577317. DOI: 10.1145/3458723
2021 arXiv
-
[24]
Audio Set: An ontology and human-labeled dataset for audio events
Jort F. Gemmeke et al. “Audio Set: An ontology and human-labeled dataset for audio events”. In: 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). ISSN: 2379-190X. Mar. 2017, pp. 776–780. DOI: 10.1109/ICASSP.2017.7952261. URL: https://ieeex...
2017
- [25]
-
[26]
Speaker recognition by machines and humans: A tutorial review
John H. L. Hansen and Taufiq Hasan. “Speaker recognition by machines and humans: A tutorial review”. In: IEEE Signal Processing Magazine 32.6 (2015), pp. 74–99. ISSN : 10535888. DOI: 10.1109/MSP.2015.2462851
2015
-
[27]
ActivityNet: A large-scale video benchmark for human activity understanding
Fabian Caba Heilbron et al. “ActivityNet: A large-scale video benchmark for human activity understanding”. In: 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). ISSN: 1063-6919. June 2015, pp. 961–970. DOI: 10.1109/CVPR.2015.7298698. URL: https://ieeexplo...
2015
-
[28]
Design Research in Information Systems
Alan Hevner and Samir Chatterjee. Design Research in Information Systems. Ed. by Ramesh Sharda and Stefan V oß. Springer, 2010.ISBN : 978-1-4419-5652-1. DOI: 10.1007/978-1- 4419-5653-8. URL: http://www.springer.com/series/6157
2010 doi
-
[29]
URL: https://partnershiponai
Improving Conditions for Data Enrichment Workers. URL: https://partnershiponai. org/responsible-sourcing-library/ (visited on 04/07/2025)
2025
-
[30]
A Standardized Machine-readable Dataset Documentation Format for Responsible AI
Nitisha Jain et al. A Standardized Machine-readable Dataset Documentation Format for Responsible AI. en. arXiv:2407.16883 [cs]. June 2024. URL: http://arxiv.org/abs/2407. 16883 (visited on 08/21/2024)
2024 arXiv
-
[31]
Libri-Light: A Benchmark for ASR with Limited or No Supervision
Jacob Kahn et al. “Libri-Light: A Benchmark for ASR with Limited or No Supervision”. In: ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). arXiv:1912.07875 [cs]. May 2020, pp. 7669–7673. DOI: 10.1109/ ICASSP40776.2020.9052942...
2020 arXiv
-
[32]
Libriheavy: a 50,000 hours ASR corpus with punctuation casing and context
Wei Kang et al. Libriheavy: a 50,000 hours ASR corpus with punctuation casing and context. arXiv:2309.08105 [eess]. Jan. 2024. DOI: 10 . 48550 / arXiv . 2309 . 08105. URL: http : //arxiv.org/abs/2309.08105 (visited on 04/28/2025)
- [33]
-
[34]
A Data Perspective on Ethical Challenges in V oice Biometrics Research
Anna Leschanowsky et al. “A Data Perspective on Ethical Challenges in V oice Biometrics Research”. In: IEEE Transactions on Biometrics, Behavior, and Identity Science 7.1 (Jan. 2025), pp. 118–131. ISSN : 2637-6407. DOI: 10.1109/TBIOM.2024.3446846. URL: https: //ieeexplore.ieee...
2025
- [35]
-
[36]
Foundations & Trends in Multi- modal Machine Learning: Principles, Challenges, and Open Questions
Paul Pu Liang, Amir Zadeh, and Louis-Philippe Morency. “Foundations & Trends in Multi- modal Machine Learning: Principles, Challenges, and Open Questions”. en. In: ACM Comput- ing Surveys (Apr. 2024), p. 3656580. ISSN : 0360-0300, 1557-7341. DOI: 10.1145/3656580. URL: https://...
2024 doi
- [37]
-
[38]
Data Authenticity, Consent, & Provenance for AI are all broken: what will it take to fix them? en
Shayne Longpre et al. Data Authenticity, Consent, & Provenance for AI are all broken: what will it take to fix them? en. arXiv:2404.12691 [cs]. Apr. 2024. URL: http://arxiv.org/ abs/2404.12691 (visited on 05/16/2024)
2024 arXiv
-
[39]
The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI
Shayne Longpre et al. The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI. en. arXiv:2310.16787 [cs]. Nov. 2023. URL: http://arxiv. org/abs/2310.16787 (visited on 05/07/2024)
2023 arXiv
-
[40]
Data Statements: From Technical Concept to Community Practice
Angelina McMillan-Major, Emily M. Bender, and Batya Friedman. “Data Statements: From Technical Concept to Community Practice”. en. In: ACM Journal on Responsible Computing 1.1 (Mar. 2024), pp. 1–17. ISSN : 2832-0565. DOI: 10.1145/3594737 . URL: https://dl. acm.org/doi/10.1145/...
2024 doi
-
[41]
Gender Artifacts in Visual Datasets
Nicole Meister et al. Gender Artifacts in Visual Datasets. en. arXiv:2206.09191 [cs]. Sept
-
[42]
Documenting Computer Vision Datasets: An Invitation to Reflexive Data Practices
Milagros Miceli et al. “Documenting Computer Vision Datasets: An Invitation to Reflexive Data Practices”. In: ACM Fairness Accountability and Transparency (FAccT) ’21 . ACM, 2021, p. 12. ISBN : 978-1-4503-8309-7. DOI: 10.1145/3442188.3445880 . URL: https: //doi.org/10.1145/344...
2021
-
[43]
Measuring Data
Margaret Mitchell et al. Measuring Data. en. arXiv:2212.05129 [cs]. Feb. 2023. URL: http: //arxiv.org/abs/2212.05129 (visited on 03/12/2024)
2023 arXiv
-
[44]
On responsible machine learning datasets emphasizing fairness, privacy and regulatory norms with examples in biometrics and healthcare
Surbhi Mittal et al. “On responsible machine learning datasets emphasizing fairness, privacy and regulatory norms with examples in biometrics and healthcare”. en. In: Nature Machine Intelligence (Aug. 2024). Publisher: Nature Publishing Group, pp. 1–14.ISSN : 2522-5839. DOI: 1...
2024 doi
-
[45]
Data Collection in Music Generation Training Sets: A Critical Analysis
Fabio Morreale, Megha Sharma, and I-Chieh Wei. “Data Collection in Music Generation Training Sets: A Critical Analysis”. en. In: Milan, Italy, 2023
2023
-
[46]
V oxceleb: A large-scale speaker identification dataset
Arsha Nagrani, Joon Son Chung, and Andrew Zisserman. “V oxceleb: A large-scale speaker identification dataset”. In: arXiv (2017), pp. 2616–2620. ISSN : 23318422
2017
-
[47]
DMLR: Data-centric Machine Learning Research – Past, Present and Future
Luis Oala et al. DMLR: Data-centric Machine Learning Research – Past, Present and Future. en. arXiv:2311.13028 [cs, eess]. Nov. 2023. URL: http://arxiv.org/abs/2311.13028 (visited on 05/06/2024)
2023 arXiv
-
[48]
Librispeech: An ASR corpus based on public domain audio books
Vassil Panayotov et al. “Librispeech: An ASR corpus based on public domain audio books”. In: ICASSP , IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings 2015-Augus (2015). Publisher: IEEE ISBN: 9781467369978, pp. 5206–5210. ISSN : 15206149. ...
2015
-
[49]
MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations
Soujanya Poria et al. MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations. en. arXiv:1810.02508 [cs]. June 2019. DOI: 10.48550/arXiv.1810.02508. URL: http://arxiv.org/abs/1810.02508 (visited on 04/25/2025)
-
[50]
Large image datasets: A pyrrhic win for computer vision? en
Vinay Uday Prabhu and Abeba Birhane. Large image datasets: A pyrrhic win for computer vision? en. arXiv:2006.16923 [cs]. July 2020. DOI: 10.48550/arXiv.2006.16923 . URL: http://arxiv.org/abs/2006.16923 (visited on 05/12/2025)
-
[51]
MLS: A Large-Scale Multilingual Dataset for Speech Research
Vineel Pratap et al. “MLS: A Large-Scale Multilingual Dataset for Speech Research”. en. In: Interspeech 2020. arXiv:2012.03411 [eess]. Oct. 2020, pp. 2757–2761. DOI: 10.21437/ Interspeech . 2020 - 2826. URL: http : / / arxiv . org / abs / 2012 . 03411(visited on 04/25/2025). 12
2020 arXiv
-
[52]
Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI
Mahima Pushkarna, Andrew Zaldivar, and Oddur Kjartansson. “Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI”. en. In: 2022 ACM Conference on Fairness, Accountability, and Transparency . Seoul Republic of Korea: ACM, June 2022, pp. 1776–1826. ISBN...
2022
-
[53]
Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 laying down harmonised rules on artificial intelligence and amending Regulations (EC) No 300/2008, (EU) No 167/2013, (EU) No 168/2013, (EU) 2018/858, (EU) 2018/1139 and (EU) 2019/2144 and D...
2024
-
[54]
Open Datasheets: Machine-readable Documentation for Open Datasets and Responsible AI Assessments
Anthony Cintron Roman et al. Open Datasheets: Machine-readable Documentation for Open Datasets and Responsible AI Assessments. en. arXiv:2312.06153 [cs]. Mar. 2024. DOI: 10. 48550/arXiv.2312.06153 . URL: http://arxiv.org/abs/2312.06153 (visited on 04/06/2025)
-
[55]
Measuring Social Biases in Grounded Vision and Language Embeddings
Candace Ross, Boris Katz, and Andrei Barbu. “Measuring Social Biases in Grounded Vision and Language Embeddings”. en. In:Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Online: Asso...
2021
-
[56]
About V oice: A Longitudinal Study of Speaker Recognition Dataset Dynamics
Casandra Rusti et al. “About V oice: A Longitudinal Study of Speaker Recognition Dataset Dynamics”. In: (Apr. 2023). arXiv: 2304.03858. URL: http://arxiv.org/abs/2304. 03858
2023 arXiv
-
[57]
Second Draft of the General-Purpose AI Code of Practice. en. 2025. URL: https://digital- strategy.ec.europa.eu/en/library/second-draft-general-purpose-ai-code- practice-published-written-independent-experts (visited on 05/12/2025)
2025
-
[58]
Representation Bias in Data: A Survey on Identification and Resolution Techniques
Nima Shahbazi et al. “Representation Bias in Data: A Survey on Identification and Resolution Techniques”. en. In: ACM Computing Surveys 55.13s (Dec. 2023), pp. 1–39. ISSN : 0360-0300, 1557-7341. DOI: 10.1145/3588433. URL: https://dl.acm.org/doi/10.1145/3588433 (visited on 03/19/2024)
2023 doi
-
[59]
The Great Scrape: The Clash Between Scraping and Privacy
Daniel J. Solove and Woodrow Hartzog. “The Great Scrape: The Clash Between Scraping and Privacy”. en. In: SSRN Electronic Journal (2024). ISSN : 1556-5068. DOI: 10.2139/ssrn. 4884485. URL: https://www.ssrn.com/abstract=4884485 (visited on 07/19/2024)
2024 doi
-
[60]
Mitigating Gender Bias in Captioning Systems
Ruixiang Tang et al. “Mitigating Gender Bias in Captioning Systems”. en. In: Proceedings of the Web Conference 2021. Ljubljana Slovenia: ACM, Apr. 2021, pp. 633–645. ISBN : 978- 1-4503-8312-7. DOI: 10.1145/3442381.3449950. URL: https://dl.acm.org/doi/10. 1145/3442381.3449950 (...
2021
-
[61]
The Belmont Report. en. Tech. rep. US Department of Health, Education, and Welfare, 1979. URL: https://www.hhs.gov/ohrp/sites/default/files/the- belmont- report- 508c_FINAL.pdf
1979
-
[62]
URL: https://keithito.com/LJ-Speech-Dataset (visited on 04/25/2025)
The LJ Speech Dataset. URL: https://keithito.com/LJ-Speech-Dataset (visited on 04/25/2025)
2025
-
[63]
Third Draft of the General-Purpose AI Code of Practice. en. 2025. URL: https://digital- strategy.ec.europa.eu/en/library/third-draft-general-purpose-ai-code- practice-published-written-independent-experts (visited on 05/12/2025)
2025
-
[64]
YFCC100M: The New Data in Multimedia Research
Bart Thomee et al. “YFCC100M: The New Data in Multimedia Research”. In:Communications of the ACM 59.2 (Jan. 2016). arXiv:1503.01817 [cs], pp. 64–73. ISSN : 0001-0782, 1557-7317. DOI: 10 . 1145 / 2812802. URL: http : / / arxiv . org / abs / 1503 . 01817(visited on 04/25/2025)
2016 arXiv
-
[65]
Interim Measures for the Management of Generative Artificial In- telligence Services
China Law Translate. Interim Measures for the Management of Generative Artificial In- telligence Services . en. July 2023. URL: https : / / www . chinalawtranslate . com / generative-ai-interim/ (visited on 05/12/2025). 13
2023
-
[66]
V oxPopuli: A Large-Scale Multilingual Speech Corpus for Represen- tation Learning, Semi-Supervised Learning and Interpretation
Changhan Wang et al. “V oxPopuli: A Large-Scale Multilingual Speech Corpus for Represen- tation Learning, Semi-Supervised Learning and Interpretation”. In: Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint ...
2021 doi
-
[67]
Contrastive Language-Vision AI Models Pretrained on Web-Scraped Multimodal Data Exhibit Sexual Objectification Bias
Robert Wolfe et al. “Contrastive Language-Vision AI Models Pretrained on Web-Scraped Multimodal Data Exhibit Sexual Objectification Bias”. en. In: 2023 ACM Conference on Fairness, Accountability, and Transparency. Chicago IL USA: ACM, June 2023, pp. 1174–
2023
-
[68]
Sierra Wyllie, Ilia Shumailov, and Nicolas Papernot.Fairness Feedback Loops: Training on Synthetic Data Amplifies Bias. en. arXiv:2403.07857 [cs]. Mar. 2024. URL: http://arxiv. org/abs/2403.07857 (visited on 04/30/2024)
2024 arXiv
-
[69]
Being ’Seen’ vs
Alice Xiang. Being ’Seen’ vs. ’Mis-Seen’: Tensions between Privacy and Fairness in Computer Vision. en. SSRN Scholarly Paper. Rochester, NY, Feb. 2022. DOI: 10.2139/ssrn.4068921. URL: https://papers.ssrn.com/abstract=4068921 (visited on 05/12/2025)
2022 doi
-
[70]
CSTR VCTK Corpus: English Multi-speaker Corpus for CSTR V oice Cloning Toolkit (version 0.92)
Junichi Yamagishi, Christophe Veaux, and Kirsten MacDonald. “CSTR VCTK Corpus: English Multi-speaker Corpus for CSTR V oice Cloning Toolkit (version 0.92)”. eng. In:The Rainbow Passage which the speakers read out can be found in the International Dialects of English Archive: (...
2019 doi
-
[71]
Navigating Dataset Documentations in AI: A Large-Scale Analysis of Dataset Cards on Hugging Face
Xinyu Yang, Weixin Liang, and James Zou. Navigating Dataset Documentations in AI: A Large-Scale Analysis of Dataset Cards on Hugging Face . en. arXiv:2401.13822 [cs]. Jan
-
[72]
Position: Measure Dataset Diversity, Don’t Just Claim It
Dora Zhao et al. “Position: Measure Dataset Diversity, Don’t Just Claim It”. en. In:Proceedings of the 41 st International Conference on Machine Learning. Vienna, Austria, 2024. 14 A Appendix A.1 Supporting Material: Trustworthy and Ethical Dataset Indicators Table 4: Detailed...
2024
- [76]
-
[1185]
DOI: 10.1145/3593013.3594072
ISBN : 9798400701924. DOI: 10.1145/3593013.3594072 . URL: https://dl.acm. org/doi/10.1145/3593013.3594072 (visited on 04/17/2024)
2024
-
[2019]
URL: https : / / digital - strategy . ec . europa . eu / en / library / ethics - guidelines-trustworthy-ai
- [2023]
-
[2024]
URL: https://foundation.mozilla.org/en/research/library/generative- ai-training-data/common-crawl/
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.