Pith. sign in

REVIEW

Casual Conversations v2: Designing a large consent-driven dataset to measure algorithmic bias and robustness

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2211.05809 v1 pith:JN2N7HNO submitted 2022-11-10 cs.CV cs.AIcs.CLcs.CY

classification cs.CVcs.AIcs.CLcs.CY
keywords categoriesalgorithmicbiascasualcollectingcomprehensiveconsent-drivenconversations
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Developing robust and fair AI systems require datasets with comprehensive set of labels that can help ensure the validity and legitimacy of relevant measurements. Recent efforts, therefore, focus on collecting person-related datasets that have carefully selected labels, including sensitive characteristics, and consent forms in place to use those attributes for model testing and development. Responsible data collection involves several stages, including but not limited to determining use-case scenarios, selecting categories (annotations) such that the data are fit for the purpose of measuring algorithmic bias for subgroups and most importantly ensure that the selected categories/subcategories are robust to regional diversities and inclusive of as many subgroups as possible. Meta, in a continuation of our efforts to measure AI algorithmic bias and robustness (https://ai.facebook.com/blog/shedding-light-on-fairness-in-ai-with-a-new-data-set), is working on collecting a large consent-driven dataset with a comprehensive list of categories. This paper describes our proposed design of such categories and subcategories for Casual Conversations v2.

Discussion (0). Continue with ORCID to comment.

Pith tools