{"id":"9101c289-f220-46a3-8b3d-6005a520a948","arxiv_id":"2506.04305","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A survey of 796 AI/ML professionals finds that disabled employees, women, gender minorities, and sexual minorities report worse workplace experiences and more microaggressions than majority groups.","lead":"A pilot survey of 796 AI/ML professionals shows that disabled employees, women, gender minorities, and sexual minorities report worse workplace experiences and more frequent microaggressions than their majority counterparts. The study is a first step toward measuring workplace transparency and DEI outcomes in the AI community, though its self-selected sample limits generalization.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The disability-disparity claim rests on unadjusted comparisons; age or employer confounding could drive the reported gap, and no effect sizes are reported.","rationale":"The reader's weakest assumption focused on sample representativeness: the convenience sample recruited through AI affinity groups cannot support generalization to the broader AI/ML workforce. That is a real limitation and is acknowledged in the paper's own limitations section. However, the most load-bearing concern for the specific headline claim about disability is internal validity: the disability comparison is unadjusted and could be confounded by age, employer, or seniority. This is a distinct issue from external generalizability. The paper already receives a CONDITIONAL verdict with moderate confidence, and this additional internal-validity concern reinforces that conditionality rather than changing the verdict. The requested check is feasible with the existing data and would directly test whether the disability gap survives adjustment, which is necessary before the claim can be treated as robust.","tokens_in":21130,"tokens_out":2536,"duration_ms":26620,"concrete_test":"Re-analyze the disability versus non-disabled comparison for the DEI-leaving item and the performance-fairness item using ordinal logistic regression (or stratified/rank-based tests) with covariates for age group, gender, race/ethnicity, academia versus industry, seniority, and employer. Report adjusted odds ratios and 95% confidence intervals. If the disability coefficient attenuates to non-significance or a small effect, the central claim should be softened to describe unadjusted associations in a convenience sample.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's central claim that 'disabled employees have a worse workplace experience than their non-disabled colleagues' is supported primarily by a single unadjusted comparison: disabled participants scored significantly higher on 'I have thought about leaving my Employer due to DEI-related reasons' (p=0.000048, §4.2), plus descriptive accessibility items with no non-disabled comparator (§4.4). This comparison is vulnerable to confounding. The paper itself shows that unfavorable responses increase with age (Figure 3), and disability status is typically associated with age. If disabled respondents are older, or concentrated in particular employer types or seniority bands, the observed gap could reflect age or employer effects rather than disability. No adjustment, stratification, interaction analysis, or effect size is reported; all tests treat ordinal Likert scales as interval and no multiple-comparison correction is applied. Because the headline claim is broader than any single item, the evidence as presented does not securely establish a disability-specific workplace-experience deficit.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a pilot online survey of 1,260 AI/ML professionals (796 complete responses) recruited through AI affinity groups, NeurIPS 2023 events, and personal outreach, conducted from November 2023 to March 2024. The survey covers workplace DEI initiatives, belonging, accessibility, microaggressions, misconduct, performance and compensation, growth, well-being, and overall satisfaction, with responses on Likert scales. The authors compute descriptive statistics and unadjusted statistical tests (Student's t-tests and chi-squared tests) to compare majority and minority demographic groups. The central claim is that workplace disparities persist for underrepresented and marginalized subgroups, with particular emphasis on accessibility and on disabled employees having a worse workplace experience than non-disabled colleagues. The paper also presents intersectional analyses, employer-level comparisons for three large companies, and a non-disclosure analysis. The authors explicitly acknowledge in Sections 3.4 and 5.1 that the sample is not unbiased, that women are over-represented (60.55% versus roughly 22% in the field), and that the results should be interpreted as a pilot.","tokens_in":21275,"tokens_out":2794,"duration_ms":29863,"significance":"If the results are taken as descriptive evidence within the surveyed sample, the paper provides a useful first large-scale, multi-dimensional snapshot of workplace experiences in the AI/ML community, going beyond representation counts to probe lived experiences across several identity axes. The intersectional splits, the accessibility module for disabled respondents, the non-disclosure analysis, and the inclusion of the full survey instrument in the appendix are genuine strengths. The paper is also transparent about its sampling limitations and avoids over-claiming employer-level findings. However, the headline claims in the abstract and Section 5 are considerably stronger than the statistical evidence presented: the disability-disparity claim rests on a single unadjusted comparison plus descriptive items without a comparator, and the sampling frame is a self-selected convenience sample. The paper is best viewed as hypothesis-generating; its value would be enhanced by reweighting or external benchmarking, adjusted analyses, and effect sizes.","major_comments":[{"comment":"The central claim that 'disabled employees have a worse workplace experience than their non-disabled colleagues' is not securely established by the evidence provided. The only direct statistical comparison is a single unadjusted t-test on the item 'I have thought about leaving my Employer due to DEI-related reasons' (p=0.000048), while the accessibility items in §4.4 have no non-disabled comparator. Given that Figure 3 shows unfavorable responses increasing with age, and disability status is plausibly associated with age, employer type, or seniority, the paper needs adjusted analyses (e.g., stratification, regression, or propensity-based weighting), effect sizes with confidence intervals, or a substantial reframing of the result as a preliminary, sample-specific finding.","section":"Abstract and §4.2"},{"comment":"The statistical reporting is not adequate for the number and scope of claims. Likert-scale responses are treated as interval data, many unadjusted tests are run, and no correction for multiple testing is applied; p-values such as p=0.000048, p=0.0000117, and p=0.0077 are reported without effect sizes or confidence intervals. Because dozens of comparisons are implicit in the demographic splits and item-level analyses, the reader cannot judge which findings would survive a multiple-comparison correction. The authors should either provide a correction, report sensitivity analyses, or explicitly label all results as exploratory.","section":"§3.5 and §4"},{"comment":"The sampling frame is a self-selected convenience sample recruited through AI affinity groups and conference games, with women substantially over-represented relative to the AI/ML workforce (60.55% versus about 22%). The authors acknowledge this in §5.1, but the abstract and Section 5 nonetheless generalize to 'the AI/ML community' and 'disabled employees' without qualification. The over-representation of women can confound demographic comparisons, and self-selection into the survey may correlate with workplace attitudes in ways that bias within-sample disparity estimates. The manuscript should either reweight to known population benchmarks, compare against external data, or explicitly confine all headline claims to the surveyed sample.","section":"§3.4, §4.1, and §5.1"}],"minor_comments":[{"comment":"There are typos: 'respondants' should be 'respondents' in §4.5, and 'their is an interest' should be 'there is an interest' in §5.2.","section":"§4.5 and §5.2"},{"comment":"The bar plots and line plots in Figures 3 and 5 would benefit from error bars or confidence intervals; without them, the visual comparisons overstate precision, especially for groups such as disabled respondents (n=92) and subgroups split by age and gender.","section":"Figures 3 and 5"},{"comment":"The race/ethnicity abbreviations (W, EA, SA, AAB, H/L, ME, OR) are defined only in the caption text; please define them explicitly in the table caption or in a footnote so the table is self-contained.","section":"Table 3"},{"comment":"The phrase 'men and gender minorities' in the text and figure caption seems to refer to gender majority versus gender minority groups after binarization; please align the wording with the binarized variable described in §4.3 to avoid ambiguity.","section":"Figure 4 and §4.3"},{"comment":"The sentence reporting a non-significant age difference (p=0.8057) among respondents experiencing microaggressions is unclear about which comparison is being made; please specify the contingency table and the groups being compared.","section":"§4.4"}],"recommendation":"major_revision","confidential_remarks":"The paper is honest about its limitations, but the gap between the abstract's strong claims and the unadjusted, self-selected-sample evidence is the main barrier to acceptance. If the authors reframe the central claims as sample-specific and exploratory, or add reweighting and adjusted analyses, the paper would be a valuable descriptive contribution. Scope-wise, cs.CY is appropriate."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing to know: this paper gives the AI/ML community something it didn't have—a multi-construct workplace survey across gender, race/ethnicity, sexual orientation, disability, and intersections, with n=796 completions. The instrument is in the appendix, ethics approval is documented, and the authors are candid that the sample is convenience-based and over-represents women and minority groups. That transparency is real and not perfunctory; they even report non-disclosure patterns, which most surveys ignore.\n\nWhat's new is the dataset and the intersectional breakdowns. The finding that marginalized groups report worse experiences is consistent with decades of workplace research in other fields, which is exactly why the pilot is useful rather than surprising.\n\nSoft spots, in order of importance. The abstract's statement that 'disabled employees have a worse workplace experience than their non-disabled colleagues' is supported by unadjusted t-tests on a few items (e.g., 'thought about leaving for DEI reasons', p=0.000048) plus descriptive accessibility percentages. The stress-test worry about age confounding is fair: their own Figure 3 shows unfavorable responses rise with age, and they never show the age distribution of disabled vs. non-disabled respondents. A stratified analysis or a small regression adjusting for age, seniority, and employer type would answer this. I don't think the worry sinks the paper—multiple items point the same way and prior literature supports the direction—but the headline claim is broader than the statistics as presented.\n\nThe broader statistical limitations are standard for a pilot: Likert scales treated as interval, no confidence intervals or effect sizes, no multiple-comparison correction, and a self-selected sample. These limit the magnitudes, not the direction. The authors use hedged language most of the time, though the abstract could use a 'suggests' instead of an unqualified 'indicate.'\n\nBottom line: this deserves a serious referee. It is a first empirical baseline for a relevant community, the methods are transparent enough to reproduce, and the limitations are acknowledged rather than buried. I'd send it out and ask for adjusted analyses, effect sizes, and softened generalization language in revision. I'd cite it as a pilot baseline in my own work.","headline":"A welcome, honestly limited pilot baseline for AI/ML workplace disparities; the disability gap is plausible but not yet adjusted for age or employer.","tokens_in":21821,"tokens_out":2994,"would_cite":true,"duration_ms":28518,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A survey of 1,260 AI/ML professionals finds that disabled employees and other underrepresented groups report persistently worse workplace experiences than their colleagues.","keywords":["workplace disparities","AI/ML workforce","disability and accessibility","DEI initiatives","microaggressions","pilot survey","intersectionality","responsible AI"],"falsifier":"A reweighted or externally benchmarked replication using a probability-based or employer-partnered sample of AI/ML professionals, matched to known field demographics, would settle the claim: if disabled and non-disabled employees showed no significant difference on leaving-for-DEI reasons or on accessibility comfort after controlling for age, gender, seniority, and employer, the central conclusion would fail.","tokens_in":1830,"feed_emoji":"♿","tokens_out":2179,"duration_ms":112280,"temperature":0.7,"pith_summary":"This paper sets out to show that workplace disparities in the AI/ML community are not just a matter of who shows up, but of how different employees actually experience the workplace, and that those gaps persist even when employers say they value diversity. The authors fielded a pilot survey of 1,260 AI/ML professionals, with 796 complete responses, covering belonging, accessibility, DEI initiatives, microaggressions, misconduct, performance, compensation, growth, well-being, and overall satisfaction. The headline result is that disabled employees report a worse workplace experience than non-disabled colleagues: 45.65% of disabled respondents are not comfortable discussing physical or cognitive challenges, and 41.30% are not confident that requesting accommodations will not harm their career. The paper also documents significantly higher rates of unfavorable experiences for women, gender minorities, sexual minorities, and older workers, and a gap between employers stating that they value DEI and visibly committing to it. For a field concerned with responsible AI, this suggests that the internal work environment is part of the responsibility.","feed_headline":"Disabled AI staff report worse workplaces, survey of 1,260 finds","feed_subtitle":"45.65% feel uncomfortable discussing challenges; 41.30% fear accommodations could backfire.","key_machinery":"The machine carrying the argument is the survey instrument itself: up to 55 self-report questions across nine constructs plus demographics, answered on four-point and five-point Likert scales with no neutral option. All items are positively framed, and options are ordered from most positive to most negative. The analysis encodes responses in ascending order (strongly agree = 1 through strongly disagree = 4, and never = 1 through multiple times a day = 5), so lower scores mean better workplace experience, then reports medians, modes, and the fraction of unfavorable responses. To compare subgroups, the authors binarize demographic dimensions into majority and minority groups with a minimum cell size of $n = 30$ (and $n = 10$ within employers) and use unpaired or paired Student's $t$-tests and $\\chi^2$ tests. This aggregation machinery is what lets the paper convert raw Likert answers into the disparity claims about disability, gender, race, sexual orientation, and intersectional groups.","core_discovery":"On its own terms, the paper's central discovery is that disparities in AI workplace experience are enduring and measurable, and that the clearest signal is disability and accessibility. Participants reporting a disability had significantly higher scores on considering leaving their employer for DEI-related reasons ($p = 0.000048$), and among disabled respondents 45.65% do not feel comfortable discussing physical or cognitive challenges while 41.30% are not confident that requesting accommodations will not harm their career. Microaggressions also track identity: 81.51% of respondents who reported experiencing workplace microaggressions identify as gender minorities, 30.99% identify with a minority sexual orientation, and 15.63% have a disability or chronic condition, all at statistically significant rates relative to the respondent pool. In the DEI category, unfavorable responses rise significantly between the item 'my employer states that they value DEI' and 'my employer demonstrates a visible commitment' ($p < 0.00001$), and DEI efforts are reported to be carried disproportionately by members of minority groups. The paper presents these results as the first extensive, multi-dimensional survey of AI/ML professionals across roles, locations, employers, and seniority levels, and as a baseline for tracking whether DEI interventions actually change workplace experience.","pith_inferences":["If the disability findings replicate, employers should treat fear of requesting accommodations as a retention metric rather than a compliance checkbox; the 41.30% who fear career harm are likely to leak out of the field over time.","Because the sample over-represents women and was recruited through identity-based affinity groups, the reported race-by-gender gaps may be partly an artifact of who self-selected; a matched or weighted sample could change the magnitudes even if the directions hold.","Respondents with undisclosed identities report more favorable scores than any disclosed group, which suggests non-disclosure may systematically inflate apparent workplace satisfaction; future surveys should model \"prefer not to say\" as a substantive category, not just missing data.","The same instrument, repeated annually, could serve as an early-warning system for the effects of DEI rollbacks and sociopolitical changes, complementing company-reported diversity statistics."],"forward_implications":["Accessibility is a measured weakness for AI employers: roughly 46% of disabled respondents are uncomfortable discussing challenges and roughly 41% fear that requesting accommodations could hurt their career.","Employers' DEI messaging and DEI practice diverge: unfavorable responses rise significantly between the \"states they value DEI\" item and the \"demonstrates visible commitment\" item, and DEI initiatives are reported to be led mainly by minority-group members.","Microaggression patterns in AI workplaces resemble broader society, with gender minorities, sexual minorities, and disabled employees experiencing higher rates, so workplace-based DEI interventions cannot be treated as separate from societal bias.","The instrument offers a reusable baseline for tracking DEI effectiveness over time, which the authors argue is urgent as DEI programs face political backlash and rollback."],"supporting_citations":[{"why":"Supplies the 22% field-wide share of women used to flag the survey sample's over-representation of women.","marker":"Interface 2024"},{"why":"Provides the definition of microaggressions that the survey's microaggression items are built on.","marker":"Sue et al. 2007"},{"why":"Provides the method for aggregating multi-item Likert responses into a single scale per participant, which the analysis relies on.","marker":"Svensson 2001"},{"why":"Supplies the intersectional framework the authors use when splitting race results by gender.","marker":"Crenshaw 1991"},{"why":"Validates the sense-of-belonging questionnaire approach the survey adapts for AI/ML contexts.","marker":"Knekta, Chatzikyriakidou, and McCartney 2020"},{"why":"Documents disability-based workplace bias in another profession, providing precedent for expecting disability disparities in AI.","marker":"Blanck, Hyseni, and Wise 2021"},{"why":"Reports a multi-profession U.S. survey on microaggressions that motivates and frames the paper's microaggression category.","marker":"Gassam 2019"}],"fun_headline_variants":["Disabled AI staff face worse workplace experiences, survey finds","Survey of 1,260 reveals disability gaps in AI workplaces","AI workplace disparities hit disabled employees hardest","Accessibility gap persists in AI industry, survey shows","Disability disparities endure in AI, poll of professionals finds"],"cache_read_input_tokens":24064,"weakest_assumption_plain":"The central finding rests on the assumption that a self-selected convenience sample, recruited heavily through AI affinity groups and conference games, represents the broader AI/ML workforce well enough for within-sample comparisons, despite women being over-represented at 60.55% versus about 22% in the field.","fun_headline_variants_meta":{"raw":{"variants":["Disabled AI staff face worse workplace experiences, survey finds","Survey of 1,260 reveals disability gaps in AI workplaces","AI workplace disparities hit disabled employees hardest","Accessibility gap persists in AI industry, survey shows","Disability disparities endure in AI, poll of professionals finds"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000245,"raw_usage":{"total_tokens":1607,"prompt_tokens":1090,"completion_tokens":517,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":706,"completion_tokens_details":{"reasoning_tokens":441}},"tokens_in":706,"tokens_out":517,"duration_ms":5493,"temperature":1.0,"reasoning_tokens":441,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T10:45:51.146667+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A reweighted or externally benchmarked replication using a probability-based or employer-partnered sample of AI/ML professionals, matched to known field demographics, would settle the claim: if disabled and non-disabled employees showed no significant difference on leaving-for-DEI reasons or on accessibility comfort after controlling for age, gender, seniority, and employer, the central conclusion would fail.","supporting_citations":[],"review_version":1}