Pith. sign in

REVIEW 5 major objections 4 minor 2 cited by

AI developers' own social-impact reports are sparse, shallow, and declining, while independent evaluations are broader and more rigorous—yet only developers can supply data on labor, provenance, and costs, so critical gaps remain.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review

2026-08-03 23:39 UTC pith:54DFY7NL

load-bearing objection A useful first map of who reports social-impact evaluations, backed by a released dataset and honest caveats; the first-vs-third-party gap is probably real, but the paper overclaims 'rigor' and the visibility-biased sampling frame deserves more weight. the 5 major comments →

arxiv 2511.05613 v2 pith:54DFY7NL submitted 2025-11-06 cs.CY cs.AIcs.LG

Who Evaluates AI's Social Impacts? Mapping Coverage and Gaps in First and Third Party Evaluations

classification cs.CY cs.AIcs.LG
keywords social impact evaluationfoundation modelsfirst-party reportingthird-party evaluationmodel cardsAI transparencyAI governanceevaluation gaps
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper tries to establish, with the first large-scale quantitative audit of social-impact evaluation reporting, a division of labor: first-party developer reports on bias, environment, privacy, and similar dimensions are thin and have gotten thinner since 2023, while independent academic and nonprofit evaluations provide more frequent and more detailed coverage of bias, sensitive content, and performance disparities. It also claims that the categories third parties cannot cover without insider access—data provenance, content-moderation labor, financial costs, training infrastructure—are exactly the ones developers deprioritize. If true, this means current evaluation practice is complementary but incomplete: regulation and shared evaluation infrastructure are needed, not just more benchmarks. The claim rests on a manually assembled dataset of 186 release reports and 183 post-release sources, scored on seven dimensions using a standardized rubric, plus interviews with model developers.

Core claim

The paper's central claim is that AI's social impacts are systematically under-evaluated where accountability matters most. First-party reporting averages 0.72 on a 0-3 detail scale, versus 2.62 for third-party evaluations, and a Bayesian regression shows first-party reports are far less likely to reach higher detail across every measured dimension except one. The same data show environmental-cost and bias reporting declined sharply after late 2023 even though it had been feasible earlier. Interviews attribute the decline to strategic deprioritization: evaluations are run when they support product adoption, affect business metrics, or are legally required, while privacy and labor reporting a

What carries the argument

The load-bearing instrument is a standardized 0-3 scoring rubric applied to seven social-impact dimensions: bias and representational harms, sensitive content, disparate performance, environmental costs, privacy and data protection, financial costs, and data/content-moderation labor. Each evaluation instance is classified as first-party or third-party, and the rubric scores specificity and reproducibility—from no mention, to vague mention, to concrete results without methodology, to fully contextualized and reproducible reporting. This classification lets the authors compare coverage and detail across 186 release reports and 183 post-release sources, and a Bayesian hierarchical ordinal regre

Load-bearing premise

The entire measurement—the size of the gaps, the declining trends, the first/third-party contrast—depends on the manual search and deduplication having captured the real population of evaluation reports; if low-visibility, non-English, or duplicate reports were systematically missed, the gap estimates could be artifacts of the sampling frame.

What would settle it

Take the 50 least-downloaded models in the annotated dataset and independently inspect every release document and evaluation source in their home languages: if first-party social-impact reporting there matches third-party detail levels, the claimed gap collapses. A second check: re-run the third-party count after requiring that each evaluation be traced to its original dataset or experiment, and confirm whether the post-release sources remain distinct.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X LinkedIn Reddit HN

If this is right

  • Policies that merely ask developers to self-report are unlikely to close the gaps; the paper's evidence that reporting drops when incentives shift points toward mandated disclosure.
  • Independent evaluation ecosystems should be strengthened, since third parties currently carry the most detailed coverage of bias, sensitive content, and performance disparities.
  • Shared infrastructure to aggregate and compare third-party evaluations is needed; the paper shows current evaluation reports are fragmented and static, making model risk profiles hard to assemble.
  • Low-resource and low-visibility models receive far less third-party scrutiny; even accounting for sampling bias, the paper finds whole categories of models left underexamined.
  • Post-release first-party reports tend to be more detailed but arrive too late to inform adoption decisions, so timely reporting, not just eventual reporting, matters.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • By extension, the measured decline in environmental and bias reporting suggests disclosure tracks the perceived reputational or legal cost of negative findings; if so, safe-harbor protections and regulatory sandboxes would be more effective than voluntary templates.
  • The nearly nonexistent third-party coverage of moderation labor and data provenance implies that third-party evaluation cannot substitute for mandated insider disclosure; jurisdictions with such mandates should show higher coverage in these categories.
  • The popularity-driven allocation of third-party scrutiny likely creates a Matthew effect, where less visible models become even less evaluated; subsidizing evaluations for neglected languages and regions would be a natural policy experiment.
  • Because the paper scored presence and detail rather than methodological soundness, its coverage-gap findings could overstate how much is known even where reports exist; a follow-up adequacy review might reveal the true information gap is larger.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

5 major / 4 minor

Summary. The paper presents a large-scale empirical map of social-impact evaluation reporting for foundation models. The authors manually compile 186 first-party release reports and 183 post-release sources (152 fully third-party, 13 fully first-party, 18 mixed), code them on a 0–3 scale across seven social-impact dimensions drawn from Solaiman et al. (2023), and supplement the resulting dataset with ten semi-structured developer interviews. A Bayesian hierarchical ordinal regression is used to estimate how first-party status, openness, sector, and time affect reporting detail. The central claims are that first-party reporting is sparse and often superficial, that it has declined since roughly 2023Q3 for environmental and bias evaluations, and that third-party evaluations provide broader and more detailed coverage for most categories, while data provenance, moderation labor, and cost disclosures remain neglected because only developers can provide them and they are strategically deprioritized. The paper releases its annotated dataset and analysis code.

Significance. If the empirical picture is accurate, the paper makes a valuable contribution to AI governance and accountability: it is the first systematic, cross-provider comparison of first- versus third-party social-impact evaluation reporting, and it supplies a reusable, openly released dataset and code. The statistical model is carefully specified with convergence diagnostics (R-hat, ESS, no post-warmup divergences), and the interview material adds useful context on organizational incentives. The main findings have clear policy relevance, supporting arguments for mandated transparency and independent evaluation infrastructure. However, the central claims rest on the completeness and reliability of a manually constructed sample and manually assigned codes; the paper's own caveats about visibility-biased selection and its explicit statement that no formal inter-annotator agreement was calculated mean that the quantitative core is not yet fully validated.

major comments (5)
  1. [Sec. 4, Sec. 6, App. 8.7] The sampling frame is load-bearing for the headline claims about sparse and declining first-party reporting and comprehensive third-party coverage. The model list is triangulated from FMTI, SaferAI, Hugging Face Hub, LMArena, and Concordia AI, and the third-party search is restricted to Paperfinder peer-reviewed works from 2024 onward plus leaderboard searches; all search terms (App. 8.7) are English, and App. 8.6 concedes that underspecified reporting may hide duplicates and inflate counts. The paper acknowledges in Sec. 6 that the selection 'favors models with higher visibility' but does not provide a sensitivity analysis. If low-visibility or non-English first-party reports contain more environmental/bias evaluations, or if third-party coverage of non-frontier models is systematically missing, the measured first-party/third-party gap and the post-2023Q3 decline could be artifacts. Ple
  2. [App. 8.3] The paper states that annotations were performed by individual researchers with 'manual spot checks' and that 'no formal inter-annotator agreement was calculated.' Since every quantitative conclusion (averages, LORs, time trends) is derived from these 0–3 codes, the measurement reliability of the central variable is unestablished. Even a small dual-coded subset with a reported Cohen's kappa or quadratic-weighted kappa would substantially strengthen the claims. Without this, readers cannot distinguish coding noise from true reporting differences.
  3. [Sec. 5 and Table 3] Sec. 5 states that 'Regression analysis confirms this disparity persists across all seven social impact dimensions,' but Table 3 shows the coefficient for first-party status on category 7 (Data and Content Moderation Labor) is -0.019 with a 95% HDI of [-3.050, 2.526], i.e., not statistically distinguishable from zero. The text lists LORs for only six categories, silently omitting the non-significant seventh. This is an internal inconsistency in a central quantitative finding; the claim should be restricted to the six categories with credible intervals excluding zero, or the inference for category 7 should be explained and interpreted.
  4. [Abstract vs. Sec. 1 / Sec. 4] The abstract states the analysis covers '186 first-party release reports and 248 third-party evaluation sources,' while the introduction and Sec. 4 consistently report '186 first-party release-time reports and 183 post-release reports,' with footnote 3 specifying that only 152 of the 183 are fully third-party. The 248 figure does not appear anywhere in the body and appears to be an error. Since the paper's scope claim is quantitative, this inconsistency should be corrected and the final numbers reconciled.
  5. [Figs. 1–3, 5, 7, 10] The descriptive tables and heatmaps report average scores without any uncertainty intervals, despite small cell counts (e.g., Fig. 3 has quarters with one or two models). The paper's narrative of 'declining' first-party reporting relies heavily on these raw averages, and the claim that Meta's most recent release only included 'vague mentions' is based on score differences that could be within coding noise. Please add credible intervals, standard errors, or annotated sample sizes to the main descriptive figures, or explicitly state that the regression results (which do include uncertainty) are the only basis for the time-trend claims.
minor comments (4)
  1. [App. 8.4] The stratified sample list includes 'AI Singapore' but Table 1 and Figs. 1–2 do not contain this provider; the intended entry is probably 'Ai2' or a similar name. Please align the names across the appendix and main text.
  2. [App. 8.8.1, Eq. (12)] The notation L(q) for the Cholesky factor is used but never explicitly defined in the equation block; a one-line definition would improve reproducibility.
  3. [Table 4] Typo: 'unexpectely' should be 'unexpectedly' in the Nonprofit row.
  4. [Figure 4] The layout of the figure is confusing: the provider names 'Google' and 'Meta' appear to be placed as separate rows rather than as labels for the two heatmaps. Please restructure the figure so each heatmap has an unambiguous header.

Circularity Check

0 steps flagged

No significant circularity: the analysis is an empirical coding of external reports, not a derivation from fitted inputs or a self-citation chain.

full rationale

The paper's central claims—first-party social-impact reporting is sparser and shallower than third-party reporting, with declines over time in categories such as environmental costs and bias—rest on manual annotation of 186 first-party reports and 183 post-release sources against a 0–3 rubric. There is no fitted parameter that is later renamed as a prediction: the Bayesian ordinal regressions only summarize the annotated scores, and the LOR coefficients do not construct the scores they explain. The main self-reference is the choice of the seven social-impact dimensions from Solaiman et al. (2023), whose author list overlaps with the present paper through senior author Irene Solaiman. This is load-bearing in the weak sense that the dimensions define what counts as 'social impact,' but the paper explicitly acknowledges that 'the taxonomy choice is subjective' and gives independent reasons for preferring it over alternatives, rather than importing a uniqueness theorem or unverified prior result to force its conclusions. The coding of individual reports is an independent empirical artifact, and the conclusions about gaps and trends could in principle be wrong if the sampling frame is biased; the paper itself flags that its selection strategy 'favors models with higher visibility' and that duplicates may inflate counts. Those are external-validity limitations concerning the completeness of the source population, not circular reductions of the conclusions to their inputs. There is no step where an equation or fitted value is shown to be equivalent by construction to the reported finding, so the paper is not significantly circular; the minor author-overlapping taxonomy citation justifies at most a score of 1.

Axiom & Free-Parameter Ledger

0 free parameters · 5 axioms · 0 invented entities

This is an observational measurement paper, so the central claims rest on measurement assumptions—taxonomy choice, sampling, annotation reliability, and interview representativeness—rather than on mathematical axioms or new physical entities.

axioms (5)
  • domain assumption Solaiman et al. (2023)'s seven-dimension taxonomy is the appropriate set of social-impact dimensions for foundation models.
    The whole coding scheme is grounded in this taxonomy (Sec. 3). The authors acknowledge other taxonomies exist and that the choice is subjective.
  • domain assumption Absence of reporting in the collected public sources indicates absence of evaluation practice.
    The paper equates 'not reported' with 'gap' throughout Sec. 5. The authors did not have access to internal evaluations, and interviews suggest internal work may go unpublished.
  • domain assumption The manual search and deduplication procedure yields a representative sample of first- and third-party evaluations.
    Model lists are drawn from FMTI, SaferAI, Hugging Face Hub, LMArena, and Concordia AI (Sec. 4). The authors note the selection favors visible models.
  • domain assumption Manual annotation with spot checks is reliable without formal inter-annotator agreement.
    App. 8.3 states no formal inter-annotator agreement was calculated. All downstream scores depend on this subjective coding.
  • domain assumption Quotes from 11 interviewees are representative of developer incentives.
    The interview sample is small and snowball-sampled; the authors state in Sec. 7 that it is not representative of smaller organizations or global majority countries.

reviewed 2026-08-03 · how reviews work

0 comments
Cite this review

Pith. "Pith review of Who Evaluates AI's Social Impacts? Mapping Coverage and Gaps in First and Third Party Evaluations." pith.science (2026). https://pith.science/paper/54DFY7NL

@misc{pith2026251105613,
  author       = {Pith},
  title        = {Pith review of: Who Evaluates AI's Social Impacts? Mapping Coverage and Gaps in First and Third Party Evaluations},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/54DFY7NL}},
  note         = {Machine review of arXiv:2511.05613}
}
Share X LinkedIn Reddit HN
read the original abstract

Foundation models are increasingly central to high-stakes AI systems, and governance frameworks now depend on evaluations to assess their risks and capabilities. Although general capability evaluations are widespread, social impact assessments covering bias, fairness, privacy, environmental costs, and labor remain uneven. To characterize this landscape, we conduct the first comprehensive analysis of social impact evaluation reporting, examining 186 first-party release reports and 248 third-party evaluation sources, supplemented by developer interviews. We find a stark division of labor: first-party reporting is sparse, often superficial, and declining in areas like environmental impact and bias, while third-party evaluators provide broader, more rigorous coverage of bias, harmful content, and performance disparities. However, only developers can authoritatively report on data provenance, content moderation labor, costs, and infrastructure, yet interviews reveal these disclosures are deprioritized unless tied to product adoption or compliance. Current practices leave major gaps in assessing societal impacts, underscoring the need for policies that mandate developer transparency, strengthen independent evaluation ecosystems, and create shared infrastructure for aggregating third-party evaluations.

Figures

Figures reproduced from arXiv: 2511.05613 by Anastassia Kornilova, Andrew Tran, Angelina Wang, Anka Reuel, Arushi Saxena, Avijit Ghosh, Eliya Habba, Felix Friedrich, Hossein A. Rahmani, Irene Solaiman, Jan Batzner, Jeba Sania, Jennifer Mickel, Jenny Chim, Kevin Klyman, Kevin Wei, Michael Alexander Riegler, Mowafak Allaham, Mubashara Akhtar, Mykel Kochenderfer, Olivia Beyer Bruvik, Pawan Sasanka Ammanamanchi, Pouya Sadeghi, Prajna Soni, Robert Scholz, Sanmi Koyejo, Srishti Yadav, Stella Biderman, Subramanyam Sahoo, Sujata Goswami, Usman Gohar, Yacine Jernite, Yanan Long, Yohan Mathew, Zeerak Talat.

Figure 1
Figure 1. Figure 1: Average scores for first-party social impact reporting per provider. Color indicates the reporting detail level (lightest green = lowest scores, medium green = mid scores, darkest green = highest scores) (see Scoring in Section 4 for details). In the case of multiple models per provider, we report the average detail level of the evaluation reporting. For clarity, we present results for a stratified sample … view at source ↗
Figure 2
Figure 2. Figure 2: Average scores and counts for third party social impact evaluations for stratified sample of providers. Rectangle size corresponds log-linearly to the number of evaluations, and color indicates the average reporting detail level (lightest blue = lowest scores, medium blue = mid scores, darkest blue = highest scores). Each cell displays the score (bold) and evaluation count (in parentheses) (see Scoring in … view at source ↗
Figure 3
Figure 3. Figure 3: Average scores for first-party social impact reporting over time per release quarter. The number after the release quarter in parentheses denotes the number of models released that quarter in our dataset. Color indicates the average reporting detail level (lightest green = lowest scores, medium green = mid scores, darkest green = highest scores) (see Scoring in Section 4 for details). The full results for … view at source ↗
Figure 4
Figure 4. Figure 4: Reporting detail level across social impact categories within select providers over model releases. Color indicates the reporting detail level (lightest green = lowest scores, medium green = mid scores, darkest green = highest scores). (see Scoring in Section 4 for details) Sensitive content evaluations see the most attention across first- and third-party reports. Sensitive content was, on average, the mos… view at source ↗
Figure 5
Figure 5. Figure 5: Average scores for first-party social impact reporting per sector. Color indicates the average reporting detail level (lightest green = lowest scores, medium green = mid scores, darkest green = highest scores) (see Scoring in Section 4 for details). Full analysis results in App. 8.1. not a lot of evaluation or real discussion happening around [data and content moderation labor reporting], though it deserve… view at source ↗
Figure 6
Figure 6. Figure 6: Average scores and counts for social impact evaluations by provider geographic region. Rectangle size corresponds log-linearly to the number of evaluations, and color indicates the average reporting detail level (lightest blue = lowest scores, medium blue = mid scores, darkest blue = highest scores) (see Scoring in Section 4 for details). Each cell displays the score (bold) and evaluation count (in parenth… view at source ↗
Figure 7
Figure 7. Figure 7: First-party social impact reporting by model provider country. Color indicates the average reporting detail level (lightest green = lowest scores, medium green = mid scores, darkest green = highest scores) (see Scoring in Section 4 for details). "EU" in Figures 7 and 8 denote a Multi-Country EU Consortium. Bias & Harm Sensitive Content Performance Disparity Env. Costs & Emissions Privacy & Data Financial C… view at source ↗
Figure 8
Figure 8. Figure 8: Third-party social impact reporting by model provider country. Rectangle size corresponds log-linearly to the number of evaluations for each country (EU denotes Multi Country EU Consortium), and color indicates the average reporting detail level (lightest blue = lowest scores, medium blue = mid scores, darkest blue = highest scores). Each cell displays the score (bold) and evaluation count (in parentheses)… view at source ↗
Figure 9
Figure 9. Figure 9: Evaluation scores for third-party social impact evaluation by level of openness: either open (both open￾source and open-weight) and closed. Rectangle size corresponds log-linearly to the number of evaluations, and color indicates the average reporting detail level (lightest blue = lowest scores, medium blue = mid scores, darkest blue = highest scores). Each cell displays the score (bold) and evaluation cou… view at source ↗
Figure 10
Figure 10. Figure 10: Average scores for first-party social impact reporting per provider. Color indicates the average reporting detail level (lightest green = lowest scores, medium green = mid scores, darkest green = highest scores) (see Scoring in Section 4 for details). In the case of multiple models per provider, we report the average detail level of the evaluation reporting. 18 [PITH_FULL_IMAGE:figures/full_fig_p018_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: Evaluation scores for third-party social impact evaluations per provider for all providers (contd.). Rectangle size corresponds log-linearly to the number of evaluations, and color indicates the average reporting detail level (lightest blue = lowest scores, medium blue = mid scores, darkest blue = highest scores). Each cell displays the score (bold) and evaluation count (in parentheses). (see Scoring in S… view at source ↗
Figure 12
Figure 12. Figure 12: Stacked bar chart of variance decomposition by group-level effects contribution, shown per outcome and [PITH_FULL_IMAGE:figures/full_fig_p028_12.png] view at source ↗
Figure 13
Figure 13. Figure 13: 95% high density interval (HDI) of the posterior distribution of the regression coefficients of the slopes. Black parallelograms indicate the posterior medians. The outcomes are coded as above: 1 = Bias, Stereotypes, and Representational Harms, 2 = Sensitive Content, 3 = Disparate Performance, 4 = Environmental Costs and Carbon Emissions, 5 = Privacy and Data Protection, 6 = Financial Costs, 7 = Data and … view at source ↗
Figure 14
Figure 14. Figure 14: Standardized effect of year on the linear predictor on the logistic link scale. [PITH_FULL_IMAGE:figures/full_fig_p030_14.png] view at source ↗
Figure 15
Figure 15. Figure 15: Standardized effect of year on the linear predictor by outcome category. Each panel shows posterior mean [PITH_FULL_IMAGE:figures/full_fig_p030_15.png] view at source ↗
Figure 15
Figure 15. Figure 15: Standardized effect of year on the linear predictor by outcome category [PITH_FULL_IMAGE:figures/full_fig_p031_15.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Informing AI Policy Assessment using Large-Scale Simulation of Interventions

    cs.CY 2026-04 conditional novelty 6.5

    A genetic algorithm optimizes weighted combinations of LLM-perceived harm mitigation, expert costs, and participatory scores over stakeholder-action pairs to surface viable AI policy packages for media harms.

  2. Informing AI Policy Assessment using Large-Scale Simulation of Interventions

    cs.CY 2026-04 conditional novelty 5.0

    A genetic algorithm exploring billions of policy combinations, scored by LLM-evaluated harm mitigation, expert cost, and participatory ratings, identifies viable AI policy options under different weighting schemes.

Reference graph

Works this paper leans on

40 extracted references · 1 linked inside Pith · cited by 1 Pith paper

  1. [1]

    Reactions to observed trends in social impact evaluation reporting as discussed in Section 5

  2. [2]

    Incentives and motivations for conducting social impact evaluations and reporting results

  3. [3]

    Barriers to running social impact evaluations and reporting results

  4. [4]

    The specific questions asked to our interviewees were designed to support the following key objectives:

    Ideal future directions in social impact evaluations. The specific questions asked to our interviewees were designed to support the following key objectives:

  5. [5]

    What would incentivize your organization to report more in this dimension?

  6. [6]

    How does granularity and specificity impact feasibility with regard to running social impact evaluations? Future Directions

  7. [7]

    Uncover qualitative insights regarding the quantitative results reported in Section 5

  8. [8]

    Understand practitioner viewpoints on why social impact evaluations are lacking

  9. [9]

    Interviewees were recruited based on their experience conducting or developing evaluations

    Identify incentives to encourage increased social impact evaluation reporting. Interviewees were recruited based on their experience conducting or developing evaluations. These participants were sourced through the authors’ collective professional contacts and subsequent snowball sampling. All participating interviewees (see Section 8.10.2) had direct exp...

  10. [10]

    With respect to conducting and reporting evaluations, what are relevant responsibilities and job tasks that fall within your scope of responsibility?

  11. [11]

    Does your role involve one or more of the following: model specification design, model development, deployment, model evaluation, general model governance, and/or acceptable use? 32 Preprint

  12. [12]

    Have you been directly involved in designing, implementing, reviewing, or reporting evaluations in the past 2–3 years? (Or have you had visibility into the decision processes leading up to these activities?) Incentives/Motivations for Reporting

  13. [13]

    Why do you think the observed trends (demonstrated in the plots developed by team #4) in reporting social impact evaluations are occurring? Interviewer note: Show the plots prior to asking this question

  14. [14]

    What currently motivates, if anything, your organization (or you) to conduct evaluations?

  15. [15]

    Barriers to Adoption

    Are there any incentives that would encourage you or your organization to conduct more thorough evaluations for social impact? Notes: Possible incentives include positive media attention, regulation, higher internal capacity, or knowledge that customers/users care about specific evaluations. Barriers to Adoption. This is the central portion of the intervi...

  16. [16]

    Given the following Likert scale, where would you rank each of these categories in terms of importance and feasibility? (a) 1 = Very easy / very feasible (b) 2 = Easy / feasible (c) 3 = Somewhat easy / somewhat feasible (d) 4 = Neutral (e) 5 = Somewhat hard / somewhat infeasible (f) 6 = Hard / infeasible (g) 7 = Very hard / very infeasible Interviewer not...

  17. [17]

    What barriers do you face to reporting these categories (e.g., time, legal risk, reputational/investor harm, resources)?

    In our analysis of evaluations reported across social impact categories, we observed missing evaluations in certain areas (e.g., ____). What barriers do you face to reporting these categories (e.g., time, legal risk, reputational/investor harm, resources)?

  18. [18]

    Which is the most pressing barrier?

  19. [19]

    Which factors or underlying reasons need to change to enable greater reporting?

  20. [22]

    What does the ideal social impact evaluation look like in terms of breadth and specificity?

  21. [23]

    Can you provide an example?

  22. [24]

    Are there any social impact categories missing that should be reported?

  23. [25]

    Who should be responsible for evaluations?

  24. [26]

    Who should be responsible for developing, running, or setting standards for evaluations?

  25. [27]

    To what extent would a standardized reporting template help enable more comprehensive social impact evaluation reporting? Interviewer note: Show the current evaluation card design

  26. [28]

    magic wand,

    (Optional) If you had a “magic wand,” what would the ideal process/tool look like for reporting social impact evaluations? Current Practices (Organizational). This section can be skipped if not relevant

  27. [29]

    Who is ultimately responsible for conducting social impact evaluations in your organization?

  28. [30]

    Which social impact evaluations are conducted? What criteria/processes shape the choice, and who are the stakeholders?

  29. [31]

    What criteria/processes exist for deciding which evaluation results get publicly reported? 33 Preprint

  30. [32]

    Describe the types of social impact evaluations, if any, that are reported when documenting AI systems

  31. [33]

    Are there evaluations conducted internally that were not reported? If so, which, and why not?

    Based on your most recent model/system card, we found these social impact evaluations: [share screen]. Are there evaluations conducted internally that were not reported? If so, which, and why not?

  32. [34]

    Are social impact evaluations broad or granular (e.g., examining bias in specific contexts)?

  33. [35]

    use-case-independent?

    For the social impact evaluations displayed, to what extent are they context-specific vs. use-case-independent?

  34. [36]

    general/agnostic tasks?

    How do you select or design evaluations? To what extent do you include/prioritize use-case-specific vs. general/agnostic tasks?

  35. [37]

    To what extent do you rely on third-party evaluators or contractors?

  36. [38]

    To what extent do internal evaluations differ from those released externally? If differences exist, what practices/criteria determine release (e.g., legal, PR)? Closing Questions

  37. [39]

    Is there any question we should have asked, or anything with respect to social impact evaluations that we have not covered yet?

  38. [40]

    Is there anything else you think we should know before we end this call? 8.11 CODE ANDDATASET Our annotated social impact eval dataset is available at https://huggingface.co/datasets/evaleval/s ocial_impact_eval_annotations , and the analysis code to reproduce the results and plots in this paper is accessible athttps://github.com/evaleval/social_impact_ev...

  39. [2021]

    scheming

    doi: 10.1145/3411764.3445518. URLhttps://doi.org/10.1145/3411764.3445518. Roy Schwartz, Jesse Dodge, Noah A Smith, and Oren Etzioni. Green ai.Communications of the ACM, 63(12):54–63, 2020. Reza Shokri, Marco Stronati, Congzheng Song, and Vitaly Shmatikov. Membership inference attacks against machine learning models. In2017 IEEE symposium on security and p...

  40. [2024]

    everyone wants to do the model work, not the data work

    URLhttps://arxiv.org/abs/2405.15802. Henk A. Becker. Social impact assessment.European Journal of Operational Research, 128(2):311–321, 2001. doi: 10.1016/S0377-2217(00)00074-6. Emily M. Bender and Batya Friedman. Data statements for natural language processing: Toward mitigating system bias and enabling better science.Transactions of the Association for ...

This paper was first reviewed by deepseek-v4-flash on August 3, 2026.