Pith. sign in

REVIEW 4 major objections 4 minor 25 references

The Societal Impact of Foundation Models: Advancing Evidence-based AI Policy

T0 review · 4 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read This dissertation claims that AI governance becomes tractable once foundation models are named as a paradigm, evaluated at the model level, and scored at the developer level, and it demonstrates policy tools that use that evidence.

desk verdict A candid, well-integrated dissertation that repackages the author's influential prior work into a coherent framework for AI governance; the empirical anchor is weaker than the rhetoric, but the author knows it. read the letter →

arxiv 2506.23123 v1 pith:VI5ARAKX submitted 2025-06-29 cs.AI cs.CYcs.ET

classification cs.AIcs.CYcs.ET
keywords foundationmodelsAIgovernanceevidence-basedpolicymodelevaluationtransparencyindexalgorithmicmonocultureemergentcapabilitiessupplychain
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This dissertation argues that the social harms and benefits of foundation models are not something to be debated only from first principles: they can be studied, measured, and fed into policy. The central claim is that three ingredients—a clear concept of foundation models, model-level evaluation, and organization-level transparency measurement—together give governments and the public a workable evidence base for governing AI. The dissertation builds each ingredient: a definition of foundation models organized around emergence and homogeneity, the HELM evaluation platform that measures models on accuracy plus robustness, fairness, calibration, toxicity, and efficiency, and the Foundation Model Transparency Index that scores developers on 100 indicators across the supply chain. It then shows how that evidence can enter policymaking through structured evidence review, compliance pre-assessment against the EU AI Act, and policy designs that generate new evidence. A sympathetic reader would take the dissertation to establish that evidence-based AI governance is feasible now, rather than contingent on future breakthroughs.

What carries the argument

The load-bearing object is the concept of a foundation model, defined as a model trained on broad data via self-supervision at scale and adapted to many downstream tasks. Around it, the dissertation builds two measurement instruments: HELM, an evaluation platform that assesses every model on the same scenarios and desiderata (accuracy, calibration, robustness, fairness, bias, toxicity, and efficiency), and the Foundation Model Transparency Index, a composite index that turns the abstract construct of transparency into 100 scored indicators spanning upstream (data, compute, labor), model, and downstream (distribution, usage, impact) domains. The bridging mechanism is the research-policy interface: evidence reviews that map marginal risk, compliance pre-assessments that grade current practice against proposed law, and policy designs that generate new evidence. The concept does the work of making different AI systems comparable as instances of one paradigm, while the instruments make the paradigm's societal exposure measurable.

What would settle it

Observation: build a failure matrix from a cohort of deployed systems that all draw on the same foundation model across independent organizations, and compare the observed rate at which every system fails for the same individual against the Poisson-binomial baseline computed from each system's overall failure rate. If the observed systemic failure rate does not exceed the baseline, the dissertation's monoculture-harm argument for foundation models would be refuted.

Watch

Extended reading notes

Core claim

The central discovery is that the foundation model paradigm's two defining properties—emergence, by which capabilities appear abruptly at scale, and homogeneity, by which many systems come to depend on the same model—are not just technical curiosities; they are the points where society becomes exposed, and the places where transparency instruments can intervene. On this basis, the dissertation claims that measuring models through standardized third-party evaluation (HELM) and measuring developers through a 100-indicator transparency index (FMTI) makes the opaque ecosystem legible enough to regulate. The concrete payoff is demonstrated in the policy chapter: an evidence review of open foundation models that separates speculation from documented harm, a compliance pre-assessment that scored ten providers against the EU AI Act's requirements, and a call for evidence-generating policy that treats disclosure requirements as instruments for producing better information. The dissertation does not claim to resolve disagreements about AI's future; it claims to create the shared factual basis on which such disagreements can be productive.

Load-bearing premise

The load-bearing premise is that failure patterns found in older, task-specific commercial AI systems also hold for foundation models; the dissertation openly says the direct evidence for foundation models had not yet arrived.

Editorial extensions

If this is right

  • If the dissertation is right, regulators do not need to wait for perfect foresight: standardized evaluation of models and developers can give them an operational view of capabilities, risks, and supply-chain exposure.
  • The EU AI Act's transparency and accountability requirements are actionable now, because the compliance pre-assessment shows that current providers score far below the achievable maximum, so enforcement would materially change ecosystem behavior.
  • Open versus closed foundation model debates can be grounded in marginal-risk analysis rather than ideology, since the evidence review separates documented harm from unsubstantiated speculation.
  • Supply-chain monitoring through Ecosystem Graphs means a faulty upstream asset (data, compute, or model) can in principle be traced to the applications and users affected, following the same logic as automobile recalls.
  • The improvement in FMTI scores between 2023 and 2024 indicates that public scoring itself pressures developers to disclose, meaning measurement can change behavior.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If HELM-style evaluation and FMTI-style indexing were institutionalized inside government safety institutes, they could function like clinical-trial registries, making safety claims about models auditable over time.
  • The homogeneous-outcomes metric is a ready-made early-warning indicator: regulators watching a shared foundation model could track whether the systemic failure rate rises above the independence baseline and treat that as a market-structure red flag.
  • Compliance pre-assessment against the EU AI Act could be re-run as the law evolves, producing a longitudinal record of whether regulation actually changes developer behavior.
  • The FMTI's correlation structure across developers hints that opacity is partly a sector-wide equilibrium, which would argue for coordinated disclosure mandates rather than firm-by-firm persuasion.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The manuscript is a PhD dissertation posted on arXiv that synthesizes the author's prior research on foundation models and AI governance. It argues that foundation models constitute a distinct AI paradigm and that their societal impact can be understood through three interlocking contributions: a conceptual framing of capabilities, risks, and supply chains; empirical transparency mechanisms at the model level (HELM) and organizational level (FMTI); and policy-facing analyses including an open-models risk review and an EU AI Act compliance pre-assessment. The abstract claims that together these contributions 'make inroads into achieving better societal outcomes' by building the scientific foundations and research-policy interface for evidence-based AI policy. The dissertation is a narrative compilation rather than a presentation of new derivations or data analyses; each chapter summarizes previously published works and adds reflective discussion of their policy uptake.

Significance. If the claims are accepted, the dissertation provides a widely adopted vocabulary ('foundation model'), a standardized model-evaluation infrastructure (HELM), a replicable transparency-index methodology (FMTI), and a template for pre-legislative compliance assessment. A notable strength is the author's explicit acknowledgment of limitations, particularly in the homogeneous-outcomes vignette, and the manuscript gives a clear, honest account of what was and was not possible to measure during the PhD. The main significance is integrative: it shows a plausible pathway from technical measurement to policy engagement, and several of its constituent works have demonstrably entered policy discourse (e.g., NIST guidance, CMA market surveillance). However, because the dissertation itself contains no new empirical analyses, the significance for a research venue depends on the credibility of the underlying prior works and on the strength of the extrapolations that connect them to the central claim of advancing evidence-based AI policy.

major comments (4)
  1. [§3.2.2, Eq. (3.2)] The systemic-risk argument for foundation-model-specific policy rests on an extrapolation from HAPI, which audits three commercial task-specific ML APIs per modality (sentiment analysis, facial emotion recognition, spoken command recognition). The author explicitly acknowledges that 'the conditions to study this rigorously were not available during the time of my PhD' and that the evidence concerns 'other older forms of widespread AI.' Because foundation models share pretraining data, architectures, and deployment patterns in ways that task-specific APIs do not, the independence baseline in Eq. (3.2) does not establish that homogeneous outcomes will be as prevalent for foundation models. This is a load-bearing gap for the claim that the dissertation supplies the empirical foundation for foundation-model-specific governance. The author should either add a direct replication on foundation-model APIs (the data conditions have partly materialized by 2025) or carefully re-scope the policy conclusions to AI systems in general rather than foundation models in particular.
  2. [§1.3, §3.4, §6] The evidence that the framework 'advances evidence-based AI policy' consists largely of the author's own descriptions of briefings, media coverage, and citations in policy documents. This creates a self-referential loop: the dissertation's central claim is supported by the impact of the very works that the dissertation summarizes. Independent evidence of impact would strengthen the claim, for example traceable changes in company disclosure practices, specific provisions in the EU AI Act or the G7 Code of Conduct that can be causally linked to the framework, or a third-party evaluation of the FMTI's effect on disclosure behavior. At minimum, the author should distinguish clearly between 'uptake' (citations, invitations, coverage) and 'impact' (documented change in policy or practice) and acknowledge the limits of self-reported engagement as evidence for the central claim.
  3. [§5.2.3, §5.3] The FMTI scoring uses 100 equally weighted binary indicators, and the overall ranking is a direct function of this equal weighting. Equal weighting is a free parameter: the manuscript does not provide a sensitivity analysis or a justification grounded in a measurement model. The same issue applies to the selection of HELM scenarios (§4.2.1) and metrics (§4.2.2), where the 'systematic' procedure ultimately depends on the author's judgments (e.g., what counts as user-facing, the desiderata taxonomy, and the mapping from venues to desiderata). For the index to support the claim that it 'concretizes' the nebulous construct of transparency, the author should provide robustness checks (e.g., alternative weightings, leaving-one-out analysis over indicators) and a discussion of the construct validity of the indicator set.
  4. [§6.2] The EU AI Act compliance pre-assessment assigns grades on a 0-4 scale to ten providers across twelve requirements, with the resulting scores presented in Figure 6.1. No inter-rater reliability, independent legal validation, or sensitivity analysis is reported, and the scoring rubric appears to be defined by the author and collaborators. Given that this assessment was offered as 'immediate feedback for policy decisions in the legislative process,' the method should be either validated against independent legal analysis or explicitly presented as a research proposal rather than a compliance measurement.
minor comments (4)
  1. [§1.2] The phrase 'the the primary mechanism' should read 'the primary mechanism'.
  2. [Table 3.1] The columns 'Train. FLOPs', 'Params.', and 'Model' are not clearly separated in the rendered text; consider reformatting the table for readability.
  3. [Figure 3.4b] The axes are not labeled; the caption refers to the line y=x, but the reader cannot verify this without axis labels.
  4. [§3.4.1] The claim that policymakers 'woefully misunderstand poorly communicated science' is a strong reflective assertion; consider supporting it with a direct quotation or citation to the House SST conclusion rather than a footnote URL.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: empirical chapters are grounded in external data and benchmarks; the acknowledged extrapolation in the homogeneous-outcomes vignette is a limitation, not a circular derivation.

full rationale

The dissertation presents a conceptual framework (foundation models, emergence, homogeneity, supply chain), two empirical instruments (HELM, FMTI), and policy applications. The load-bearing empirical steps are not defined in terms of the conclusions. HELM evaluates 30 models on external scenarios and metrics; FMTI scores developers against 100 independently specified disclosure indicators. The homogeneous-outcomes argument explicitly states that 'the conditions to study this rigorously were not available' and instead reports HAPI observations of commercial ML APIs; this is an extrapolation and acknowledged limitation, not a fitted input renamed as a prediction. The policy-impact claims reference external bodies (NIST, UK CMA, US Congress, EU AI Act) rather than deriving solely from the author's own indices. Self-citations are pervasive, as expected for a dissertation 'largely based on my prior writings,' but none functions as a uniqueness theorem or as an unverified premise on which the conclusion depends. No equation or metric reduces to the paper's own inputs by construction, so the derivation chain is self-contained enough to avoid circularity.

Assumptions & free parameters 2 free parameters · 4 assumptions · 3 invented entities

The dissertation rests on several conceptual and measurement choices that are not derived from first principles: the definition of foundation models, the assumption that benchmark scores and transparency indices capture socially meaningful properties, and the extrapolation of findings from older AI systems to foundation models.

free parameters (2)
  • FMTI indicator set and equal weighting = 100 binary indicators, equal weight (implicit)
    The Foundation Model Transparency Index scores each developer on 100 hand-picked binary indicators, aggregated with equal weight. The choice of indicators and weighting is a modeling decision that determines all FMTI scores and policy conclusions drawn from them.
  • HELM scenario and metric selection = 16 core scenarios, 7 metrics (2022)
    The HELM evaluation platform selects a small set of scenarios and metrics based on authorial judgment about user-facing tasks and feasibility. This selection determines the measured capabilities and risks.
assumptions (4)
  • domain assumption Foundation models constitute a coherent, meaningful technological paradigm
    The dissertation's entire framing rests on the claim, introduced in Bommasani et al. (2021), that foundation models are a distinct class with shared properties. This is a conceptual choice, not a theorem.
  • domain assumption Benchmark evaluations are valid proxies for real-world capabilities and risks
    HELM relies on benchmark scenarios (e.g., MMLU, RAFT) despite the author's own acknowledgment that 'standard benchmarks are poor proxies' (§3.1.1). The policy recommendations assume these measurements matter.
  • domain assumption Transparency, as operationalized by FMTI, is a meaningful and desirable property
    The FMTI assumes that the 100 disclosed items capture the relevant aspects of transparency and that more transparency is beneficial. No external validation is provided.
  • domain assumption Homogeneous outcomes in older commercial APIs will generalize to foundation models
    The evidence for homogeneous outcomes comes from HAPI (2020-2022 commercial APIs), not foundation models. The dissertation explicitly notes this gap and asserts it will manifest in foundation models due to monoculture.
invented entities (3)
  • Foundation model (concept) independent evidence
    purpose: Rename and reframe a class of large pretrained models as a technological paradigm
    The term has been adopted in academic literature and government documents (e.g., US Executive Order 14110), providing external uptake. However, it remains a definitional concept, not a falsifiable entity.
  • Foundation Model Transparency Index (FMTI)
    purpose: To measure organizational transparency of AI developers
    FMTI is a composite score constructed by the author. It predicts nothing outside itself; its scores are the measurement, with no independent outcome to validate them.
  • Homogeneous outcomes metric
    purpose: To quantify systemic failure across multiple AI systems
    The metric is used to assess commercial APIs, but it is not a physical or biological entity. Its independence is unclear; the concept is an interpretation of data.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The Societal Impact of Foundation Models: Advancing Evidence-based AI Policy." pith.science (2026). https://pith.science/paper/VI5ARAKX

@misc{pith2026250623123,
  author       = {Pith},
  title        = {Pith review of: The Societal Impact of Foundation Models: Advancing Evidence-based AI Policy},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/VI5ARAKX}},
  note         = {Machine review of arXiv:2506.23123}
}
read the original abstract

Artificial intelligence is humanity's most promising technology because of the remarkable capabilities offered by foundation models. Yet, the same technology brings confusion and consternation: foundation models are poorly understood and they may precipitate a wide array of harms. This dissertation explains how technology and society coevolve in the age of AI, organized around three themes. First, the conceptual framing: the capabilities, risks, and the supply chain that grounds foundation models in the broader economy. Second, the empirical insights that enrich the conceptual foundations: transparency created via evaluations at the model level and indexes at the organization level. Finally, the transition from understanding to action: superior understanding of the societal impact of foundation models advances evidence-based AI policy. View together, this dissertation makes inroads into achieving better societal outcomes in the age of AI by building the scientific foundations and research-policy interface required for better AI governance.

Figures

Figures reproduced from arXiv: 2506.23123 by the authors.

Figure 4.15
Figure 4.15. [PITH_FULL_IMAGE:figures/full_fig_p020_4_15.png] view at source ↗
Figure 1.1
Figure 1.1. The rapid adoption of ChatGPT. Within two months of release, OpenAI’s ChatGPT accumulated 100 million users, reflecting a faster rate of adoption than any prior digital technology. ChatGPT accumulated 100 million global active users in 2 months, outpacing the world’s most iconic digital technologies to date ( [PITH_FULL_IMAGE:figures/full_fig_p027_1_1.png] view at source ↗
Figure 1.2
Figure 1.2. The growth in AI company stocks. Following ChatGPT’s launch, val￾uations of AI companies (especially Nvidia) skyrocketed. Figure used with permission. billions, better resembling the large industrial machinery seen in more mature product sectors than the simple conceptual frameworks taught in academic machine learning courses. In the 2010s, companies espoused ideals (e.g. Google, IBM, and Microsoft published respons… view at source ↗
Figures from the paper (66 more)
Figure 2.1
Figure 2.1. Figure 2.1: Emergence and homogenization as a lens to identify AI paradigms. The trajectory of AI is marked by increasing emergence and homogenization. Machine learning demonstrates that the ability to perform tasks emerge from data and that the methods to develop models homogen…
Figure 2.2
Figure 2.2. Figure 2.2: The foundation model paradigm. A foundation model can centralize the information from all the data from various modalities. This one model can then be adapted to a wide range of downstream tasks. 2025; Kokotajlo et al., 2025; Bengio et al., 2025) [PITH_FULL_IMAGE:fi…
Figure 2.3
Figure 2.3. Figure 2.3: The foundation model ecosystem. The foundation model ecosystem stretches from data creation to deployment. At both ends, we highlight the role of people as the ultimate source of data into training of a foundation model, but also as the downstream recipients of any b…
Figure 3.1
Figure 3.1. Figure 3.1: Emergent capabilities. We plot the performance of different models to depict emergent capabilities. Figure taken from Wei et al. (2022a) [PITH_FULL_IMAGE:figures/full_fig_p056_3_1.png]
Figure 3.2
Figure 3.2. Figure 3.2: Emergent interventions. We depict the performance of different models when using different adaptation techniques to demonstrate the emergent meta-level capability to make use of these interventions. Figure taken from Wei et al. (2022a) [PITH_FULL_IMAGE:figures/full_…
Figure 3
Figure 3. Figure 3: C, on 8-digit addition, using a scratchpad only improves program execution [PITH_FULL_IMAGE:figures/full_fig_p057_3.png]
Figure 3
Figure 3. Figure 3: D–H) [PITH_FULL_IMAGE:figures/full_fig_p059_3.png]
Figure 3.3
Figure 3.3. Figure 3.3: Ecosystem-level analysis. Individuals interact with decision-makers (left), receiving outcomes that constitute the failure matrix (right). or systems. For example, when a candidate applies to jobs, they typically apply to several firms. The decision each company make…
Figure 3.4
Figure 3.4. Figure 3.4: Evidence of homogeneous outcomes in deployed AI systems. Ecosystem-level analysis surfaces the general trend of homogeneous outcomes: the observed rates that all models succeed/fail consistently exceeds the corresponding baseline rates [PITH_FULL_IMAGE:figures/full_…
Figure 3.5
Figure 3.5. Figure 3.5: Examples of homogeneous outcomes. Instances that are sampled uniformly at random from “0 models correct” (top row) or ”3 models correct” (bottom row) in fer+. The systemic failures (top row) do not appear to be inherently harder for humans to classify; more extensive…
Figure 3.6
Figure 3.6. Figure 3.6: The supply chain begins with the upstream resources that are used to build [PITH_FULL_IMAGE:figures/full_fig_p075_3_6.png]
Figure 3.6
Figure 3.6. Figure 3.6: The canonicalized foundation model supply chain. A conceptual depiction of the foundation model supply chain, beginning with the primary upstream resources (i.e. data, compute) and transitioning to the foundation model, subsequent hosts (or distribution channels), an…
Figure 3.7
Figure 3.7. Figure 3.7: Spectrum of release options. In releasing a foundation model, devel￾opers choose what to release, ranging from nothing (i.e. the model is only accessible to developer-internal employees) to everything (i.e. the weights, data, and code are all available without restri…
Figure 4.1
Figure 4.1. Figure 4.1: Example of question answering. An example instance for question an￾swering from MMLU. Different QA scenarios can have significantly different properties, but this example captures the overall structure of question answering. In QA, given a question (e.g. “Where was t…
Figure 4.2
Figure 4.2. Figure 4.2: Example of information retrieval (passage ranking). An example instance for information retrieval from MS MARCO. We focus here on the passage ranking task: given a query q and a large corpus C of passages, systems must output a list of the top-k passages from C in de…
Figure 4.3
Figure 4.3. Figure 4.3: Example of summarization. An example instance for summariza￾tion from CNN/DailyMail. Different summarization scenarios can have significantly different properties, but this example captures the overall structure of summarization. where a document (e.g. a CNN news art…
Figure 4.4
Figure 4.4. Figure 4.4: Example of sentiment analysis. An example instance for sentiment analysis from IMDB. Sentiment analysis. Sentiment analysis is an iconic task in NLP (see Jurafsky and Martin, 2000, §4) that has led to widespread deployment in finance, health, social media, with appli…
Figure 4.5
Figure 4.5. Figure 4.5: Example of toxicity detection. An example instance for toxicity detection from CivilComments. concern that automated moderation could exacerbate the problem. Akin to sentiment analysis, for toxicity detection we consider the binary classifica￾tion problem of determin…
Figure 4.6
Figure 4.6. Figure 4.6: Example of miscellaneous text classification. An example instance for miscellaneous text classification from RAFT (subset=Banking77). Miscellaneous text classification. Text classification and categorization refers to the family of NLP tasks where an input sequence (…
Figure 4.6
Figure 4.6. Figure 4.6: Unlike other tasks, essentially by construction, it is near-impossible to enumerate, let alone represent, all the non-standard text classification tasks that are useful. For this reason, we turn to RAFT (Alex et al., 2021), which is a collection of 11 ecologically￾va…
Figure 4.7
Figure 4.7. Figure 4.7: Calibration metrics. A demonstration of how we measure calibration and selective classification. The model probabilities refer to the probabilities the model assigns to its prediction. For simplicity, the figure uses 2 bins for ECE computation, but we use 10 bins in …
Figure 4.8
Figure 4.8. Figure 4.8: Robustness perturbations. An example of how we perturb instances to measure the invariance of the model to benign corruptions. 2021; Dhole et al., 2021; Wang et al., 2021b). Towards this goal, we measure the robustness of different models by evaluating them on transf…
Figure 4.9
Figure 4.9. Figure 4.9: Fairness Perturbations. An example of how we perturb examples to measure fairness with respect to subject properties (e.g. the gender of the entities mentioned in the text). the original BoolQ contrast sets). Fairness. The disparate treatment and disparate impact (Ba…
Figure 4.10
Figure 4.10. Figure 4.10: Bias metrics. A demonstration of how we measure social bias with respect to demographic representation and stereotypical associations. technologies. We reiterate the point raised in Rauh et al. (2022) that the norms for language technologies need not be the same as …
Figure 4.11
Figure 4.11. Figure 4.11: Toxicity metrics. A demonstration of how we measure toxicity of language model predictions [PITH_FULL_IMAGE:figures/full_fig_p120_4_11.png]
Figure 4.12
Figure 4.12. Figure 4.12: Inference efficiency metrics. A demonstration of how we measure inference efficiency. We compute two metrics: denoised inference runtime and idealized inference runtime. F returns the runtime of encoding a prompt of given size, and g is the runtime of generating eac…
Figure 4.13
Figure 4.13. Figure 4.13: Accuracy vs. X. The relationship between accuracy (x-axis) and each of the 6 metrics (calibration, robustness, fairness, social bias, toxicity, efficiency) we study in this work across all core scenarios and for all models. For calibration error, we measure ECE-10; …
Figure 4.14
Figure 4.14. Figure 4.14 [PITH_FULL_IMAGE:figures/full_fig_p129_4_14.png]
Figure 4.15
Figure 4.15. Figure 4.15: Head-to-head win rate per each model. We report the fraction of head-to-head comparisons between the given model and all other models, across all scenarios, where the given model is higher along the metric (e.g. more accurate in the accuracy subfigure). If a model w…
Figure 4.16
Figure 4.16. Figure 4.16: Cumulative accuracy over time. The relationship between time (x-axis) and the accuracy of the most accurate model released up to that point (y-axis) across 16 core scenarios. That is, the graph tracks the progress in the state-of-the-art (SOTA) accuracy over time fo…
Figure 4.17
Figure 4.17. Figure 4.17: Accuracy as a function of model access. The relationship between access (open vs. limited vs. closed) and model accuracy for each of the 16 core scenarios. Shaded bars indicate the performance of the best model for that scenario, whereas the solid bars indicate the …
Figure 4
Figure 4. Figure 4: , we report model accuracies as a function of time (i.e. model release date). [PITH_FULL_IMAGE:figures/full_fig_p135_4.png]
Figure 4.18
Figure 4.18. Figure 4.18: Model size vs. accuracy. The relationship between model parameter size (x-axis) and the accuracy of the most accurate model released up to that scale on each core scenario. That is, the graph tracks the progress in the state-of-the-art (SOTA) accuracy as a function …
Figure 4.19
Figure 4.19. Figure 4.19: The Pile loss vs. accuracy. The relationship between log bits-per-byte (BPB) on The Pile and the accuracy on each core scenario [PITH_FULL_IMAGE:figures/full_fig_p137_4_19.png]
Figure 4.20
Figure 4.20. Figure 4.20: Variance across seeds. For a subset of models and scenarios, we evaluate each scenario with three different random sets of in-context examples. We compute the range of the accuracy metric (maximum minus minimum value over the three random seeds) and visualize across…
Figure 4.21
Figure 4.21. Figure 4.21: Number of in-context examples. For each model, we set the maximum number of in-context examples to [0, 1, 2, 4, 8, 16] and fit as many in-context examples as possible within the context window. We plot performance as a function of the average number of in-context ex…
Figure 4.22
Figure 4.22. Figure 4.22: Multiple-choice adaptation. For each adaptation method (joint, separate, and separate calibrated), we compare models across scenarios. an output_prefix of “Answer:” Formulation of multiple choice scenarios. Beyond the details of the prompt, we can conceptually imagi…
Figure 4.23
Figure 4.23. Figure 4.23: Metric spread for core scenarios. Metrics for every model on every core scenario as a means for indicating the spread on a per-metric basis. Since we organize the 16 core scenarios by the broader task, we highlight findings at the task level. To provide a sense of t…
Figure 4.24
Figure 4.24. Figure 4.24: Robustness–equivariance via contrast sets. For the two scenarios where we have access to hand-crafted contrast sets, for each model, we plot the robustness of the model on that scenario (worst-case performance across perturbations of each instance) as a function of …
Figure 3.6
Figure 3.6. Figure 3.6: indicators that are upstream of the model, indicators that relate to the model itself, and indicators that are downstream of the model. Upstream indicators. The upstream indicators identify the ingredients and pro￾cesses involved in building a foundation model. There…
Figure 5.1
Figure 5.1. Figure 5.1: Indicators. The 100 indicators of the Foundation Model Transparency Index spanning the 3 domains: upstream, model, and downstream. [PITH_FULL_IMAGE:figures/full_fig_p181_5_1.png]
Figure 5.2
Figure 5.2. Figure 5.2: Upstream indicators. The 32 upstream indicators that span Data, Data Labor, Data Access, Compute, Methods, and Data Mitigations [PITH_FULL_IMAGE:figures/full_fig_p182_5_2.png]
Figure 5.3
Figure 5.3. Figure 5.3: Model indicators. The 33 model indicators that span Model Basics, Model Access, Capabilities, Limitations, Risks, Model Mitigations, Trustworthiness, and Inference [PITH_FULL_IMAGE:figures/full_fig_p185_5_3.png]
Figure 5.4
Figure 5.4. Figure 5.4: Downstream indicators. The 35 downstream indicators that span Dis￾tribution, Usage Policy, Model Behavior Policy, User Interface, User Data Protection, Model Updates, Feedback, Impact, and Downstream Documentation [PITH_FULL_IMAGE:figures/full_fig_p188_5_4.png]
Figure 5.5
Figure 5.5. Figure 5.5: Overall 2023 FMTI scores. The overall 2023 Foundation Model Transparency Index score and ranking across all 100 indicators. 5.3 Results The finalized results of the 2023 Foundation Model Transparency Index are the scores for each of the 100 indicators across all 10 c…
Figure 5.6
Figure 5.6. Figure 5.6: Scores by domain. The aggregate 2023 FMTI score of each developer broken down by the three domains: upstream, model, and downstream. 40% 60% 20% 40% 20% 0% 20% 0% 0% 0% 29% 86% 14% 14% 0% 29% 0% 0% 0% 0% 57% 14% 14% 57% 14% 0% 14% 0% 0% 0% 75% 100% 50% 100% 75% 75% 0…
Figure 5.7
Figure 5.7. Figure 5.7: Scores by major dimensions of transparency. The fraction of achieved indicators in each of the 13 major dimension of transparency in the 2023 FMTI. Major dimension of transparency are large subdomains within the 23 subdomains [PITH_FULL_IMAGE:figures/full_fig_p197_5…
Figure 5.8
Figure 5.8. Figure 5.8: Upstream scores by indicator. The 2023 FMTI scores for each of the 32 upstream indicators [PITH_FULL_IMAGE:figures/full_fig_p204_5_8.png]
Figure 5.9
Figure 5.9. Figure 5.9: Model scores by indicator. The 2023 FMTI scores for each of the 33 model indicators [PITH_FULL_IMAGE:figures/full_fig_p205_5_9.png]
Figure 5.10
Figure 5.10. Figure 5.10: Downstream scores by indicator. The 2023 FMTI scores for each of the 35 downstream indicators. most lightweight of these indicators. In general, we highlight an important mismatch between the many risks that are enumerated and the relatively few mitigations that are…
Figure 5.11
Figure 5.11. Figure 5.11: Open developers score higher in aggregate and on every domain. We establish a clear trend that the open developers score higher overall, with all three being among the four highest-scoring developers (see [PITH_FULL_IMAGE:figures/full_fig_p215_5_11.png]
Figure 5.11
Figure 5.11. Figure 5.11: Open vs. closed by subdomain. The 2023 FMTI average score for the 3 open developers (Meta, Hugging Face, Stability AI) and the 7 closed developers (OpenAI, Anthropic, Google, Cohere, AI21 Labs, Inflection, Amazon) across each of the 23 subdomains. Note: the number o…
Figure 5.12
Figure 5.12. Figure 5.12: Correlations between Companies. The correlation between the 2023 FMTI scores for pairs of companies across all indicators. Correlation is measured using the simple matching coefficient (i.e. agreement rate), which is the fraction of all indicators for which both com…
Figure 5.13
Figure 5.13. Figure 5.13: Correlations between companies (upstream). The correlation between the 2023 FMTI scores for pairs of companies across all indicators when only considering upstream indicators. Correlation is measured using the simple matching coefficient (i.e. agreement rate), which…
Figure 5.14
Figure 5.14. Figure 5.14: Correlations between companies (model). The correlation between the 2023 FMTI scores for pairs of companies across all indicators when only considering model indicators. Correlation is measured using the simple matching coefficient (i.e. agreement rate), which is th…
Figure 5.15
Figure 5.15. Figure 5.15: Correlations between companies (downstream). The correlation between the 2023 FMTI scores for pairs of companies across all indicators when considering only downstream indicators. Correlation is measured using the simple matching coefficient (i.e. agreement rate), w…
Figure 5.16
Figure 5.16. Figure 5.16: Scores by domain. The overall 2024 FMTI scores disaggregated into the three domains: upstream, model, and downstream. These initial scores were sent to developers to permit rebuttal: in contrast to the 2023 FMTI, the rebuttal process was a more iterative multi-week …
Figure 5.17
Figure 5.17. Figure 5.17: Scores by major dimensions of transparency. The fraction of achieved indicators in each of the 13 major dimensions of transparency in the 2024 FMTI. Major dimensions of transparency are large subdomains within the 23 subdo￾mains. The upstream domain is the most opaq…
Figure 5.18
Figure 5.18. Figure 5.18: Overall scores by release strategy. The overall 2024 FMTI scores for the 6 open developers (Adept, BigCode/Hugging Face/ServiceNow, Meta, Microsoft, Mistral, Stability AI) and the 8 closed developers (AI21 Labs, Aleph Alpha, Amazon, Anthropic, Google, IBM, OpenAI, W…
Figure 5.19
Figure 5.19. Figure 5.19: Change in overall scores. The 2023 FMTI and 2024 overall scores for the eight developers assessed in both versions. Transparency increased across the board from 2023 to 2024. Foundation model developers significantly improved their scores between October 2023 and Ma…
Figure 5.20
Figure 5.20. Figure 5.20: Change in subdomain scores from 2023 FMTI to 2024 FMTI. This figure shows the percentage point change in scores for major subdomains for the eight developers that are included in both the October 2023 and May 2024 versions of the Foundation Model Transparency Index …
Figure 5.21
Figure 5.21. Figure 5.21: Scores by new information status. The overall 2024 FMTI scores disaggregated based on whether the information was newly disclosed. were included in 2023 FMTI had the benefit of having already evaluated their own transparency practices in relation to this initiative …
Figure 5
Figure 5. Figure 5: totals the amount of new information across all developers [PITH_FULL_IMAGE:figures/full_fig_p241_5.png]
Figure 5.22
Figure 5.22. Figure 5.22: Aggregate new information by major dimensions of transparency. The number of pieces of new information, aggregated across all developers, for each of the 13 major domains of transparency in the 2024 FMTI. Note: major dimensions of transparency each have a different …
Figure 6
Figure 6. Figure 6: shows the results of the compliance pre-assessment. We see four areas [PITH_FULL_IMAGE:figures/full_fig_p264_6.png]
Figure 6.1
Figure 6.1. Figure 6.1: Compliance pre-assessment scores. We assess 10 major foundation model providers (and their flagship models) for the 12 AI Act requirements on a scale from 0 (worst) to 4 (best). The best possible score is 48 as a result. We present the final scores in the above figur…

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

25 extracted references · 7 canonical work pages

  1. [7]

    The platform accountability and transparency act.Congressional Bill. A. Feder Cooper, David Mimno, Madiha Choksi, and Katherine Lee. 2023. Machine learning and artificial intelligence: Legal concepts.Generative AI and Law Workshop at the International Conference of Machine Learning. Eric Corbett and Emily Denton. 2023. Interrogating the t in facct. InProc...

  2. [11]

    arXiv preprint arXiv:2009.11462

    Realtoxicityprompts: Evaluating neural toxic degeneration in language models. arXiv preprint arXiv:2009.11462. Sebastian Gehrmann, Tosin Adewumi, Karmanya Aggarwal, Pawan Sasanka Am- manamanchi, Anuoluwapo Aremu, Antoine Bosselut, Khyathi Raghavi Chandu, Miruna-Adriana Clinciu, Dipanjan Das, Kaustubh Dhole, Wanyu Du, Esin Durmus, Ondřej Dušek, Chris Chine...

  3. [12]

    Lauren Kogen

    Racial disparities in automated speech recognition.Proceedings of the National Academy of Sciences of the United States of America, 117:7684 – 7689. Lauren Kogen. 2022. From statistics to stories: Indices and indicators as com- munication tools for social change.The International Journal of Press/Politics, 0(0):19401612221094246. Pang Wei Koh, Shiori Saga...

  4. [14]

    O’Reilly Media, Inc

    Quantifying the carbon emissions of machine learning. arXiv preprint arXiv:1910.09700. Faisal Ladhak, Esin Durmus, He He, Claire Cardie, and Kathleen McKeown. 2022. Faithful or extractive? on mitigating the faithfulness-abstractiveness trade-off in abstractive summarization. InProceedings of the 60th Annual Meeting of the Association for Computational Lin...

  5. [15]

    Christopher A

    Textattack: A framework for adversarial attacks, data augmentation, and adversarial training in nlp.arXiv preprint arXiv:2005.05909. Christopher A. Mouton, Caleb Lucas, and Ella Guest. 2024. The Operational Risks of AI in Large-Scale Biological Attacks: Results of a Red-Team Study. Technical report, RAND Corporation. Mozilla. 2023. Joint Statement on AI S...

  6. [17]

    open science

    Adversarial nli: A new benchmark for natural language understanding. In Association for Computational Linguistics (ACL). Helen Nissenbaum. 2024. Privacy as contextual integrity. InWashington Law Review. Safiya Umoja Noble. 2018.Algorithms of Oppression. New York University Press. Rodrigo Nogueira and Kyunghyun Cho. 2019. Passage re-ranking with BERT.arXiv...

  7. [19]

    Scaling language models: Methods, analysis & insights from training gopher. arXiv. Colin Raffel. 2023. Building Machine Learning Models Like Open Source Software. Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2019. Exploring the limits of transfer learning with a unified text-...

  8. [20]

    everyone wants to do the model work, not the data work

    Understanding and mitigating the tradeoff between robustness and accuracy. In International Conference on Machine Learning (ICML). Deb Raji. 2022. Mozilla open source audit tooling (oat) project. Project documentation or description at Mozilla. Inioluwa Deborah Raji and Joy Buolamwini. 2019. Actionable auditing: Investigating the impact of publicly naming...

Show all 25 references
  1. [21]

    Raphael Satter

    Privacy leakage in discrete time updating systems. Raphael Satter. 2023. FBI says artificial intelligence being used for ’sextortion’ and harassment. Reuters. Joseph R. Saveri, Cadio Zirpoli, Christopher K.L. Young, and Kathleen J. McMahon

  2. [22]

    openai, inc., et al

    Paul tremblay, mona awad vs. openai, inc., et al. Case 3:23-cv-03223-AMO Document 1 Filed 06/28/23, UNITED STATES DISTRICT COURT, NORTHERN DISTRICT OF CALIFORNIA, SAN FRANCISCO DIVISION. Rylan Schaeffer, Brando Miranda, and Sanmi Koyejo. 2023. Are emergent abilities of large l...

  3. [23]

    David Thiel, Melissa Stroebel, and Rebecca Portnoff

    URL https://purl .... David Thiel, Melissa Stroebel, and Rebecca Portnoff. 2023. Generative ML and CSAM: Implications and Mitigations. Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam M. Shazeer, Apoorv Kul- shreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, ...

  4. [24]

    arXiv preprint arXiv:2112.04359

    Ethical and social risks of harm from language models. arXiv preprint arXiv:2112.04359. Laura Weidinger, Deb Raji, Hanna Wallach, Margaret Mitchell, Angelina Wang, Olawale Salaudeen, Rishi Bommasani, Sayash Kapoor, Deep Ganguli, Sanmi Koyejo, and William Isaac. 2025. Toward an...

  5. [25]

    Ethical AI

    Taxonomy of risks posed by language models. InProceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, pages 214–229. Barry R. Weingast and William J. Marshall. 1988. The industrial organization of congress; or, why legislatures, like firms, are no...

  6. [68]

    Max Bartolo, Alastair Roberts, Johannes Welbl, Sebastian Riedel, and Pontus Stene- torp

    PMID: 31313636. Max Bartolo, Alastair Roberts, Johannes Welbl, Sebastian Riedel, and Pontus Stene- torp. 2020. Beat the AI: Investigating adversarial human annotation for reading comprehension. Transactions of the Association for Computational Linguistics, 8:662–678. Adrien Ba...

  7. [2012]

    InProceedings of the 3rd Innovations in Theoret- ical Computer Science Conference, ITCS ’12, page 214–226, New York, NY, USA

    Fairness through awareness. InProceedings of the 3rd Innovations in Theoret- ical Computer Science Conference, ITCS ’12, page 214–226, New York, NY, USA. Association for Computing Machinery. Josh Dzieza. 2023. Ai is a lot of work: As the technology becomes ubiquitous, a vast t...

  8. [2015]

    InEmpirical Methods in Natural Language Processing (EMNLP)

    A large annotated corpus for learning natural language inference. InEmpirical Methods in Natural Language Processing (EMNLP). Samuel R. Bowman and George Dahl. 2021. What will it take to fix benchmarking in natural language understanding? InProceedings of the 2021 Conference o...

  9. [2016]

    InProceedings of the 10th International Workshop on Semantic Evaluation (SemEval-2016), pages 1–18, San Diego, California

    SemEval-2016 task 4: Sentiment analysis in Twitter. InProceedings of the 10th International Workshop on Semantic Evaluation (SemEval-2016), pages 1–18, San Diego, California. Association for Computational Linguistics. Shashi Narayan, Shay B. Cohen, and Mirella Lapata. 2018. Do...

  10. [2017]

    InAdvances in Neural Information Processing Systems (NeurIPS), pages 5684–5693

    On fairness and calibration. InAdvances in Neural Information Processing Systems (NeurIPS), pages 5684–5693. Christopher Potts, Zhengxuan Wu, Atticus Geiger, and Douwe Kiela. 2021. DynaSent: A dynamic benchmark for sentiment analysis. InProceedings of the 59th Annual Meeting o...

  11. [2018]

    ArXiv:1802.07228 [cs]

    The Malicious Use of Artificial Intelligence: Forecasting, Prevention, and Mitigation. ArXiv:1802.07228 [cs]. Miles Brundage, Shahar Avin, Jasmine Wang, Haydn Belfield, Gretchen Krueger, BIBLIOGRAPHY 268 Gillian Hadfield, Heidy Khlaaf, Jingying Yang, Helen Toner, Ruth Fong, Te...

  12. [2019]

    In World Wide Web (WWW), pages 491–500

    Nuanced metrics for measuring unintended bias with real data for text classification. In World Wide Web (WWW), pages 491–500. Samuel Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning

  13. [2020]

    Laurent Fleury

    Principled Artificial Intelligence: Mapping Consensus in Ethical and Rights- Based Approaches to Principles for AI.Berkman Klein Center Research Publication, (2020-1). Laurent Fleury. 2014.Sociology of culture and cultural practices: The transformative power of institutions. L...

  14. [2021]

    In 30th USENIX Security Symposium (USENIX Security 21), pages 2633–2650

    Extracting training data from large language models. In 30th USENIX Security Symposium (USENIX Security 21), pages 2633–2650. Nancy Cartwright and Jeremy Hardie. 2012.Evidence-Based Policy: A Practical Guide to Doing It Better. Oxford University Press. Stephen Casper, Xander D...

  15. [2022]

    Rishi Bommasani, Drew A

    Picking on the same person: Does algorithmic monoculture lead to outcome homogenization? In Advances in Neural Information Processing Systems. Rishi Bommasani, Drew A. Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S. Bernstein, Jeannette Bohg, Antoine...

  16. [2023]

    Technical report

    Supporting Open Source and Open Science in the EU AI Act. Technical report. Kimberlé Crenshaw. 1989. Demarginalizing the intersection of race and sex: A black feminist critique of antidiscrimination doctrine, feminist theory and antiracist politics. University of Chicago Legal...

  17. [2025]

    Noam Kolt, Markus Anderljung, Joslyn Barnhart, Asher Brass, Kevin Esvelt, Gillian K

    AI 2027. Noam Kolt, Markus Anderljung, Joslyn Barnhart, Asher Brass, Kevin Esvelt, Gillian K. Hadfield, Lennart Heim, Mikel Rodriguez, Jonas B. Sandbrink, and Thomas Wood- side. 2024. Responsible Reporting for Frontier AI Development. BIBLIOGRAPHY 292 Alex Krizhevsky, Ilya Sut...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.