Pith. sign in

REVIEW 10 cited by

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.00561 v2 pith:JQDGEE2V submitted 2025-02-01 cs.CY

classification cs.CY
keywords systemsevaluatinggenaimeasurementsocialpositionchallengeconceptual
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The measurement tasks involved in evaluating generative AI (GenAI) systems lack sufficient scientific rigor, leading to what has been described as "a tangle of sloppy tests [and] apples-to-oranges comparisons" (Roose, 2024). In this position paper, we argue that the ML community would benefit from learning from and drawing on the social sciences when developing and using measurement instruments for evaluating GenAI systems. Specifically, our position is that evaluating GenAI systems is a social science measurement challenge. We present a four-level framework, grounded in measurement theory from the social sciences, for measuring concepts related to the capabilities, behaviors, and impacts of GenAI systems. This framework has two important implications: First, it can broaden the expertise involved in evaluating GenAI systems by enabling stakeholders with different perspectives to participate in conceptual debates. Second, it brings rigor to both conceptual and operational debates by offering a set of lenses for interrogating validity.

Discussion (0). Sign in to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations)

    cs.HC 2026-08 conditional novelty 6.0 of 10

    Chat UI and API access to the same chatbot produce different accuracy, consistency, citation, and refusal behaviors on safety benchmarks, and web search changes these patterns further.

  2. On the Convergent Validity of Offline Evaluation Designs for Recommender Systems

    cs.IR 2026-07 conditional novelty 6.0 of 10

    Sparse offline recommender rankings correlate only weakly—and sometimes negatively—with dense ground-truth rankings, and no evaluation design is uniformly best.

  3. Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Small hyperbolic models (146M–3B) report 100% creative-seed preference, 90.7% compliance-gap detection, and a selective-gating skeleton–wallpaper memory pilot as a companion-AI stack.

  4. Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering

    cs.SE 2026-06 unverdicted novelty 6.0 of 10

    Coding benchmarks misalign with agentic software engineering because they conflate model and harness, grade against single references, and provide no component-level iteration signals.

  5. Grounded Chess Reasoning in Language Models via Master Distillation

    cs.AI 2026-03 unverdicted novelty 6.0 of 10

    Master Distillation turns opaque chess-engine search into natural-language CoT, lifting a 4B model (C1) from near-zero to 48.1% accuracy that beats its teacher and most larger LMs.

  6. Toward Valid Measurement Of (Un)fairness For Generative AI: A Proposal For Systematization Through The Lens Of Fair Equality of Chances

    cs.CY 2025-07 accept novelty 6.0 of 10

    A Fair Equality of Chances-based framework decomposes GenAI unfairness into harms/benefits, morally arbitrary factors, and morally decisive factors to improve measurement validity.

  7. Understanding and Meeting Practitioner Needs When Measuring Representational Harms Caused by LLM-Based Systems

    cs.CY 2025-06 conditional novelty 6.0 of 10

    Practitioners trying to measure representational harms in LLM-based systems often cannot use public measurement instruments, either because the instruments lack validity, specificity, interpretability, or actionabilit...

  8. Long-term Measurements: Towards a Longitudinal Understanding of Human-AI Interactions

    cs.AI 2026-08 accept novelty 5.0 of 10

    The paper sets out a research agenda for NLP to measure how prolonged language-model use changes human behavior over long time horizons, replacing single-session safety evaluations with longitudinal tracking.

  9. Against 'softmaxing' culture

    cs.HC 2025-06 unverdicted novelty 5.0 of 10

    A position paper arguing that AI evaluations should shift from defining culture to understanding when culture becomes relationally valid.

  10. Reality Check: A New Evaluation Ecosystem Is Necessary to Understand AI's Real World Effects

    cs.CY 2025-05 conditional novelty 4.0 of 10

    A position paper argues that understanding AI's second-order effects requires moving from static benchmarks to an ecosystem of field testing, red teaming, and contextual evaluation.

Pith tools