Pith. sign in

REVIEW 5 cited by

Rethinking Model Evaluation as Narrowing the Socio-Technical Gap

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.03100 v4 pith:OMZ77NWP submitted 2023-06-01 cs.HC cs.AI

classification cs.HCcs.AI
keywords evaluationmodelmethodssocio-technicalchallengescommunitydiversehomogenization
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The recent development of generative large language models (LLMs) poses new challenges for model evaluation that the research community and industry have been grappling with. While the versatile capabilities of these models ignite much excitement, they also inevitably make a leap toward homogenization: powering a wide range of applications with a single, often referred to as ``general-purpose'', model. In this position paper, we argue that model evaluation practices must take on a critical task to cope with the challenges and responsibilities brought by this homogenization: providing valid assessments for whether and how much human needs in diverse downstream use cases can be satisfied by the given model (\textit{socio-technical gap}). By drawing on lessons about improving research realism from the social sciences, human-computer interaction (HCI), and the interdisciplinary field of explainable AI (XAI), we urge the community to develop evaluation methods based on real-world contexts and human requirements, and embrace diverse evaluation methods with an acknowledgment of trade-offs between realisms and pragmatic costs to conduct the evaluation. By mapping HCI and current NLG evaluation methods, we identify opportunities for evaluation methods for LLMs to narrow the socio-technical gap and pose open questions.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. From Queries to Criteria: Understanding How Astronomers Evaluate LLMs

    cs.CL 2025-07 conditional novelty 7.0 of 10

    A user study of an astronomy RAG bot identifies the question types and evaluation criteria astronomers actually use, and turns them into a 40-item benchmark.

  2. Position: Evaluation Scores Are Perishable Knowledge Claims

    cs.AI 2026-07 conditional novelty 6.0 of 10

    Evaluation scores should be treated as perishable knowledge claims with explicit formality, scope, and validity windows; weakest-link (min) aggregation is the proposed conservative endpoint for combining them.

  3. The Foreign Policy AI Evaluation Gap

    cs.CY 2026-07 conditional novelty 6.0 of 10

    Public technical AI governance almost never evaluates real foreign-policy AI workflows; the paper maps that gap and proposes task-scoped, human-recombined evaluation instead of model leaderboards.

  4. A Conceptual Framework for AI Capability Evaluations

    cs.AI 2025-06 conditional novelty 5.0 of 10

    A descriptive conceptual framework with seven elements (target, task, subject, inputs, instance, measurement, result analysis) for systematizing analysis of AI capability evaluations.

  5. On the Dual-Use Dilemma in Physical Reasoning and Force

    cs.RO 2025-05 conditional novelty 5.0 of 10

    Adding Asimov-style safety prompts to vision-language models lowers both harmful and helpful force generation for contact-rich robotic tasks.

Pith tools