Pith. sign in

REVIEW 6 cited by

Evaluating Human-Language Model Interaction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2212.09746 v5 pith:HSVQEW3W submitted 2022-12-19 cs.CL

classification cs.CL
keywords interactionevaluationhuman-lmnon-interactiveinteractivebetterhaliemetrics
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Many real-world applications of language models (LMs), such as writing assistance and code autocomplete, involve human-LM interaction. However, most benchmarks are non-interactive in that a model produces output without human involvement. To evaluate human-LM interaction, we develop a new framework, Human-AI Language-based Interaction Evaluation (HALIE), that defines the components of interactive systems and dimensions to consider when designing evaluation metrics. Compared to standard, non-interactive evaluation, HALIE captures (i) the interactive process, not only the final output; (ii) the first-person subjective experience, not just a third-party assessment; and (iii) notions of preference beyond quality (e.g., enjoyment and ownership). We then design five tasks to cover different forms of interaction: social dialogue, question answering, crossword puzzles, summarization, and metaphor generation. With four state-of-the-art LMs (three variants of OpenAI's GPT-3 and AI21 Labs' Jurassic-1), we find that better non-interactive performance does not always translate to better human-LM interaction. In particular, we highlight three cases where the results from non-interactive and interactive metrics diverge and underscore the importance of human-LM interaction for LM evaluation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 45 citations worldwide. Full citation record

  1. Beyond the Golden Record: Toward a Design Theory for Trustworthy Master Data Management with Self-Sovereign Identity

    cs.SE 2026-04 unverdicted novelty 6.0 of 10

    A design theory is derived for trustworthy master data management based on self-sovereign identity to support reliable, sovereign, and accountable data sharing in data ecosystems.

  2. Observable Social Life Spaces: Exploring User Interpretations of agent-side life context in human-agent interaction

    cs.HC 2026-03 conditional novelty 6.0 of 10

    Seeing an AI agent's autonomous virtual life increased users' perceived equality with it in a small study, but the effect needs replication.

  3. PersonaFuse: A Personality Activation-Driven Framework for Enhancing Human-LLM Interactions

    cs.CL 2025-09 reject novelty 5.0 of 10

    A post-training framework with persona-specific LoRA experts and a situation-aware router improves LLM emotional responses, but the evidence on preserving general ability is undercut by missing base-model comparisons.

  4. Interaction as Intelligence: Deep Research With Human-AI Partnership

    cs.CL 2025-07 reject novelty 5.0 of 10

    A human-in-the-loop deep research system with transparent, interruptible interaction is claimed to outperform commercial baselines, but the evidence is weakened by small samples and biased instructions.

  5. Automated Novelty Evaluation of Academic Paper: A Collaborative Approach Integrating Human and Large Language Model Knowledge

    cs.CL 2025-07 reject novelty 5.0 of 10

    Method novelty prediction from peer-review novelty sentences and ChatGPT method summaries improves accuracy on ICLR 2022 data, but the benchmark leaks reviewer opinions into the input.

  6. The Science of Evaluating Foundation Models

    cs.CL 2025-02 conditional novelty 3.0 of 10

    A survey-and-checklist proposal that organizes LLM evaluation into an ABCD framework (Algorithm, Big Data, Computation, Domain Expertise) for context-aware, documented assessment.

Pith tools