Pith. sign in

REVIEW 3 major objections 5 minor 1 cited by

Human-Centric Evaluation for Foundation Models

T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read A human-centric evaluation framework, built from nine subjective ratings gathered during open-ended human-AI research collaborations, claims that Grok 3 currently gives the best assistance among four major foundation models.

desk verdict Real dataset and rubric, but the abstract's ranking contradicts the paper's own tables and the stats are too thin to support the leaderboard. read the letter →

arxiv 2506.01793 v1 pith:76GERYO5 submitted 2025-06-02 cs.CL

classification cs.CL
keywords human-centricevaluationfoundationmodelssubjectivebenchmarkhuman-AIcollaborationproblem-solvingabilityinformationqualityinteractionexperienceopen-endedresearchtasks
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Nearly all foundation-model benchmarks grade models on quiz-style objective accuracy, which the authors argue misses what it feels like to use a model. This paper proposes replacing that with a Human-Centric subjective Evaluation (HCE) framework, in which people work with a model on an open-ended research task for twenty minutes and then rate it on nine sub-dimensions grouped into problem-solving ability, information quality, and interaction experience. Using over 540 such sessions across eight disciplines and two languages, the paper finds that Grok 3 receives the highest human ratings, followed by DeepSeek R1 and Gemini 2.5, with OpenAI o3 mini rating lowest. The contribution is a reusable framework and dataset for capturing authentic human experience rather than a static quiz score.

What carries the argument

The carrying object is the HCE framework's nine-item rating instrument, grouped into three dimensions: problem-solving ability (analytical accuracy, comprehensiveness, assistance efficiency), information quality (reliability, exploration depth), and interaction experience (content relevance, feedback adaptability, expression naturalness, response timeliness). Raters choose a task from eight disciplines, work with a model for a limited time on literature-synthesis or innovation-driven problems, and score each item on a five-point scale; the mean becomes the model's score. This turns evaluation from a right-or-wrong exercise into a report on whether the model helped a person think.

What would settle it

Run a blinded replication where the same tasks are completed with model identities hidden and with at least five raters per model per language, then check whether the 0.2-0.4 point gaps between Grok 3, DeepSeek R1, Gemini 2.5, and o3 mini survive; if the intervals overlap, the ranking is rater noise rather than model differences. A cheaper check is to recompute the reported means with per-rater variances and confidence intervals from the released dataset.

Watch

Extended reading notes

Core claim

The central claim is that subjective human ratings, gathered through free-form collaboration on realistic research tasks, give a different and more faithful account of foundation-model usefulness than objective benchmarks do. On the paper's own results, Grok 3 is the best research collaborator of the four tested models, with a total score of 4.30 out of 5, ahead of Gemini 2.5 (4.09) and DeepSeek R1 (4.08), while OpenAI o3 mini trails at 3.91. The paper also reports that each model has a distinct profile: Gemini 2.5 is strongest in information reliability and content relevance, DeepSeek R1 is stable but slower, o3 mini is the most timely but shallowest, and Grok 3 leads in exploration depth and feedback adaptability.

Load-bearing premise

The load-bearing premise is that the average of ratings from a small, self-selected group of student raters who know which model they are using is a stable and comparable measure of real-world helpfulness across models.

Editorial extensions

If this is right

  • A model that wins a human-centric leaderboard can be recommended for open-ended research support even if quiz benchmarks put it lower.
  • Dimension-level scores locate the weakest behaviour per model: DeepSeek R1's timeliness, o3 mini's exploration depth, and Gemini 2.5's education-domain performance.
  • The framework turns evaluation into a task-generic process, so new models can be inserted into the same 20-minute collaboration-and-rating protocol without new question sets.
  • Because the collected 540-session dataset pairs open-ended interactions with fine-grained ratings, it can serve as training ground for automated subjective scoring.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Not tested in this paper: whether the ranking survives blinding. Because raters know which model they are using and the gaps are close to rater-scale noise, brand expectations could account for part of Grok 3's lead.
  • An analysis the paper does not report: per-rater variance. If ratings from the five evaluators per model vary widely, the reported mean differences may not be distinguishable from noise.
  • A testable extension: an automated scorer trained on this dataset would inherit the raters' preferences, including the paper's observation that encouraging language and rhetorical questions improved experience, so automation should be validated against blinded human ratings.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes a Human-Centric subjective Evaluation (HCE) framework for foundation models, built on three dimensions: problem-solving ability, information quality, and interaction experience. The authors conduct a subjective evaluation of four models (Deepseek R1, OpenAI o3 mini, Grok 3, and Gemini 2.5) through over 540 participant-driven collaborations on open-ended research tasks across eight disciplines and two languages. The headline claim is that Grok 3 performs best, followed by DeepSeek R1 and Gemini 2.5, with OpenAI o3 mini lagging. The authors also release a dataset and argue that their framework addresses gaps in objective benchmarks by capturing authentic human experience.

Significance. If the results were properly supported, the HCE framework would be a timely contribution to the growing call for human-centered evaluation of LLMs, and the released dataset could be a useful resource for future work on subjective benchmarking. The paper correctly identifies that objective benchmarks such as MMLU and SWE-bench do not capture user experience, and the proposed three-dimensional rubric is a plausible starting point. The open dataset is a concrete asset. However, the central ranking claim is currently undermined by an internal inconsistency and a lack of statistical grounding, so the contribution is not yet ready for publication in its present form.

major comments (3)
  1. [Abstract and Table 2] The abstract states that Grok 3 is 'followed by Deepseek R1 and Gemini 2.5', which implies that DeepSeek R1 is ranked second and Gemini 2.5 third. However, Table 2a and Table 2b report total scores of 4.09/4.09 for Gemini 2.5 and 4.08/4.08 for DeepSeek R1, i.e., Gemini 2.5 is second and DeepSeek R1 third by a 0.01 margin. Section 4.2's phrase 'second tier' blurs this distinction instead of resolving it. This is an internal contradiction in the paper's headline claim, and the authors must either correct the abstract and Section 4.2 to match their own data or provide a principled justification for reordering the rankings.
  2. [Section 4.1 and Table 2] The ranking in Table 2 is based on mean scores only; no variance, confidence intervals, inter-rater reliability, or significance tests are reported. Section 4.1 states that Gemini 2.5 had only three raters per language while the other models had five, and the raters were not blinded to model identity. The difference between second and third place in both leaderboards is 0.01 on a 5-point scale, which is far smaller than the rater noise one would expect from such a small, self-selected, unblinded sample. The paper's central claim about model ordering is therefore not statistically supported; the authors need to provide error bars or other measures of uncertainty, or soften the claimed ranking to match the evidence.
  3. [Sections 3.3, 3.4, and 4.2] The paper does not include the exact evaluation instrument: the questionnaire items corresponding to the nine sub-dimensions, the specific task prompts for each discipline and problem type, the rating instructions, or a breakdown of the number of raters per cell. Section 3.4 only describes a generic five-point scale, and Section 4.2 reports aggregated means. Since the dataset is a core contribution and every score depends on the subjective questionnaire, the omission prevents replication and independent verification of the results. The full protocol and instruments should be provided in an appendix or supplementary material.
minor comments (5)
  1. [Section 4.2] The text says results are 'presented in Figure 2 and Table 4', but no Table 4 exists in the manuscript; the relevant results are in Table 2.
  2. [Section 1] The introduction refers to 'GPT o3 mini' while the rest of the paper uses 'OpenAI o3 mini'; please standardize the model name.
  3. [Section 4.1] The paper says 'over 540 participant-driven evaluations' and later 'Based on 540 individual evaluations', but the exact number and the per-condition sample sizes are not stated; please clarify the total count and how it relates to the 3-5 raters per model per language.
  4. [Table 1] The comparison table would benefit from a reference for the 'Our method' row, and the meaning of 'App. Cov.' and the check marks could be stated more explicitly in the caption.
  5. [Section 4.2] The observation about Grok 3's rhetorical questions is anecdotal and unquantified; if it is intended as a finding, it should be supported by data or moved to the limitations discussion.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the leaderboard is direct aggregation of human ratings, not a derived or fitted prediction, and the framework dimensions are adopted from external prior work rather than self-citation.

full rationale

The paper's central claim is an empirical ranking produced by averaging human ratings. Section 3.4 states that 'final scores are calculated as the arithmetic mean of all ratings to ensure objectivity,' so the leaderboard in Table 2 is a direct measurement protocol, not a prediction derived from the HCE framework. No parameter is fitted to reproduce the ranking, and no equation makes the framework's dimensions numerically equivalent to the reported scores. The three core dimensions are asserted from prior literature in Section 3.2 ('We synthesize their findings and innovatively introduce a novel framework'), and the cited works on information quality, interaction experience, and problem-solving ability are external references rather than load-bearing self-citations by this paper's authors. Calling the evaluation 'human-centric' because it uses human ratings is definitional, but not circular: the goal is to measure subjective experience, and the ratings are the data, not a re-description of the framework. The paper's internal inconsistency between the abstract ordering of DeepSeek R1 and Gemini 2.5 and the 4.09-vs-4.08 totals in Tables 2a and 2b, together with unequal rater counts and the absence of significance tests in Section 4.1, are reporting and statistical-validity concerns; they do not constitute circularity by construction. No specific circular step could be exhibited from the text, so under the requirement to avoid speculation, the appropriate finding is no circularity.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

No free parameters are fitted; the framework rests on the domain assumptions that Likert-scale ratings are meaningful as interval data and that the proposed dimensions cover the relevant aspects of user experience. These assumptions are plausible but unvalidated, and they are load-bearing for the framework's conclusions.

assumptions (2)
  • domain assumption Subjective ratings on a 1-5 Likert scale can be treated as interval data and averaged across raters to compare models.
    The entire leaderboard in Table 2 is computed as arithmetic means of ratings from 3-5 raters per model per language.
  • domain assumption The three evaluation dimensions (problem-solving ability, information quality, interaction experience) are comprehensive and sufficiently non-overlapping to characterize human experience of model assistance.
    The framework in Section 3.2 defines these dimensions; no factor analysis or validation is provided.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Human-Centric Evaluation for Foundation Models." pith.science (2026). https://pith.science/paper/76GERYO5

@misc{pith2026250601793,
  author       = {Pith},
  title        = {Pith review of: Human-Centric Evaluation for Foundation Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/76GERYO5}},
  note         = {Machine review of arXiv:2506.01793}
}
read the original abstract

Currently, nearly all evaluations of foundation models focus on objective metrics, emphasizing quiz performance to define model capabilities. While this model-centric approach enables rapid performance assessment, it fails to reflect authentic human experiences. To address this gap, we propose a Human-Centric subjective Evaluation (HCE) framework, focusing on three core dimensions: problem-solving ability, information quality, and interaction experience. Through experiments involving Deepseek R1, OpenAI o3 mini, Grok 3, and Gemini 2.5, we conduct over 540 participant-driven evaluations, where humans and models collaborate on open-ended research tasks, yielding a comprehensive subjective dataset. This dataset captures diverse user feedback across multiple disciplines, revealing distinct model strengths and adaptability. Our findings highlight Grok 3's superior performance, followed by Deepseek R1 and Gemini 2.5, with OpenAI o3 mini lagging behind. By offering a novel framework and a rich dataset, this study not only enhances subjective evaluation methodologies but also lays the foundation for standardized, automated assessments, advancing LLM development for research and practical scenarios. Our dataset link is https://github.com/yijinguo/Human-Centric-Evaluation.

Figures

Figures reproduced from arXiv: 2506.01793 by the authors.

Figure 1
Figure 1. Motivation of our HCE. The traditional model-centric evaluation focuses on the quiz performance of foundation [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Framework of our HCE experiments. Participants choose a task based on their major and interests, then interact freely [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. Failure cases. Corresponding to each dimension included in the HCE framework, we present an example of the [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Brief performance comparison for problem types [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Affordance Benchmark for MLLMs

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A new 2,000-question benchmark finds multimodal AI models recognize object affordances far worse than humans, with top model Gemini-2.0-Pro at 18.05% versus 85.34% human best.

Reference graph

Works this paper leans on

26 extracted references · 6 canonical work pages · cited by 1 Pith paper

  1. [1]

    Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, Wei Xue, Shanghang Zhang, Jie Fu, and Zhiyuan Liu. 2023. Chateval: Towards better llm-based evalu- ators through multi-agent debate. arXiv preprint arXiv:2308.07201 (2023)

  2. [2]

    Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, et al . 2024. A survey on evaluation of large language models. ACM transactions on intelligent systems and technology 15, 3 (2024), 1–45

  3. [3]

    Zhiyu Chen, Wenhu Chen, Charese Smiley, Sameena Shah, Iana Borova, Dylan Langdon, Reema Moussa, Matt Beane, Ting-Hao Huang, Bryan Routledge, et al

  4. [4]

    Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Banghua Zhu, Hao Zhang, Michael Jordan, Joseph E Gonzalez, et al. 2024. Chatbot arena: An open platform for evaluating llms by human preference. In Forty-first International Conference on Machine Learning

  5. [5]

    Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. 2021. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168 (2021)

  6. [6]

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers). 4171–4186

  7. [7]

    Yann Dubois, Balázs Galambosi, Percy Liang, and Tatsunori B Hashimoto. 2024. Length-controlled alpacaeval: A simple way to debias automatic evaluators.arXiv preprint arXiv:2404.04475 (2024)

  8. [8]

    Google. 2025. Gemini 2.5: Our most intelligent AI model . https: //blog.google/technology/google-deepmind/gemini-model-thinking-updates- march-2025/#gemini-2-5-thinking

Show all 26 references
  1. [9]

    Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al . 2025. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948 (2025)

  2. [10]

    Zishan Guo, Renren Jin, Chuang Liu, Yufei Huang, Dan Shi, Linhao Yu, Yan Liu, Jiaxuan Li, Bojian Xiong, Deyi Xiong, et al . 2023. Evaluating large language models: A comprehensive survey. arXiv preprint arXiv:2310.19736 (2023)

  3. [11]

    Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2020. Measuring massive multitask language under- standing. arXiv preprint arXiv:2009.03300 (2020)

  4. [12]

    Yuzhen Huang, Yuzhuo Bai, Zhihao Zhu, Junlei Zhang, Jinghan Zhang, Tangjun Su, Junteng Liu, Chuancheng Lv, Yikai Zhang, Yao Fu, et al. 2023. C-eval: A multi- level multi-discipline chinese evaluation suite for foundation models. Advances in Neural Information Processing System...

  5. [13]

    Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. 2023. Swe-bench: Can language models resolve real-world github issues? arXiv preprint arXiv:2310.06770 (2023)

  6. [14]

    Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michi- hiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al

  7. [15]

    Aidar Myrzakhan, Sondos Mahmoud Bsharat, and Zhiqiang Shen. 2024. Open- llm-leaderboard: From multi-choice to open-style questions for llms evaluation, benchmark, and arena. arXiv preprint arXiv:2406.07545 (2024)

  8. [16]

    OpenAI. 2025. Introducing OpenAI o3 and o4-mini . https://openai.com/index/ introducing-o3-and-o4-mini/

  9. [17]

    Long Phan, Alice Gatti, Ziwen Han, Nathaniel Li, Josephina Hu, Hugh Zhang, Chen Bo Calvin Zhang, Mohamed Shaaban, John Ling, Sean Shi, et al . 2025. Humanity’s Last Exam. arXiv preprint arXiv:2501.14249 (2025)

  10. [18]

    David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R Bowman. 2024. Gpqa: A graduate-level google-proof q&a benchmark. In First Conference on Language Modeling

  11. [19]

    Ashish Sharma, Inna W Lin, Adam S Miner, David C Atkins, and Tim Althoff

  12. [20]

    Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al. 2022. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. arXiv prep...

  13. [21]

    Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B Hashimoto. 2023. Alpaca: A strong, replicable instruction-following model.Stanford Center for Research on Foundation Models. https://crfm. stanford. edu/2023/03/...

  14. [22]

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971 (2023)

  15. [23]

    xAI. 2025. Grok 3 Beta — The Age of Reasoning Agents . https://x.ai/news/grok-3

  16. [2021]

    arXiv preprint arXiv:2109.00122 (2021)

    Finqa: A dataset of numerical reasoning over financial data. arXiv preprint arXiv:2109.00122 (2021)

  17. [2022]

    arXiv preprint arXiv:2211.09110 (2022)

    Holistic evaluation of language models. arXiv preprint arXiv:2211.09110 (2022)

  18. [2023]

    Nature Machine Intelligence 5, 1 (2023), 46–57

    Human–AI collaboration enables more empathic conversations in text- based peer-to-peer mental health support. Nature Machine Intelligence 5, 1 (2023), 46–57

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.