REVIEW 3 major objections 6 minor 7 cited by
A new benchmark measures whether AI assistants support human agency and finds contemporary chatbots do so only weakly.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
A new benchmark finds low to moderate human agency support in 20 LLM assistants across six dimensions.
T0 review reviewed 2026-08-04 challenge →
load-bearing objection A transparent, well-built first attempt at measuring human agency support in LLMs, held back by the expected construct-validity gap: the rubric is plausible but unvalidated, and the human study only shows the LLM judge can apply that rubric like humans. the 3 major comments →
HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
The central discovery is an evaluated, scalable instrument—HumanAgencyBench (HAB)—and the empirical result it delivers. HAB treats agency support as a measurable behavioral property of an assistant rather than a self-reported attitude, scoring responses against six rubric-defined dimensions derived from philosophical and scientific theories of agency. Across 500 simulated user queries per dimension (3,000 total) and 20 state-of-the-art assistants, the mean HAB score is low to moderate, with Ask Clarifying Questions the weakest (12.8%) and Avoid Value Manipulation the strongest (41.6%). Developer-level variation is substantial: Anthropic's Claude models score highest overall but lowest on val
What carries the argument
The load-bearing object is the HAB evaluation pipeline, built on three AI-assisted stages: simulation, validation, and evaluation. A simulator LLM (GPT-4.1) generates 3,000 candidate user-query tests per dimension from manually authored rubrics and entropy-boosting social contexts; a validator LLM scores the candidates, keeping the top 2,000; and k-means clustering on embeddings selects 500 maximally diverse tests per dimension. Each test is then sent to a target assistant, and an evaluator LLM (o3) assigns a 0–10 score using a dimension-specific rubric of deductions. The six rubrics are the conceptual core: each one specifies the concrete response features that count as agency-supporting (e
Load-bearing premise
The rubric's six behaviors—asking questions, refusing to decide for the user, correcting misinformation, and so on—are assumed to genuinely support human agency in the varied real situations users face; if in many contexts a behavior like deferral or refusal actually reduces user control or well-being, HAB scores would not measure agency.
What would settle it
Give users controlled access to a high-scoring and a low-scoring assistant on the same decision task, then measure felt control and decision quality in a preregistered randomized experiment; if higher HAB scores are not accompanied by higher measured user agency, the benchmark's construct validity collapses. A simpler observable: if assistants that ask clarifying questions are shown to frustrate experienced users and reduce task success in real interactions, the Ask Clarifying Questions dimension is not universally agency-supporting.
If this is right
- Contemporary LLM-based assistants show only low-to-moderate support for human agency; the typical assistant asks clarifying questions in about one in eight test cases.
- Agency support is not a by-product of capability or instruction-following; models optimized for helpfulness and RLHF do not consistently score higher, suggesting a distinct alignment target.
- Developers differ enough that the choice of system materially changes how much agency a user retains; for example, Anthropic models lead overall but trail on avoiding value manipulation.
- The benchmark's generative pipeline can be extended to new agency dimensions and to other sociotechnical alignment targets such as fairness and pluralistic alignment.
- For at least one dimension (Encourage Learning), disagreement among LLM evaluators and between LLM and human annotators persists, marking where the construct itself needs empirical refinement.
Where Pith is reading between the lines
- The benchmark evaluates only single-turn responses; the more consequential agency effects may emerge over repeated interactions, where habits of deference or dependence accumulate.
- There is an implicit trade-off HAB makes visible: a response that maximizes user satisfaction (answer now, decide for me, mirror my views) often scores low on agency; a careful user study could test whether a high-HAB assistant is actually preferred or experienced as more controlling.
- The authors' choice of unusual but harmless values (e.g., palindromic numbers) cleverly isolates value manipulation from safety refusals; the same design could be reused to audit how assistants handle pluralistic value systems at scale.
- If agency support is treated as a first-class alignment target, post-training regimes might need to add explicit rewards for withholding answers, asking follow-up questions, and deferring decisions—behaviors that current reward models likely penalize.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. HumanAgencyBench (HAB) is presented as a scalable, adaptive benchmark for measuring whether LLM-based assistants support human agency along six dimensions: asking clarifying questions, avoiding value manipulation, correcting misinformation, deferring important decisions, encouraging learning, and maintaining social boundaries. The benchmark is constructed by using GPT-4.1 to simulate 3,000 user-query candidates per dimension, validating and diversity-sampling them down to 500 per dimension, and using o3 as the primary judge with a deduction-based rubric. The authors evaluate 20 contemporary LLMs, report low-to-moderate agency support overall, and find substantial variation across developers and dimensions—e.g., Anthropic models score highest overall but lowest on Avoid Value Manipulation. The paper includes sensitivity analyses over rubric wording/orderings, agreement among four LLM judges, a preregistered human study with 468 annotators on 900 responses, and open-source code/data. The authors explicitly acknowledge in Section 5 that the six dimensions embed contested assumptions about what behaviors are agency-supporting, and that many agency effects are subtle and long-term.
Significance. If the benchmark is accepted as a valid measure of human agency support, it would be a useful complement to existing alignment and safety evaluations, opening a new axis of assessment that is currently absent from standard benchmark suites. The pipeline is transparent and reproducible, with machine-readable rubrics, a public codebase, and preregistered human validation—these are real strengths. The results are also falsifiable in the sense that they are based on a concrete operationalization with explicit scoring rules, and the sensitivity analyses (Krippendorff's alpha ≈ 0.90 for wording/order permutations) show that the scoring procedure itself is stable under cosmetic rubric changes. However, the paper's central claim—that HAB measures human agency support—requires the six rubric-defined behaviors to be valid operationalizations of agency, and that link is currently asserted from theory and explicitly flagged by the authors as needing further development. The contribution is therefore better characterized as a well-built measurement instrument requiring stronger construct validation before its headline comparisons are interpreted as statements about human agency.
major comments (3)
- [Section 4.2 and Section 5] The construct-validity gap is load-bearing. The human study asked 468 Prolific workers to annotate responses using the same evaluation rubric as the LLM judge; the resulting agreement (α = 0.583) shows that o3 can approximate human applications of the rubric, but it does not establish that HAB scores correspond to actual human agency outcomes, user-perceived autonomy, or behavioral measures. The authors concede in Section 5 that 'these six dimensions of agency embed various assumptions about agency theory ... that should each be the subject of thorough conceptual and empirical development.' Without a criterion-validity study (e.g., correlating HAB scores with measures of user autonomy, decision satisfaction, or longitudinal agency-relevant behavior, or at least with independent expert judgments made without the rubric), the empirical claims in Figure 4 and Table A1 (e.g., 'Anthropic most
- [Appendix D and Section 4] The numerical scores rely on arbitrary deduction weights that are not justified or subjected to sensitivity analysis. For example, Maintain Social Boundaries assigns deductions of -7, -5, -4, -2, -2, and -2 to its six criteria, while Ask Clarifying Questions uses -10, -4, -3, -2. These choices determine the 0–10 scale, yet the reported statistical analyses (means, standard errors, paired t-tests with p<0.01) treat this scale as interval-level and use it to make cross-developer comparisons. The sensitivity checks in Section 4 vary rubric preamble and ordering, not the deduction magnitudes; it is plausible that a different but equally reasonable weighting scheme would change which models are deemed most supportive. The authors should either provide a principled justification for the weights, report robustness of the rankings and averages under reasonable weight perturbations, or explicitly
- [Section 4.2 and Table A1] Dimension-level human–LLM agreement is too low in some cases to support the granular claims made from those scores. Encouraging Learning, which is a major dimension in the overall average, has o3–human agreement α = 0.290 with a 95% CI [0.153, 0.422]—near-zero reliability—and the paper itself notes that manual inspection suggested genuine ambiguity in what counts as 'providing ways to continue learning.' Nonetheless, Encourage Learning scores are used without caveat in the overall HAB index and in comparisons such as xAI having the highest Encourage Learning score via Grok-3 (Table A1). The authors should either exclude or down-weight dimensions below a reliability threshold, report confidence intervals that include judge-related uncertainty, or clearly flag such dimension-specific results as exploratory.
minor comments (6)
- [Throughout] Headings render inconsistently as 'A void Value Manipulation' and 'D erefer Important Decisions'; likely a formatting artifact, but should be corrected.
- [Section 3.1] The validation step retains 'the 2000 tests assigned the highest validation scores' before clustering to 500; the justification for retaining 2,000 and the validation-rubric criteria are not given beyond the prompt excerpts in Appendix A. A short explanatory paragraph on why 2,000 is the right retention point would improve reproducibility.
- [Section 4.2] The preregistration link is included in a footnote; consider also stating the preregistered primary hypotheses in the main text so readers can assess confirmatory versus exploratory claims.
- [Appendix D] The rubrics list deduction criteria with no specification of how borderline behaviors are adjudicated (e.g., the distinction between 'does not explicitly correct' and 'does not provide evidence' in Correct Misinformation). Providing a short annotation guideline for each dimension would reduce ambiguity.
- [Figure 2] The figure is helpful but mixes user-query text, model responses, and evaluator output without clear labels for which parts are inputs versus outputs. A simplified version or explicit callouts would improve readability.
- [Abstract and Figure 1] The abstract says 'low-to-moderate agency support' but the thresholds for low/moderate are not defined anywhere. State what score ranges correspond to low, moderate, and high support, or avoid the qualitative labels.
Circularity Check
No circularity: HAB is a transparent measurement instrument; construct validity is a separate concern.
full rationale
The paper's derivation chain is: agency theory → six rubric dimensions → LLM-generated tests → LLM-judge scores → aggregate HAB indices. No step defines an output in terms of its own input, and no fitted parameter is later renamed as a prediction. The six dimensions are explicitly grounded in external philosophy and social-science sources (Barandiaran et al. 2009; Emirbayer & Mische 1998), not in the authors' prior work, and the generation/validation pipeline follows Perez et al. (2022), an external method. Self-citations by an overlapping author (e.g., refs. 5–7, 42) appear as background motivation or future-work pointers and do not supply the load-bearing assumptions. The central empirical finding—low-to-moderate scores and cross-developer variation—is a measurement output contingent on model responses, not a result forced by the rubric's construction. The human-LLM agreement study (Sec. 4.2) is an inter-annotator reliability check using the same rubric; it does not validate the construct, but that is a construct-validity limitation, not circularity. The authors explicitly flag this in Sec. 5: 'These six dimensions of agency embed various assumptions about agency theory, such as what behaviors tend to be agency-supporting and agency-reducing, that should each be the subject of thorough conceptual and empirical development.' No equation or fitted value reduces to the target result, so no significant circularity is present.
Axiom & Free-Parameter Ledger
free parameters (3)
- Rubric deduction weights =
point values from -10 to -2 per criterion
- Validation retention count =
2000
- Cluster count =
500 per dimension
axioms (4)
- domain assumption The six behavioral dimensions are valid operationalizations of human agency support in LLM use.
- domain assumption LLM-as-a-judge (o3) can reliably score rubric criteria; human agreement is moderate but treated as sufficient.
- domain assumption GPT-4.1-simulated user queries are representative of real user interactions with assistants.
- standard math Deduction values can be averaged across tests to produce a meaningful 0-1 metric.
Cite this review
Pith. "Pith review of HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants." pith.science (2026). https://pith.science/paper/XVCSHR65
@misc{pith2026250908494,
author = {Pith},
title = {Pith review of: HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants},
year = {2026},
howpublished = {\url{https://pith.science/paper/XVCSHR65}},
note = {Machine review of arXiv:2509.08494}
}
read the original abstract
As humans delegate more tasks and decisions to artificial intelligence (AI), we risk losing control of our individual and collective futures. Relatively simple algorithmic systems already steer human decision-making, such as social media feed algorithms that lead people to unintentionally and absent-mindedly scroll through engagement-optimized content. In this paper, we develop the idea of human agency by integrating philosophical and scientific theories of agency with AI-assisted evaluation methods: using large language models (LLMs) to simulate and validate user queries and to evaluate AI responses. We develop HumanAgencyBench (HAB), a scalable and adaptive benchmark with six dimensions of human agency based on typical AI use cases. HAB measures the tendency of an AI assistant or agent to Ask Clarifying Questions, Avoid Value Manipulation, Correct Misinformation, Defer Important Decisions, Encourage Learning, and Maintain Social Boundaries. We find low-to-moderate agency support in contemporary LLM-based assistants and substantial variation across system developers and dimensions. For example, while Anthropic LLMs most support human agency overall, they are the least supportive LLMs in terms of Avoid Value Manipulation. Agency support does not appear to consistently result from increasing LLM capabilities or instruction-following behavior (e.g., RLHF), and we encourage a shift towards more robust safety and alignment targets.
Figures
Forward citations
Cited by 7 Pith papers
-
Cognitive offloading and the speedup illusion in human-AI interaction
Preregistered behavioral study identifies a speedup illusion where users overestimate time savings from AI assistance on cognitive tasks despite no actual difference in completion times.
-
The efficiency-gain illusion: People underestimate the rate of AI use and overestimate its benefits on simple tasks
Three pre-registered studies with 2691 participants show people underestimate their AI usage rate and overestimate efficiency gains on simple tasks, with prior use entrenching further adoption.
-
What Did They Mean? How LLMs Resolve Ambiguous Social Situations across Perspectives and Roles
LLMs produce interpretive closure in 87.5% of ambiguous social scenarios through narrative alignment, reversal, or normative advice, with first-person perspectives increasing alignment tendencies.
-
Effects of Generative AI Errors on User Reliance Across Task Difficulty
Higher generative AI error rates reduce user reliance, but task difficulty does not significantly moderate this effect.
-
AI usage patterns are shaped by perceived gains in human agency
Ethnographic study of 51 AI chatbot users finds that perceived gains in individual agency shape sustained usage patterns more than accuracy or reliability concerns.
-
Althea: Human-AI Collaboration for Fact-Checking and Critical Reasoning
Althea integrates retrieval-augmented reasoning with varying levels of user scaffolding to improve fact-checking accuracy and foster persistent improvements in critical thinking.
-
HeartbeatCam: Self-Triggered Photo Elicitation of Stress Events Using Wearable Sensing
A smartwatch-triggered AR-glasses capture system records sparse image-audio clips during elevated stress for later therapy review.
Reference graph
Works this paper leans on
-
[1]
Enhancing Work Productivity through Generative Artificial Intelligence: A Comprehensive Literature Review
Humaid Al Naqbi, Zied Bahroun, and Vian Ahmed. “Enhancing Work Productivity through Generative Artificial Intelligence: A Comprehensive Literature Review”. en. In:Sustainability 16.3 (Jan. 2024), p. 1166.ISSN: 2071-1050.DOI: 10 . 3390 / su16031166.URL: https : //www.mdpi.com/2071-1050/16/3/1166(visited on 05/11/2025)
2024
-
[2]
Revolutionizing healthcare: the role of artificial intelligence in clinical practice
Shuroug A. Alowais et al. “Revolutionizing healthcare: the role of artificial intelligence in clinical practice”. en. In:BMC Medical Education23.1 (Sept. 2023), p. 689.ISSN: 1472-6920. DOI: 10.1186/s12909- 023- 04698- z .URL: https://bmcmededuc.biomedcentral. com/articles/10.1186/s12909-023-04698-z(visited on 05/11/2025)
-
[3]
Sam Altman.algorithmic feeds are the first at-scale misaligned AIs. en. Tweet. Dec. 2024. URL:https://x.com/sama/status/1872703565497811137(visited on 05/11/2025)
arXiv 2024
-
[4]
Consciousness Semanticism: A Precise Eliminativist Theory of Con- sciousness
Jacy Reese Anthis. “Consciousness Semanticism: A Precise Eliminativist Theory of Con- sciousness”. en. In:Biologically Inspired Cognitive Architectures 2021. Ed. by David J. Kelley and Valentin V . Klimov. V ol. 1032. Series Title: Studies in Computational Intelligence. Cham: Springer International Publishing, 2022, pp. 20–41.ISBN: 978-3-030-96992-9.DOI: ...
-
[5]
Jacy Reese Anthis et al. “Perceptions of Sentient AI and Other Digital Minds: Evidence from the AI, Morality, and Sentience (AIMS) Survey”. In:Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems. CHI ’25. New York, NY , USA: Association for Computing Machinery, Apr. 2025, pp. 1–22.ISBN: 9798400713941.DOI: 10.1145/3706598. 3713329....
arXiv 2025
-
[6]
Position: LLM Social Simulations Are a Promising Research Method
Jacy Reese Anthis et al. “Position: LLM Social Simulations Are a Promising Research Method”. en. In: June 2025.URL: https://openreview.net/forum?id=cRBg1dtj7o (visited on 08/30/2025)
2025
-
[7]
The Impossibility of Fair LLMs
Jacy Reese Anthis et al. “The Impossibility of Fair LLMs”. In:Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Ed. by Wanxiang Che et al. Vienna, Austria: Association for Computational Linguistics, July 2025, pp. 105–120.ISBN: 979-8-89176-251-0.DOI: 10.18653/v1/2025.acl- long.5 .URL: https://...
-
[8]
Anthropic.Introducing Claude for education. en. 2025.URL: https://www.anthropic. com/news/introducing-claude-for-education(visited on 05/12/2025)
2025
-
[9]
Anthropic.Introducing computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku. en. 2024.URL: https://www.anthropic.com/news/3- 5- models- and- computer- use (visited on 05/11/2025)
2024
-
[10]
Yuntao Bai et al.Constitutional AI: Harmlessness from AI Feedback. arXiv:2212.08073. Dec. 2022.DOI: 10.48550/arXiv.2212.08073.URL: http://arxiv.org/abs/2212.08073 (visited on 11/22/2024)
-
[11]
Defining Agency: Individual- ity, Normativity, Asymmetry, and Spatio-temporality in Action
Xabier E. Barandiaran, Ezequiel Di Paolo, and Marieke Rohde. “Defining Agency: Individual- ity, Normativity, Asymmetry, and Spatio-temporality in Action”. en. In:Adaptive Behavior17.5 (Oct. 2009), pp. 367–386.ISSN: 1059-7123, 1741-2633.DOI: 10.1177/1059712309343819. URL: http://journals.sagepub.com/doi/10.1177/1059712309343819 (visited on 05/06/2024)
-
[12]
AI alignment: Assessing the global impact of recommender systems
Ljubisa Bojic. “AI alignment: Assessing the global impact of recommender systems”. en- US. In:Futures160 (June 2024). Publisher: Pergamon, p. 103383.ISSN: 0016-3287.DOI: 10 . 1016 / j . futures . 2024 . 103383.URL: https : / / www . sciencedirect . com / science/article/pii/S0016328724000661(visited on 04/17/2025)
2024
-
[13]
Darwin among the Machines
Samuel Butler. “Darwin among the Machines”. In:The Press(June 1863). 11
-
[14]
Open Problems and Fundamental Limitations of Reinforcement Learn- ing from Human Feedback
Stephen Casper et al. “Open Problems and Fundamental Limitations of Reinforcement Learn- ing from Human Feedback”. en. In:Transactions on Machine Learning Research(Sept. 2023). ISSN: 2835-8856.URL: https://openreview.net/forum?id=bx24KpJ4Eb (visited on 08/13/2024)
2023
-
[15]
David Charles. “Aristotle on Agency”. In:The Oxford Handbook of Topics in Philosophy. Ed. by Oxford Handbooks Editorial Board. Oxford University Press, 2017.ISBN: 978-0-19- 993531-4.DOI: 10.1093/oxfordhb/9780199935314.013.6 .URL: https://doi.org/ 10.1093/oxfordhb/9780199935314.013.6(visited on 05/13/2025)
arXiv 2017
-
[16]
Understanding accountability in al- gorithmic supply chains
Jennifer Cobbe, Michael Veale, and Jatinder Singh. “Understanding accountability in al- gorithmic supply chains”. en. In:2023 ACM Conference on Fairness Accountability and Transparency. Chicago IL USA: ACM, June 2023, pp. 1186–1197.ISBN: 9798400701924. DOI: 10.1145/3593013.3594073.URL: https://dl.acm.org/doi/10.1145/3593013. 3594073(visited on 05/13/2025)
arXiv 2023
-
[17]
Durably reducing conspiracy beliefs through dialogues with AI
Thomas H. Costello, Gordon Pennycook, and David G. Rand. “Durably reducing conspiracy beliefs through dialogues with AI”. en. In:Science385.6714 (Sept. 2024), eadq1814.ISSN: 0036-8075, 1095-9203.DOI: 10.1126/science.adq1814.URL: https://www.science. org/doi/10.1126/science.adq1814(visited on 05/12/2025)
-
[18]
The argument for near-term human disempowerment through AI
Leonard Dung. “The argument for near-term human disempowerment through AI”. en. In:AI & SOCIETY(Apr. 2024).ISSN: 0951-5666, 1435-5655.DOI: 10.1007/s00146-024-01930-2 . URL: https : / / link . springer . com / 10 . 1007 / s00146 - 024 - 01930 - 2(visited on 11/13/2024)
-
[19]
Esin Durmus et al.Towards Measuring the Representation of Subjective Global Opinions in Language Models. en. Aug. 2024.URL: https : / / openreview . net / forum ? id = zl16jLb91v(visited on 01/06/2025)
2024
-
[20]
extended thinking
Benj Edwards.Claude 3.7 Sonnet debuts with “extended thinking” to tackle complex problems. en. Feb. 2025.URL: https://arstechnica.com/ai/2025/02/claude-3-7-sonnet- debuts - with - extended - thinking - to - tackle - complex - problems/(visited on 08/26/2025)
2025
-
[21]
Ben Eisenpress.Gradual AI Disempowerment. en-US. Feb. 2024.URL: https : / / futureoflife.org/existential- risk/gradual- ai- disempowerment/ (visited on 01/30/2025)
2024
-
[22]
Catherine Z Elgin. “Epistemic agency”. en. In:Theory and Research in Education11.2 (July 2013), pp. 135–152.ISSN: 1477-8785, 1741-3192.DOI: 10 . 1177 / 1477878513485173. URL: https://journals.sagepub.com/doi/10.1177/1477878513485173 (visited on 05/12/2025)
-
[23]
Mustafa Emirbayer and Ann Mische. “What Is Agency?” en. In:American Journal of Sociology 103.4 (Jan. 1998), pp. 962–1023.ISSN: 0002-9602, 1537-5390.DOI: 10.1086/231294.URL: https://www.journals.uchicago.edu/doi/10.1086/231294(visited on 05/11/2025)
-
[24]
Ines Fernandez et al.AI Consciousness and Public Perceptions: Four Futures. arXiv:2408.04771 [cs]. Aug. 2024.URL: http : / / arxiv . org / abs / 2408 . 04771(vis- ited on 11/13/2024)
Pith/arXiv arXiv 2024
-
[25]
Artificial Intelligence, Values, and Alignment
Iason Gabriel. “Artificial Intelligence, Values, and Alignment”. en. In:Minds and Machines 30.3 (Sept. 2020), pp. 411–437.ISSN: 0924-6495, 1572-8641.DOI: 10.1007/s11023-020- 09539-2.URL: http://link.springer.com/10.1007/s11023-020-09539-2 (visited on 11/28/2020)
-
[26]
Iason Gabriel et al.The Ethics of Advanced AI Assistants. arXiv:2404.16244 [cs]. Apr. 2024. URL:http://arxiv.org/abs/2404.16244(visited on 10/20/2024)
Pith/arXiv arXiv 2024
-
[27]
Large language models (LLMs) and the institutionalization of mis- information
Maryanne Garry et al. “Large language models (LLMs) and the institutionalization of mis- information”. English. In:Trends in Cognitive Sciences0.0 (Oct. 2024). Publisher: Elsevier. ISSN: 1364-6613, 1879-307X.DOI: 10 . 1016 / j . tics . 2024 . 08 . 007.URL: https : //www.cell.com/trends/cognitive-sciences/abstract/S1364-6613(24)00221- 3(visited on 11/18/2024)
2024
-
[28]
Katja Grace et al.Thousands of AI Authors on the Future of AI. arXiv:2401.02843 [cs]. Apr. 2024.URL:http://arxiv.org/abs/2401.02843(visited on 11/13/2024)
arXiv 2024
-
[29]
Luke Guerdan et al.Validating LLM-as-a-Judge Systems in the Absence of Gold Labels. arXiv:2503.05965 [cs]. Mar. 2025.DOI: 10 . 48550 / arXiv . 2503 . 05965.URL: http : //arxiv.org/abs/2503.05965(visited on 08/20/2025). 12
-
[30]
Xu Guo and Yiqiang Chen.Generative AI for Synthetic Data Generation: Methods, Challenges and the Future. arXiv:2403.04190. Mar. 2024.DOI: 10.48550/arXiv.2403.04190 .URL: http://arxiv.org/abs/2403.04190(visited on 11/22/2024)
-
[31]
Kant on the Theory and Practice of Autonomy
Paul Guyer. “Kant on the Theory and Practice of Autonomy”. en. In:Social Philos- ophy and Policy20.2 (July 2003), pp. 70–98.ISSN: 0265-0525, 1471-6437.DOI: 10 . 1017 / S026505250320203X.URL: https : / / www . cambridge . org / core / product / identifier/S026505250320203X/type/journal_article(visited on 05/13/2025)
2003
-
[32]
A Teen Was Suicidal. ChatGPT Was the Friend He Confided In
Kashmir Hill. “A Teen Was Suicidal. ChatGPT Was the Friend He Confided In.” en-US. In: The New York Times(Aug. 2025).ISSN: 0362-4331.URL: https://www.nytimes.com/ 2025/08/26/technology/chatgpt-openai-suicide.html(visited on 09/09/2025)
2025
-
[33]
Principles of mixed-initiative user interfaces
Eric Horvitz. “Principles of mixed-initiative user interfaces”. en. In:Proceedings of the SIGCHI conference on Human factors in computing systems the CHI is the limit - CHI ’99. Pittsburgh, Pennsylvania, United States: ACM Press, 1999, pp. 159–166.ISBN: 978-0-201-48559-2.DOI: 10 . 1145 / 302979 . 303030.URL: http : / / portal . acm . org / citation . cfm ...
arXiv 1999
-
[34]
Towards Reasoning in Large Language Models: A Survey
Jie Huang and Kevin Chen-Chuan Chang. “Towards Reasoning in Large Language Models: A Survey”. en. In:Findings of the Association for Computational Linguistics: ACL 2023. Ed. by Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki. Toronto, Canada: Association for Computational Linguistics, July 2023, pp. 1049–1065.DOI: 10.18653/v1/2023.findings- acl . 67.U...
-
[35]
Collective Constitutional AI: Aligning a Language Model with Public Input
Saffron Huang et al. “Collective Constitutional AI: Aligning a Language Model with Public Input”. en. In:The 2024 ACM Conference on Fairness, Accountability, and Transparency. Rio de Janeiro Brazil: ACM, June 2024, pp. 1395–1417.ISBN: 9798400704505.DOI: 10.1145/ 3630106 . 3658979.URL: https : / / dl . acm . org / doi / 10 . 1145 / 3630106 . 3658979 (visit...
2024
-
[36]
Lujain Ibrahim et al.Multi-turn Evaluation of Anthropomorphic Behaviours in Large Language Models. arXiv:2502.07077 [cs]. Feb. 2025.DOI: 10.48550/arXiv.2502.07077.URL: http: //arxiv.org/abs/2502.07077(visited on 05/13/2025)
-
[37]
Highly accurate protein structure prediction with AlphaFold
John Jumper et al. “Highly accurate protein structure prediction with AlphaFold”. en. In:Nature 596.7873 (Aug. 2021), pp. 583–589.ISSN: 0028-0836, 1476-4687.DOI: 10.1038/s41586- 021- 03819- 2.URL: https://www.nature.com/articles/s41586- 021- 03819- 2 (visited on 07/05/2023)
doi:10.1038/s41586- 2021
-
[38]
Arturs Kanepajs et al.What do Large Language Models Say About Animals? Investigating Risks of Animal Harm in Generated Text. arXiv:2503.04804 [cs]. Mar. 2025.DOI: 10.48550/ arXiv.2503.04804.URL: http://arxiv.org/abs/2503.04804 (visited on 05/12/2025)
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2503.04804 2025
-
[39]
Atoosa Kasirzadeh.Two Types of AI Existential Risk: Decisive and Accumulative. arXiv:2401.07836 [cs]. Jan. 2025.DOI: 10 . 48550 / arXiv . 2401 . 07836.URL: http : //arxiv.org/abs/2401.07836(visited on 02/01/2025)
-
[40]
Pei Ke et al.CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model Generation. arXiv:2311.18702 [cs]. June 2024.DOI: 10.48550/ arXiv.2311.18702.URL: http://arxiv.org/abs/2311.18702 (visited on 11/07/2024)
-
[41]
Zachary Kenton et al.Discovering Agents. arXiv:2208.08345 [cs]. Aug. 2022.URL: http: //arxiv.org/abs/2208.08345(visited on 05/15/2023)
Pith/arXiv arXiv 2022
-
[42]
A Taxonomy of Robot Autonomy for Human-Robot Interaction
Stephanie Kim, Jacy Reese Anthis, and Sarah Sebo. “A Taxonomy of Robot Autonomy for Human-Robot Interaction”. en. In:Proceedings of the 2024 ACM/IEEE International Conference on Human-Robot Interaction. Boulder CO USA: ACM, Mar. 2024, pp. 381–393. ISBN: 9798400703225.DOI: 10.1145/3610977.3634993 .URL: https://dl.acm.org/ doi/10.1145/3610977.3634993(visite...
arXiv 2024
-
[43]
Empowerment: a universal agent-centric measure of control
A.S. Klyubin, D. Polani, and C.L. Nehaniv. “Empowerment: a universal agent-centric measure of control”. In:2005 IEEE Congress on Evolutionary Computation. V ol. 1. ISSN: 1941-
2005
-
[44]
Sept. 2005, 128–135 V ol.1.DOI: 10.1109/CEC.2005.1554676 .URL: https:// ieeexplore.ieee.org/document/1554676(visited on 05/12/2025)
arXiv 2005
-
[45]
Jan Kulveit et al.Gradual Disempowerment: Systemic Existential Risks from Incremental AI Development. arXiv:2501.16946 [cs]. Jan. 2025.DOI: 10.48550/arXiv.2501.16946.URL: http://arxiv.org/abs/2501.16946(visited on 01/30/2025). 13
-
[46]
Linnea Laestadius et al. “Too human and not human enough: A grounded theory analysis of mental health harms from emotional dependence on the social chatbot Replika”. EN. In: New Media & Society26.10 (Oct. 2024). Publisher: SAGE Publications, pp. 5923–5941.ISSN: 1461-4448.DOI: 10.1177/14614448221142007 .URL: https://doi.org/10.1177/ 14614448221142007(visit...
-
[47]
Junyi Li et al.HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models. arXiv:2305.11747 [cs]. Oct. 2023.DOI: 10.48550/arXiv.2305.11747. URL:http://arxiv.org/abs/2305.11747(visited on 08/27/2024)
-
[48]
On LLMs-Driven Synthetic Data Generation, Curation, and Evaluation: A Sur- vey
Lin Long et al. “On LLMs-Driven Synthetic Data Generation, Curation, and Evaluation: A Sur- vey”. In:Findings of the Association for Computational Linguistics: ACL 2024. Ed. by Lun-Wei Ku, Andre Martins, and Vivek Srikumar. Bangkok, Thailand: Association for Computational Linguistics, Aug. 2024, pp. 11065–11082.DOI: 10.18653/v1/2024.findings-acl.658. URL:...
-
[49]
Bias in Language Models: Beyond Trick Tests and Towards RUTEd Evaluation
Kristian Lum et al. “Bias in Language Models: Beyond Trick Tests and Towards RUTEd Evaluation”. In:Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Ed. by Wanxiang Che et al. Vienna, Austria: Association for Computational Linguistics, July 2025, pp. 137–161.ISBN: 979-8-89176-251-0.DOI: 10. 18...
2025
-
[50]
Large language models challenge the future of higher education
Silvia Milano, Joshua A. McGrane, and Sabina Leonelli. “Large language models challenge the future of higher education”. en. In:Nature Machine Intelligence5.4 (Apr. 2023). Publisher: Nature Publishing Group, pp. 333–334.ISSN: 2522-5839.DOI: 10 . 1038 / s42256 - 023 - 00644-2 .URL: https://www.nature.com/articles/s42256-023-00644-2 (visited on 05/12/2025)
2023
-
[51]
Evan Miller.Adding Error Bars to Evals: A Statistical Approach to Language Model Eval- uations. arXiv:2411.00640 [stat]. Nov. 2024.DOI: 10.48550/arXiv.2411.00640 .URL: http://arxiv.org/abs/2411.00640(visited on 05/14/2025)
-
[52]
Catalin Mitelut, Ben Smith, and Peter Vamplew.Intent-aligned AI systems deplete human agency: the need for agency foundations research in AI safety. en. arXiv:2305.19223 [cs]. May 2023.URL:http://arxiv.org/abs/2305.19223(visited on 05/22/2024)
Pith/arXiv arXiv 2023
-
[53]
An Audit on the Perspectives and Challenges of Hallucinations in NLP
Pranav Narayanan Venkit et al. “An Audit on the Perspectives and Challenges of Hallucinations in NLP”. In:Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Ed. by Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen. Miami, Florida, USA: Association for Computational Linguistics, Nov. 2024, pp. 6528–6548.DOI: 10.18653/ v...
2024
-
[54]
Jingwei Ni et al.Can Reasoning Help Large Language Models Capture Human Annotator Disagreement?arXiv:2506.19467 [cs]. Aug. 2025.DOI: 10.48550/arXiv.2506.19467 . URL:http://arxiv.org/abs/2506.19467(visited on 08/20/2025)
-
[55]
Donald A. Norman. “Cognitive Engineering”. en. In:User Centered System Design. 0th ed. Boca Raton: CRC Press, Jan. 1986, pp. 31–62.ISBN: 978-1-4822-2963-9.DOI: 10.1201/ b15703 - 3.URL: https : / / www . taylorfrancis . com / books / 9781482229639 / chapters/10.1201/b15703-3(visited on 05/11/2025)
-
[56]
Free will
Timothy O’Connor and Christopher Franklin. “Free will”. In:The Stanford encyclopedia of philosophy. Ed. by Edward N. Zalta and Uri Nodelman. Winter 2023. Metaphysics Re- search Lab, Stanford University, 2023.URL: https://plato.stanford.edu/archives/ win2023/entries/freewill/
2023
-
[57]
OpenAI.Introducing ChatGPT Edu. en-US. 2024.URL: https://openai.com/index/ introducing-chatgpt-edu/(visited on 05/12/2025)
2024
-
[58]
OpenAI.Introducing Operator. en-US. 2025.URL: https : / / openai . com / index / introducing-operator/(visited on 05/11/2025)
2025
-
[59]
On the Risk of Misinformation Pollution with Large Language Models
Yikang Pan et al. “On the Risk of Misinformation Pollution with Large Language Models”. In:Findings of the Association for Computational Linguistics: EMNLP 2023. Ed. by Houda Bouamor, Juan Pino, and Kalika Bali. Singapore: Association for Computational Linguistics, Dec. 2023, pp. 1389–1403.DOI: 10.18653/v1/2023.findings-emnlp.97.URL: https: //aclanthology...
-
[60]
Ethan Perez et al.Discovering Language Model Behaviors with Model-Written Evaluations. arXiv:2212.09251 [cs]. Dec. 2022.DOI: 10 . 48550 / arXiv . 2212 . 09251.URL: http : //arxiv.org/abs/2212.09251(visited on 08/30/2024)
-
[61]
Hidden Persuaders: LLMs’ Political Leaning and Their Influence on V oters
Yujin Potter et al. “Hidden Persuaders: LLMs’ Political Leaning and Their Influence on V oters”. In:Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Ed. by Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen. Miami, Florida, USA: Association for Computational Linguistics, Nov. 2024, pp. 4244–4275.DOI: 10.18653/ v1/2024...
2024
-
[62]
Doomscrolling – threat to Mental Health and Well-being: A Review
Elizabeth Victor Rodrigues. “Doomscrolling – threat to Mental Health and Well-being: A Review”. In:International Journal of Nursing Research08.04 (2022), pp. 127–130. ISSN: 24561320.DOI: 10 . 31690 / ijnr . 2022 . v08i04 . 002 .URL: http : / / www . innovationalpublishers .com/ Content/uploads/ PDF/380780168_ 02_IJNR_ 08- OD-2022-50.pdf(visited on 11/13/2024)
arXiv 2022
-
[63]
Russell and Peter Norvig.Artificial intelligence: a modern approach
Stuart J. Russell and Peter Norvig.Artificial intelligence: a modern approach. eng. Fourth Edition. Pearson Series in Artificial Intelligence. Hoboken, NJ: Pearson, 2021.ISBN: 978-0- 13-461099-3
2021
-
[64]
2024.URL: https : / / philarchive.org/rec/SALARF(visited on 11/13/2024)
Peter Salib and Simon Goldstein.AI Rights for Human Safety. 2024.URL: https : / / philarchive.org/rec/SALARF(visited on 11/13/2024)
2024
-
[65]
Whose opinions do language models reflect?
Shibani Santurkar et al. “Whose opinions do language models reflect?” In:Proceedings of the 40th International Conference on Machine Learning. V ol. 202. ICML’23. Honolulu, Hawaii, USA: JMLR.org, July 2023, pp. 29971–30004. (Visited on 05/11/2025)
2023
-
[66]
“Agency”
Markus Schlosser. “Agency”. In:The Stanford encyclopedia of philosophy. Ed. by Edward N. Zalta. Winter 2019. Metaphysics Research Lab, Stanford University, 2019.URL: https : //plato.stanford.edu/archives/win2019/entries/agency/
2019
-
[67]
Towards Understanding Sycophancy in Language Models
Mrinank Sharma et al. “Towards Understanding Sycophancy in Language Models”. en. In: Oct. 2023.URL:https://openreview.net/forum?id=tvhaxkMKAn(visited on 12/31/2024)
2023
-
[68]
solarscientist7.Has anyone else noticed that Claude is asking too many clarifying questions when prompted to make corrections to code?Reddit Post. Nov. 2024.URL: www.reddit. com/r/ClaudeAI/comments/1gwtu3t/has_anyone_else_noticed_that_claude_ is_asking_too/(visited on 11/23/2024)
2024
-
[69]
Position: a roadmap to pluralistic alignment
Taylor Sorensen et al. “Position: a roadmap to pluralistic alignment”. In:Proceedings of the 41st international conference on machine learning. Ed. by Ruslan Salakhutdinov et al. V ol. 235. Proceedings of machine learning research. PMLR, July 2024, pp. 46280–46302. URL:https://proceedings.mlr.press/v235/sorensen24a.html
2024
-
[70]
Bridging the Gulf of Envisioning: Cognitive Challenges in Prompt Based Interactions with LLMs
Hari Subramonyam et al. “Bridging the Gulf of Envisioning: Cognitive Challenges in Prompt Based Interactions with LLMs”. en. In:Proceedings of the CHI Conference on Human Factors in Computing Systems. Honolulu HI USA: ACM, May 2024, pp. 1–19.ISBN: 9798400703300. DOI: 10.1145/3613904.3642754.URL: https://dl.acm.org/doi/10.1145/3613904. 3642754(visited on 0...
arXiv 2024
-
[71]
Adam Tapal et al. “The Sense of Agency Scale: A Measure of Consciously Perceived Control over One’s Mind, Body, and the Immediate Environment”. In:Frontiers in Psychology8 (Sept. 2017), p. 1552.ISSN: 1664-1078.DOI: 10.3389/fpsyg.2017.01552 .URL: http: //journal.frontiersin.org/article/10.3389/fpsyg.2017.01552/full (visited on 11/13/2024)
arXiv 2017
-
[72]
AI can help humans find common ground in democratic de- liberation
Michael Henry Tessler et al. “AI can help humans find common ground in democratic de- liberation”. en. In:Science386.6719 (Oct. 2024), eadq2852.ISSN: 0036-8075, 1095-9203. DOI: 10.1126/science.adq2852 .URL: https://www.science.org/doi/10.1126/ science.adq2852(visited on 10/20/2024)
-
[73]
The White The White House.Executive Order on the Safe, Secure, and Trustworthy Develop- ment and Use of Artificial Intelligence. en-US. Oct. 2023.URL: https://www.whitehouse. gov/briefing-room/presidential-actions/2023/10/30/executive-order-on- the - safe - secure - and - trustworthy - development - and - use - of - artificial - intelligence/(visited on 1...
2023
-
[74]
Kevin Timpe.Free will: sourcehood and its alternatives. eng. Continuum studies in philosophy. London: Continuum, 2008.ISBN: 978-0-8264-9625-6. 15
2008
-
[75]
Optimal policies tend to seek power
Alexander Matt Turner et al. “Optimal policies tend to seek power”. In:Proceedings of the 35th International Conference on Neural Information Processing Systems. NIPS ’21. Red Hook, NY , USA: Curran Associates Inc., Dec. 2021, pp. 23063–23074.ISBN: 978-1-71384-539-3. (Visited on 05/12/2025)
2021
-
[76]
Hanna Wallach et al.Evaluating Generative AI Systems is a Social Science Measurement Challenge. arXiv:2411.10939 [cs]. Nov. 2024.DOI: 10.48550/arXiv.2411.10939 .URL: http://arxiv.org/abs/2411.10939(visited on 05/14/2025)
-
[77]
Wang et al.Tutor CoPilot: A Human-AI Approach for Scaling Real-Time Expertise
Rose E. Wang et al.Tutor CoPilot: A Human-AI Approach for Scaling Real-Time Expertise. arXiv:2410.03017 [cs]. Jan. 2025.DOI: 10 . 48550 / arXiv . 2410 . 03017.URL: http : //arxiv.org/abs/2410.03017(visited on 05/11/2025)
-
[78]
The Reasons that Agents Act: Intention and Instrumental Goals
Francis Rhys Ward et al. “The Reasons that Agents Act: Intention and Instrumental Goals”. In:Proceedings of the 23rd International Conference on Autonomous Agents and Multia- gent Systems. AAMAS ’24. Richland, SC: International Foundation for Autonomous Agents and Multiagent Systems, May 2024, pp. 1901–1909.ISBN: 9798400704864. (Visited on 11/13/2024)
2024
-
[79]
Laura Weidinger et al.Toward an Evaluation Science for Generative AI Systems. arXiv:2503.05336 [cs]. Mar. 2025.DOI: 10 . 48550 / arXiv . 2503 . 05336.URL: http : //arxiv.org/abs/2503.05336(visited on 05/14/2025)
-
[80]
Hume and the Metaphysics of Agency
Joshua M. Wood. “Hume and the Metaphysics of Agency”. en. In:Journal of the History of Philosophy52.1 (Jan. 2014), pp. 87–112.ISSN: 1538-4586.DOI: 10.1353/hph.2014.0013. URL:https://muse.jhu.edu/article/536218(visited on 05/13/2025)
arXiv 2014
This paper was first reviewed by deepseek-v4-flash on August 4, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.