Pith. sign in

REVIEW 2 major objections 2 minor 85 references

AnnotateThis: Analyzing a human-LLM system for annotating social media data with the concept of climate change mitigation pessimism

T0 review · 2 major / 2 minor · reviewed 2026-06-27 · grok-4.3

Pith's one-line read Users improve LLM annotations of nuanced concepts like climate change mitigation pessimism beyond fully automated methods.

desk verdict AnnotateThis gives a practical interface for humans to inspect and fix LLM annotations on nuanced concepts, with claimed gains in two settings, but the study details are too thin to assess the numbers. read the letter →

arxiv 2606.10210 v1 pith:OZSQT5HG submitted 2026-06-08 cs.CY

classification cs.CY
keywords human-LLMcollaborationdataannotationLLMgroundingclimatechangemitigationcomputationalsocialsciencesystemspromptrefinement
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper introduces AnnotateThis, a system that lets humans inspect and refine large language model outputs for complex social science concepts. It supports two evaluation settings: one where users define the target concept at the same time as grounding the model, and another where the concept is already specified and ground truth labels exist. In both cases, human users produce final annotations that exceed those from LLMs without intervention. When ground truth is available, the approach yields an absolute gain of 0.15 in F-Measure and 0.23 in accuracy over a state-of-the-art automated prompt refinement baseline.

What carries the argument

AnnotateThis, a system of information features that let users interrogate the quality and reliability of LLM annotations during concept specification and grounding.

What would settle it

A controlled study in which users of AnnotateThis produce no higher F-Measure or accuracy than a fully automated prompt-refinement baseline when both are evaluated against the same ground truth labels.

Watch

Extended reading notes

Core claim

AnnotateThis is a human-centered system developed with computational and social scientists that supplies information features for interrogating LLM annotations, enabling a process called LLM grounding for a target concept. In the first setting, users simultaneously specify the concept and improve annotations without ground truth. In the second, with ground truth and a fixed concept, users achieve measurable gains over automation alone.

Load-bearing premise

The system's information features are enough for non-expert users to meaningfully improve LLM outputs even while they are defining the target concept at the same time.

Editorial extensions

If this is right

  • Users achieve higher quality annotations than LLMs alone in settings without ground truth labels.
  • Final annotations surpass those created without human intervention when ground truth is available.
  • The system supports existing workflows for data annotation in computational social science.
  • Gains of 0.15 F-Measure and 0.23 accuracy hold when evaluating against ground truth for the tested concept.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same human-grounding approach could apply to other difficult CSS concepts where LLMs currently underperform.
  • Wider adoption might allow larger annotation tasks with fewer expert annotators by leveraging non-experts.
  • Observed user improvements could guide the design of additional automated assistance within similar systems.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 2 minor

Summary. The paper introduces AnnotateThis, a human-centered system for inspecting and refining LLM annotations of social media data for nuanced CSS concepts such as climate change mitigation pessimism. It evaluates the system in two settings: (1) simultaneous concept specification and LLM grounding without ground truth, and (2) LLM grounding with ground truth available. The central claim is that users improve annotation quality in both settings, with final human-refined annotations substantially outperforming fully automated SOTA prompt-refinement baselines (e.g., absolute gains of 0.15 F-measure and 0.23 accuracy in the ground-truth setting).

Significance. If the user-study results are reproducible, the work offers a concrete, workflow-aligned example of human-LLM collaboration that directly responds to calls in the CSS community for systems that keep humans in the loop on difficult annotation tasks. The dual-setting design and explicit comparison to an external automated baseline are strengths that could inform future tool development.

major comments (2)
  1. [Evaluation / Abstract] The abstract and evaluation description report quantitative gains (0.15 F-measure, 0.23 accuracy) but provide no information on user-study design, participant sample size, recruitment criteria, expertise screening, task duration, or statistical tests. Without these details the reported improvements cannot be assessed for selection bias, order effects, or generalizability.
  2. [Evaluation (first setting)] The first evaluation setting simultaneously requires users to specify the target concept and ground the LLM using the provided information features. The manuscript does not describe how concept specification was operationalized, what instructions were given, or what controls prevented conflation of concept-definition effects with system-feature effects; this assumption is load-bearing for the claim that AnnotateThis enables improvement even under simultaneous specification.
minor comments (2)
  1. [Introduction] The term 'LLM grounding' is introduced without a formal definition or reference to prior usage in the CSS or HCI literature.
  2. [Results figures] Figure captions and axis labels should explicitly state the exact metric (micro/macro F1, exact-match accuracy) and the baseline method name for each comparison.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for their constructive feedback on the evaluation sections of our manuscript. We address each major comment below and have revised the paper to incorporate additional details where the original description was insufficient.

read point-by-point responses
  1. Referee: [Evaluation / Abstract] The abstract and evaluation description report quantitative gains (0.15 F-measure, 0.23 accuracy) but provide no information on user-study design, participant sample size, recruitment criteria, expertise screening, task duration, or statistical tests. Without these details the reported improvements cannot be assessed for selection bias, order effects, or generalizability.

    Authors: We agree that the manuscript as submitted did not provide sufficient methodological detail on the user study to allow full assessment of the reported gains. In the revised version we have added a new subsection (Evaluation: User Study Protocol) that specifies the participant sample size, recruitment criteria and channels, expertise screening procedures, average task duration, and the statistical tests used to evaluate improvements. These additions directly address concerns about selection bias, order effects, and generalizability. revision: yes

  2. Referee: [Evaluation (first setting)] The first evaluation setting simultaneously requires users to specify the target concept and ground the LLM using the provided information features. The manuscript does not describe how concept specification was operationalized, what instructions were given, or what controls prevented conflation of concept-definition effects with system-feature effects; this assumption is load-bearing for the claim that AnnotateThis enables improvement even under simultaneous specification.

    Authors: We acknowledge that the original manuscript provided only a high-level description of the first setting and did not explicitly operationalize concept specification or detail the controls separating definition effects from system-feature effects. The revised manuscript now includes a dedicated paragraph in the first evaluation subsection that describes the exact instructions given to participants, the base concept prompt supplied, the sequence of specification and grounding steps, and the within-subject measures used to isolate system contributions. This clarification supports the claim while preserving the dual-setting design. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity

full rationale

The paper is an empirical system evaluation reporting direct metric comparisons (F-Measure, accuracy) against external baselines and ground truth. No derivations, equations, fitted parameters presented as predictions, uniqueness theorems, or self-citation load-bearing steps exist in the described workflow. The central claims rest on observed improvements in two evaluation settings, which are externally falsifiable and do not reduce to the paper's own inputs by construction.

Assumptions & free parameters 0 free parameters · 0 assumptions · 0 invented entities

No mathematical model, free parameters, axioms, or invented entities are present; the contribution is an applied human-AI annotation interface evaluated empirically.

how reviews work

0 comments
Cite this review

Pith. "Pith review of AnnotateThis: Analyzing a human-LLM system for annotating social media data with the concept of climate change mitigation pessimism." pith.science (2026). https://pith.science/paper/OZSQT5HG

@misc{pith2026260610210,
  author       = {Pith},
  title        = {Pith review of: AnnotateThis: Analyzing a human-LLM system for annotating social media data with the concept of climate change mitigation pessimism},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/OZSQT5HG}},
  note         = {Machine review of arXiv:2606.10210}
}
read the original abstract

Large language models (LLMs) are increasingly being integrated into research workflows. However, LLMs have been shown to struggle with difficult and nuanced concepts such as those found in computational social science (CSS) research. Within the CSS community, there has been a call for new systems to be developed which center humans in LLM-supported scientific workflows. We develop AnnotateThis, a human-centered system for inspecting and improving LLM annotations, a process we refer to as LLM grounding for a target concept. AnnotateThis is developed with both computational and social scientists to reflect existing workflows for data annotation. It includes a range of information features for users to interrogate the quality and reliability of LLM annotations. We evaluate our system in two settings. In the first, we assume a researcher may not have access to ground truth data and that users of AnnotateThis have limited prior knowledge of the concept they would like an LLM to annotate. That is, they may be conducting concept specification and LLM grounding simultaneously. In the second setting, we assume access to ground truth labels and that the concept is specified for a given annotation task; here, the task of LLM grounding is more straightforward. We find that in both settings users can improve the quality of LLM annotations with AnnotateThis and that their final annotations far surpass those created without human intervention. For example, when we evaluate with ground truth labels, we see an absolute improvement of 0.15 in F-Measure and 0.23 in accuracy over a fully automated state-of-the-art method for prompt refinement.

Figures

Figures reproduced from arXiv: 2606.10210 by the authors.

Figure 1
Figure 1. ANNOTATETHIS system workflow. Participants iteratively write instructions, revisiting the instruction hub and infor￾mation features many times. A-Consistency an estimate of model consistency, A-Explanation an explanation for the AI label which is generated by the LLM. These information features were designed to provide in￾formation about label quality deemed important by social scientists, so that (1) participants h… view at source ↗
Figure 2
Figure 2. Screenshot of Annotation view with information features, the scrollable table contains video URLs and transcripts to the left. Here the expert label refers to the ground truth produced by the research team. This label was only shown in Study 2, in Study 1, we assume ground truth labels do not exist. Here, we expect that participants will use ANNOTATETHIS for concept specification while iterating on LLM grounding. Th… view at source ↗
Figure 3
Figure 3. The number of times features were used. By read [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (14 more)
Figure 4
Figure 4. Figure 4: Participants completed the SUS with respect to A [PITH_FULL_IMAGE:figures/full_fig_p008_4.png]
Figure 8
Figure 8. Figure 8: The figure shows the complete five steps in the dataset [PITH_FULL_IMAGE:figures/full_fig_p012_8.png]
Figure 6
Figure 6. Figure 6: Participants in Study 2 were asked to rate the help￾fulness of each feature. We see that the features that included natural language tended to be viewed as more helpful than the statistical features [PITH_FULL_IMAGE:figures/full_fig_p013_6.png]
Figure 5
Figure 5. Figure 5: Percentage of “Yes” labels for different groups. [PITH_FULL_IMAGE:figures/full_fig_p013_5.png]
Figure 7
Figure 7. Figure 7: Keywords selection process. The team finalized the keyword set over two stages. Newly added keywords are high [PITH_FULL_IMAGE:figures/full_fig_p014_7.png]
Figure 8
Figure 8. Figure 8: TikTok videos dataset curation process [PITH_FULL_IMAGE:figures/full_fig_p014_8.png]
Figure 9
Figure 9. Figure 9: Screenshot of S-Trend feature [PITH_FULL_IMAGE:figures/full_fig_p016_9.png]
Figure 10
Figure 10. Figure 10: Screenshot of S- Runs feature [PITH_FULL_IMAGE:figures/full_fig_p017_10.png]
Figure 11
Figure 11. Figure 11: Screenshot of S-Bucket feature [PITH_FULL_IMAGE:figures/full_fig_p017_11.png]
Figure 12
Figure 12. Figure 12: Screenshot of S- Summary feature [PITH_FULL_IMAGE:figures/full_fig_p018_12.png]
Figure 13
Figure 13. Figure 13: Screenshot of S- Agreement feature [PITH_FULL_IMAGE:figures/full_fig_p019_13.png]
Figure 14
Figure 14. Figure 14: Screenshot of instructions in ANNOTATETHIS [PITH_FULL_IMAGE:figures/full_fig_p020_14.png]
Figure 15
Figure 15. Figure 15: Screenshot of Personal LLM Instruction Hub in A [PITH_FULL_IMAGE:figures/full_fig_p021_15.png]
Figure 16
Figure 16. Figure 16: Screenshot of A-Table in ANNOTATETHIS [PITH_FULL_IMAGE:figures/full_fig_p021_16.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

85 extracted references · 5 canonical work pages

  1. [1]

    Showing They Care (or Don’t): Affective Publics and Ambivalent Climate Activism on

    Hautea, Samantha and Parks, Perry and Takahashi, Bruno and Zeng, Jing , year =. Showing They Care (or Don’t): Affective Publics and Ambivalent Climate Activism on

  2. [2]

    Mapping the Climate Change Landscape on

    Galdeman, Alessia and Aiello, Luca Maria , year =. Mapping the Climate Change Landscape on

  3. [3]

    2022 , title =

    Treen, Kathie and Williams, Hywel and O. 2022 , title =

  4. [4]

    Interpretable machine learning: fundamental principles and 10 grand challenges , journal =

    Rudin, Cynthia and Chen, Chaofan and Chen, Zhi and Huang, Haiyang and Semenova, Lesia and Zhong, Chudi , year =. Interpretable machine learning: fundamental principles and 10 grand challenges , journal =

  5. [5]

    Rethinking Interpretability in the Era of Large Language Models , howpublished =

    Chandan Singh and Jeevana Priya Inala and Michel Galley and Rich Caruana and Jianfeng Gao , year =. Rethinking Interpretability in the Era of Large Language Models , howpublished =

  6. [6]

    Human-LLM collaborative annotation through effective verification of LLM labels , booktitle =

    Wang, Xinru and Kim, Hannah and Rahman, Sajjadur and Mitra, Kushan and Miao, Zhengjie , year =. Human-LLM collaborative annotation through effective verification of LLM labels , booktitle =

  7. [7]

    Davidson , year =

    Youngjin Chae and Thomas R. Davidson , year =. Large Language Models for Text Classification: From Zero-Shot Learning to Instruction-Tuning , journal =

  8. [8]

    Why Johnny can’t prompt: how non-AI experts try (and fail) to design LLM prompts , booktitle =

    Zamfirescu-Pereira, J Diego and Wong, Richmond Y and Hartmann, Bjoern and Yang, Qian , year =. Why Johnny can’t prompt: how non-AI experts try (and fail) to design LLM prompts , booktitle =

Show all 85 references
  1. [9]

    Social media news and its impact , publisher =

    Song, Hyunjin and de Z. Social media news and its impact , publisher =. 2021 , title =

  2. [10]

    When social media attack: How exposure to political attacks on social media promotes anger and political cynicism , journal =

    Hasell, Ariel and Halversen, Audrey and Weeks, Brian E , year =. When social media attack: How exposure to political attacks on social media promotes anger and political cynicism , journal =

  3. [11]

    2025 , title =

    Lane, Daniel S and Molina-Rogers, Nancy and Gagr. 2025 , title =

  4. [12]

    If in a crowdsourced data annotation pipeline, a gpt-4 , booktitle =

    He, Zeyu and Huang, Chieh-Yang and Ding, Chien-Kuang Cornelia and Rohatgi, Shaurya and Huang, Ting-Hao Kenneth , year =. If in a crowdsourced data annotation pipeline, a gpt-4 , booktitle =

  5. [13]

    What should we engineer in prompts? training humans in requirement-driven llm use , journal =

    Ma, Qianou and Peng, Weirui and Yang, Chenyang and Shen, Hua and Koedinger, Ken and Wu, Tongshuang , year =. What should we engineer in prompts? training humans in requirement-driven llm use , journal =

  6. [14]

    Using rhetorical strategies to design prompts: a human-in-the-loop approach to make AI useful , journal =

    Ranade, Nupoor and Saravia, Marly and Johri, Aditya , year =. Using rhetorical strategies to design prompts: a human-in-the-loop approach to make AI useful , journal =

  7. [15]

    Coannotating: Uncertainty-guided work allocation between human and large language models for data annotation , howpublished =

    Li, Minzhi and Shi, Taiwei and Ziems, Caleb and Kan, Min-Yen and Chen, Nancy F and Liu, Zhengyuan and Yang, Diyi , year =. Coannotating: Uncertainty-guided work allocation between human and large language models for data annotation , howpublished =

  8. [16]

    2024 , title =

    Shin, Joongi and Hedderich, Michael A and Rey, Bart. 2024 , title =

  9. [17]

    Efficient Data Labeling by Hierarchical Crowdsourcing with Large Language Models , booktitle =

    Zhang, Haodi and Yang, Junyu and Nie, Jinyin and Liang, Peirou and Wu, Kaishun and Lian, Defu and Mao, Rui and Song, Yuanfeng , year =. Efficient Data Labeling by Hierarchical Crowdsourcing with Large Language Models , booktitle =

  10. [18]

    Fine-Tuned'Small'

    Martin Juan José Bucher and Marco Martini , year =. Fine-Tuned'Small'

  11. [19]

    Keeping humans in the loop: human-centered automated annotation with generative

    Pangakis, Nick and Wolken, Sam , year =. Keeping humans in the loop: human-centered automated annotation with generative

  12. [20]

    2024 , title =

    Nasution, Arbi Haza and Onan, Aytu. 2024 , title =

  13. [21]

    Felkner, Virginia K and Thompson, Jennifer A and May, Jonathan , year =

  14. [22]

    Karimi, Akbar and Rossi, Leonardo and Prati, Andrea , year =

  15. [23]

    Evaluating large language models for health-related text classification tasks with public social media data , journal =

    Guo, Yuting and Ovadje, Anthony and Al-Garadi, Mohammed Ali and Sarker, Abeed , year =. Evaluating large language models for health-related text classification tasks with public social media data , journal =

  16. [24]

    Exploring Persuasive Engagement to Reduce Over-Reliance on

    Raees, Muhammad and Khan, Vassilis-Javed and Papangelis, Konstantinos , year =. Exploring Persuasive Engagement to Reduce Over-Reliance on

  17. [25]

    Gilardi, Fabrizio and Alizadeh, Meysam and Kubli, Maël , year =

  18. [26]

    The sensemaking process and leverage points for analyst technology as identified through cognitive task analysis , booktitle =

    Pirolli, Peter and Card, Stuart , year =. The sensemaking process and leverage points for analyst technology as identified through cognitive task analysis , booktitle =

  19. [27]

    Zhu, Yiming and Yin, Zhizhuo and Tyson, Gareth and Haq, Ehsan-Ul and Lee, Lik-Hang and Hui, Pan , year =

  20. [28]

    , year =

    Strobelt, Hendrik and Webson, Albert and Sanh, Victor and Hoover, Benjamin and Beyer, Johanna and Pfister, Hanspeter and Rush, Alexander M. , year =

  21. [29]

    Chainforge: A visual toolkit for prompt engineering and llm hypothesis testing , booktitle =

    Arawjo, Ian and Swoopes, Chelse and Vaithilingam, Priyan and Wattenberg, Martin and Glassman, Elena L , year =. Chainforge: A visual toolkit for prompt engineering and llm hypothesis testing , booktitle =

  22. [30]

    Evaluation of an llm in identifying logical fallacies: A call for rigor when adopting llms in hci research , booktitle =

    Lim, Gionnieve and Perrault, Simon T , year =. Evaluation of an llm in identifying logical fallacies: A call for rigor when adopting llms in hci research , booktitle =

  23. [31]

    2019 , title =

    Pearce, Warren and Niederer, Sabine and. 2019 , title =

  24. [32]

    How Climate Movement Actors and News Media Frame Climate Change and Strike: Evidence from Analyzing Twitter and News Media Discourse from 2018 to 2021 , journal =

    Chen, Kaiping and Molder, Amanda L and Duan, Zening and Boulianne, Shelley and Eckart, Christopher and Mallari, Prince and Yang, Diyi , year =. How Climate Movement Actors and News Media Frame Climate Change and Strike: Evidence from Analyzing Twitter and News Media Discourse ...

  25. [33]

    Topic modelling and sentiment analysis of global warming tweets: evidence from big data analysis , journal =

    Qiao, Fang and Williams, Jago , year =. Topic modelling and sentiment analysis of global warming tweets: evidence from big data analysis , journal =

  26. [34]

    Topic modeling and sentiment analysis of global climate change tweets , journal =

    Dahal, Biraj and Kumar, Sathish AP and Li, Zhenlong , year =. Topic modeling and sentiment analysis of global climate change tweets , journal =

  27. [35]

    Tian, Katherine and Mitchell, Eric and Zhou, Allan and Sharma, Archit and Rafailov, Rafael and Yao, Huaxiu and Finn, Chelsea and Manning, Christopher , year =. Just

  28. [36]

    Large Language Models Are Human-Level Prompt Engineers , howpublished =

    Yongchao Zhou and Andrei Ioan Muresanu and Ziwen Han and Keiran Paster and Silviu Pitis and Harris Chan and Jimmy Ba , year =. Large Language Models Are Human-Level Prompt Engineers , howpublished =

  29. [37]

    Human and LLM Biases in Hate Speech Annotations: A Socio-Demographic Analysis of Annotators and Targets , booktitle =

    Giorgi, Tommaso and Cima, Lorenzo and Fagni, Tiziano and Avvenuti, Marco and Cresci, Stefano , year =. Human and LLM Biases in Hate Speech Annotations: A Socio-Demographic Analysis of Annotators and Targets , booktitle =

  30. [38]

    Large Language Model Annotation Bias in Hate Speech Detection , booktitle =

    Okpala, Ebuka and Cheng, Long , year =. Large Language Model Annotation Bias in Hate Speech Detection , booktitle =

  31. [39]

    Climate catastrophe: The value of envisioning the worst‐case scenarios of climate change , journal =

    Davidson, Joe and Kemp, Luke , year =. Climate catastrophe: The value of envisioning the worst‐case scenarios of climate change , journal =

  32. [40]

    What's in a Prompt?: A Large-Scale Experiment to Assess the Impact of Prompt Design on the Compliance and Accuracy of LLM-Generated Text Annotations , booktitle =

    Atreja, Shubham and Ashkinaze, Joshua and Li, Lingyao and Mendelsohn, Julia and Hemphill, Libby , year =. What's in a Prompt?: A Large-Scale Experiment to Assess the Impact of Prompt Design on the Compliance and Accuracy of LLM-Generated Text Annotations , booktitle =

  33. [41]

    , title =

    Lim, Gionnieve and Perrault, Simon T. , title =. 2024 , booktitle =

  34. [42]

    Arawjo, I.; Swoopes, C.; Vaithilingam, P.; Wattenberg, M.; and Glassman, E. L. 2024. Chainforge: A visual toolkit for prompt engineering and llm hypothesis testing. In Proceedings of the ACM Conference on Human Factors in Computing Systems (CHI)

  35. [43]

    Atreja, S.; Ashkinaze, J.; Li, L.; Mendelsohn, J.; and Hemphill, L. 2025. What's in a Prompt?: A Large-Scale Experiment to Assess the Impact of Prompt Design on the Compliance and Accuracy of LLM-Generated Text Annotations. In Proceedings of the AAAI Conference on Web and Soci...

  36. [44]

    B \"o hm, G.; and Pfister, H.-R. 2025. Exploring climate change discourses on the internet: a topic modeling study across ten years. Journal of risk research

  37. [45]

    Bucher, M. J. J.; and Martini, M. 2024. Fine-Tuned'Small' LLM s (Still) Significantly Outperform Zero-Shot Generative AI Models in Text Classification. arXiv:2406.08660

  38. [46]

    Chae, Y.; and Davidson, T. R. 2025. Large Language Models for Text Classification: From Zero-Shot Learning to Instruction-Tuning. Sociological Methods & Research

  39. [47]

    L.; Duan, Z.; Boulianne, S.; Eckart, C.; Mallari, P.; and Yang, D

    Chen, K.; Molder, A. L.; Duan, Z.; Boulianne, S.; Eckart, C.; Mallari, P.; and Yang, D. 2023. How Climate Movement Actors and News Media Frame Climate Change and Strike: Evidence from Analyzing Twitter and News Media Discourse from 2018 to 2021. The International Journal of Pr...

  40. [48]

    A.; and Li, Z

    Dahal, B.; Kumar, S. A.; and Li, Z. 2019. Topic modeling and sentiment analysis of global climate change tweets. Social network analysis and mining

  41. [49]

    Davidson, J.; and Kemp, L. 2023. Climate catastrophe: The value of envisioning the worst‐case scenarios of climate change. Wiley Interdisciplinary Reviews: Climate Change

  42. [50]

    D\" o rk, M.; Carpendale, S.; and Williamson, C. 2011. The information flaneur: a fresh look at information seeking. In Proceedings of the ACM Conference on Human Factors in Computing Systems (CHI)

  43. [51]

    K.; Thompson, J

    Felkner, V. K.; Thompson, J. A.; and May, J. 2024. GPT is Not an Annotator: The Necessity of Human Annotation in Fairness Benchmark Construction. arXiv:2405.15760

  44. [52]

    Galdeman, A.; and Aiello, L. M. 2025. Mapping the Climate Change Landscape on TikTok . In Proceedings of the AAAI Conference on Web and Social Media (ICWSM)

  45. [53]

    Gilardi, F.; Alizadeh, M.; and Kubli, M. 2023. ChatGPT Outperforms Crowd - Workers for Text - Annotation Tasks . Proceedings of the National Academy of Sciences (PNAS)

  46. [54]

    Giorgi, T.; Cima, L.; Fagni, T.; Avvenuti, M.; and Cresci, S. 2025. Human and LLM Biases in Hate Speech Annotations: A Socio-Demographic Analysis of Annotators and Targets. In Proceedings of the AAAI Conference on Web and Social Media (ICWSM)

  47. [55]

    A.; and Sarker, A

    Guo, Y.; Ovadje, A.; Al-Garadi, M. A.; and Sarker, A. 2024. Evaluating large language models for health-related text classification tasks with public social media data. Journal of the American Medical Informatics Association

  48. [56]

    Hasell, A.; Halversen, A.; and Weeks, B. E. 2025. When social media attack: How exposure to political attacks on social media promotes anger and political cynicism. The International Journal of Press/Politics

  49. [57]

    Hautea, S.; Parks, P.; Takahashi, B.; and Zeng, J. 2021. Showing They Care (or Don’t): Affective Publics and Ambivalent Climate Activism on TikTok . Social Media + Society

  50. [58]

    C.; Rohatgi, S.; and Huang, T.-H

    He, Z.; Huang, C.-Y.; Ding, C.-K. C.; Rohatgi, S.; and Huang, T.-H. K. 2024. If in a crowdsourced data annotation pipeline, a gpt-4. In Proceedings of the ACM Conference on Human Factors in Computing Systems (CHI)

  51. [59]

    Karimi, A.; Rossi, L.; and Prati, A. 2021. AEDA : An Easier Data Augmentation Technique for Text Classification. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP)

  52. [60]

    S.; Molina-Rogers, N.; and Gagr c in, E

    Lane, D. S.; Molina-Rogers, N.; and Gagr c in, E. 2025. Worn out & tuned out: does politics fatigue on social media foster participatory inequality among Americans? Mass Communication and Society

  53. [61]

    F.; Liu, Z.; and Yang, D

    Li, M.; Shi, T.; Ziems, C.; Kan, M.-Y.; Chen, N. F.; Liu, Z.; and Yang, D. 2023. Coannotating: Uncertainty-guided work allocation between human and large language models for data annotation. arXiv:2310.15638

  54. [62]

    Lim, G.; and Perrault, S. T. 2024. Evaluation of an LLM in Identifying Logical Fallacies: A Call for Rigor When Adopting LLMs in HCI Research. In Companion Publication of the Conference on Computer-Supported Cooperative Work and Social Computing (CSCW)

  55. [63]

    Ma, Q.; Peng, W.; Yang, C.; Shen, H.; Koedinger, K.; and Wu, T. 2025. What should we engineer in prompts? training humans in requirement-driven llm use. ACM Transactions on Computer-Human Interaction

  56. [64]

    H.; and Onan, A

    Nasution, A. H.; and Onan, A. 2024. Chatgpt label: Comparing the quality of human-generated and llm-generated annotations in low-resource language NLP tasks. IEEE Access

  57. [65]

    Okpala, E.; and Cheng, L. 2025. Large Language Model Annotation Bias in Hate Speech Detection. In Proceedings of the AAAI Conference on Web and Social Media (ICWSM)

  58. [66]

    OpenAI . 2024. GPT-4o mini (model documentation). https://platform.openai.com/docs/models/gpt-4o-mini

  59. [67]

    OpenAI . 2025. OpenAI API Reference. https://platform.openai.com/docs/api-reference/introduction

  60. [68]

    Pangakis, N.; and Wolken, S. 2025. Keeping humans in the loop: human-centered automated annotation with generative AI . In Proceedings of the AAAI Conference on Web and Social Media (ICWSM)

  61. [69]

    M.; and S \'a nchez Querub \' n, N

    Pearce, W.; Niederer, S.; \"O zkula, S. M.; and S \'a nchez Querub \' n, N. 2019. The social media life of climate change: Platforms, publics, and future imaginaries. Wiley interdisciplinary reviews: Climate change

  62. [70]

    Pirolli, P.; and Card, S. 2005. The sensemaking process and leverage points for analyst technology as identified through cognitive task analysis. In Proceedings of the Conference on Intelligence Analysis

  63. [71]

    Qiao, F.; and Williams, J. 2022. Topic modelling and sentiment analysis of global warming tweets: evidence from big data analysis. Journal of Organizational and End User Computing (JOEUC)

  64. [72]

    Raees, M.; Khan, V.-J.; and Papangelis, K. 2025. Exploring Persuasive Engagement to Reduce Over-Reliance on AI -Assistance in a Customer Classification Case. In Proceedings of the Conference on User Modeling, Adaptation and Personalization (UMAP)

  65. [73]

    Ranade, N.; Saravia, M.; and Johri, A. 2025. Using rhetorical strategies to design prompts: a human-in-the-loop approach to make AI useful. AI & SOCIETY

  66. [74]

    Rudin, C.; Chen, C.; Chen, Z.; Huang, H.; Semenova, L.; and Zhong, C. 2021. Interpretable machine learning: fundamental principles and 10 grand challenges. Statistics Surveys

  67. [75]

    A.; Rey, B

    Shin, J.; Hedderich, M. A.; Rey, B. J.; Lucero, A.; and Oulasvirta, A. 2024. Understanding human-AI workflows for generating personas. In Proceedings of the ACM Designing Interactive Systems Conference (DIS)

  68. [76]

    P.; Galley, M.; Caruana, R.; and Gao, J

    Singh, C.; Inala, J. P.; Galley, M.; Caruana, R.; and Gao, J. 2024. Rethinking Interpretability in the Era of Large Language Models. arXiv:2402.01761

  69. [77]

    news finds me

    Song, H.; de Z \'u \ n iga, H. G.; and Boomgaarden, H. G. 2021. Social media news use and political cynicism: Differential pathways through “news finds me” perception. In Social media news and its impact. Routledge

  70. [78]

    Strobelt, H.; Webson, A.; Sanh, V.; Hoover, B.; Beyer, J.; Pfister, H.; and Rush, A. M. 2023. Interactive and Visual Prompt Engineering for Ad-hoc Task Adaptation with Large Language Models . IEEE Transactions on Visualization & Computer Graphics

  71. [79]

    Tian, K.; Mitchell, E.; Zhou, A.; Sharma, A.; Rafailov, R.; Yao, H.; Finn, C.; and Manning, C. 2023. Just Ask for Calibration : Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine - Tuned with Human Feedback . In Proceedings of the Conference on Emp...

  72. [80]

    Treen, K.; Williams, H.; O ' Neill, S.; and Coan, T. G. 2022. Discussion of Climate Change on Reddit : Polarized Discourse or Deliberative Debate? Environmental Communication

  73. [81]

    Wang, X.; Kim, H.; Rahman, S.; Mitra, K.; and Miao, Z. 2024. Human-LLM collaborative annotation through effective verification of LLM labels. In Proceedings of the ACM Conference on Human Factors in Computing Systems (CHI)

  74. [82]

    D.; Wong, R

    Zamfirescu-Pereira, J. D.; Wong, R. Y.; Hartmann, B.; and Yang, Q. 2023. Why Johnny can’t prompt: how non-AI experts try (and fail) to design LLM prompts. In Proceedings of the ACM Conference on Human Factors in Computing Systems (CHI)

  75. [83]

    Zhang, H.; Yang, J.; Nie, J.; Liang, P.; Wu, K.; Lian, D.; Mao, R.; and Song, Y. 2025. Efficient Data Labeling by Hierarchical Crowdsourcing with Large Language Models. In Proceedings of the Conference on Computational Linguistics (COLING)

  76. [84]

    I.; Han, Z.; Paster, K.; Pitis, S.; Chan, H.; and Ba, J

    Zhou, Y.; Muresanu, A. I.; Han, Z.; Paster, K.; Pitis, S.; Chan, H.; and Ba, J. 2023. Large Language Models Are Human-Level Prompt Engineers. arXiv:2211.01910

  77. [85]

    Zhu, Y.; Yin, Z.; Tyson, G.; Haq, E.-U.; Lee, L.-H.; and Hui, P. 2024. APT - Pipe : A Prompt - Tuning Tool for Social Data Annotation using ChatGPT . In Proceedings of the ACM Web Conference (WWW)

Pith tools

Reviewed June 27, 2026 · model on record in the stance chip above.