REVIEW 3 major objections 6 minor 64 references
Can Third-parties Read Our Emotions?
T0 review · 3 major / 6 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read Third-party emotion labels—human or LLM—reach only low to fair agreement with authors' self-reported emotions.
desk verdict Solid empirical study showing third-party emotion labels diverge sharply from author self-reports; the 'private state' interpretation outruns the gold-standard validity, but the core finding and LLM-vs-human comparison deserve peer review. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is a paired first-party/third-party annotation study: social media users supplied their own posts and labeled them with the 27-emotion-plus-neutral taxonomy, then three demographically matched humans, three unmatched humans, and five LLMs labeled the same screenshots. Alignment is measured by Cohen's $\kappa$ and by $F_1$, recall, and precision, with majority voting to aggregate each group's labels. The design holds the text constant and varies only who labels it, so any measured gap is attributed to the annotator's position as an outsider rather than to the stimulus.
What would settle it
One decisive check would be to recompute the alignment after removing the 9.2% of first-party labels that the authors' own quality review flagged as spurious; if Cohen's $\kappa$ and $F_1$ then rose from low/fair to substantial (say $\kappa > 0.6$), the central claim would weaken into a self-report noise artifact, but if the scores remained in the same low range, the claim would stand.
Extended reading notes
Core claim
On its own terms, the paper's central discovery is that self-reported emotions and third-party judgments are systematically misaligned, and the gap is not limited to fine-grained label confusion. When the 28 emotion categories are collapsed into seven broad groups, human and LLM annotators still reach only low to moderate agreement (macro $F_1$ around 0.45–0.54), and at the fine-grained level LLM $F_1$ scores range from 0.2 to 0.6 while human scores range from 0.1 to 0.5. Treating the author's self-report as ground truth, the best third-party annotators—the LLMs—are better at saying what a text expresses than at recovering what the author says they felt. The claim that follows is that emotion-recognition models trained or evaluated on third-party labels are modeling a reader's perception of emotion, not the author's private state.
Load-bearing premise
The load-bearing assumption is that an author's self-reported emotion label faithfully captures their internal emotional state; the paper itself concedes that this cannot be externally verified and that 9.2% of first-party labels were flagged as spurious yet kept in the analysis.
Editorial extensions
If this is right
- Emotion recognition systems trained on third-party labels will tend to encode a reader's guess rather than the author's state, and are likely to misclassify the same people in high-stakes applications.
- LLM annotators' higher agreement with first-party labels does not make them a faithful replacement for first-party data, since their agreement is still only fair.
- Demographic matching of human annotators improves performance but leaves alignment far from reliable, so matching alone will not solve the private-state annotation problem.
- Because misalignment persists on coarse emotion groups, the failure is not merely confusion between near-synonymous labels; third parties pick substantially different emotion categories.
- Datasets and benchmarks that use third-party labels for emotion should be treated as measuring perceived emotion rather than felt emotion.
Reading between the lines
- An extension the authors do not draw: if the same first-party/third-party gap holds for other private states such as stance or opinion, many existing NLP benchmarks for those tasks may be measuring reader perception rather than author belief; testing that transfer is a natural next study.
- A testable extension would compare first-party labels collected immediately after posting with labels collected retrospectively; if the delayed labels agree more closely with third parties, part of the measured gap is memory or self-report timing, not inability to infer emotion.
- Because the retained spurious labels suggest a conservative estimate, a follow-up that reports alignment both with and without those labels would clarify how much of the gap is annotation error versus an unobservable private state.
- In practice, the results imply that high-stakes systems should treat emotion predictions as uncertain inferences and should surface that uncertainty or ask the user directly; this is a design implication, not a result the paper proves.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a set of human-subject experiments comparing first-party (author-provided) emotion labels on original social media posts with third-party annotations from crowd workers and five large language models. The authors measure agreement using Cohen's kappa, F1, recall, and precision, and report that all third-party annotators achieve only low-to-fair alignment with first-party labels (kappa 0 to 0.45), with LLMs outperforming human annotators on most emotions. They further test whether demographic similarity between human annotators and authors improves alignment and whether adding demographic information to LLM prompts helps. The central claim is that third-party annotations—human or machine—fail to faithfully represent authors' private states, and the paper proposes a framework for evaluating such limitations.
Significance. If valid, the result challenges a widespread assumption in subjective NLP that third-party labels are an acceptable gold standard for private states, with direct implications for emotion recognition, sentiment analysis, and related high-stakes applications. The study's strengths include a direct first-party-versus-third-party design, a reasonably large author-recruited corpus (729 posts from 123 participants), multiple LLMs and human annotators, a detailed emotion taxonomy, and both quantitative and qualitative analyses. The paper is honest about many limitations, including the unverifiability of first-party reports. However, the interpretive claim that low agreement demonstrates failure to model 'authors' private states' depends crucially on treating first-party self-reports as a valid ground truth, and the paper does not provide a sensitivity analysis to show that noise in that criterion does not drive the reported low agreement.
major comments (3)
- [Section 7 and Appendix D] The gold-standard assumption is load-bearing for the paper's central claim, and the manuscript provides no sensitivity analysis to support it. Appendix D states that 9.2% of first-party labels were flagged as spurious by two independent reviewers yet retained, and Section 5 reports that most posts with spurious first-party labels appear in the low-alignment subset. This pattern is exactly what would be observed if criterion noise, rather than third-party limitation alone, were suppressing the measured kappa and F1. The claim in Section 7 that retention 'is expected to have minimal impact on the overall findings' needs quantitative support: the analysis should be rerun after excluding flagged posts, or with a label-noise model, to show that the main conclusions survive.
- [Section 4.3 and Table 1] The text overstates the significance of the in-group advantage. It says the Wilcoxon tests show that differences 'measured based on F1, recall, precision, and Cohen's kappa, are statistically significant,' but Table 1 reports precision p = 0.050 without an asterisk and Table 2 reports annotator-level precision p = 0.052, which is not significant at the 0.05 level. Since RQ2 is central to the demographic-similarity claim, the reporting should be corrected to distinguish recall/F1 benefits from the nonsignificant precision difference.
- [Sections 4.3, 4.4, and 4.5] Multiple Wilcoxon signed-rank tests are reported without any multiple-comparison correction. This matters most for the marginal results: RQ2 precision (p = 0.050), annotator-level precision (p = 0.052), and RQ3 F1 (p = 0.0095) and precision (p = 0.0004) could be affected by correction, even though the very small p-values in Table 3 are robust. The authors should either apply a correction such as Benjamini–Hochberg and report which conclusions survive, or explicitly justify the uncorrected interpretation.
minor comments (6)
- [Section 3.2] The first-party instruction asked participants to select emotions 'they believed were expressed in the post,' which is itself an interpretive judgment rather than direct access to felt emotion. This wording should be acknowledged more prominently in the limitations, since it narrows the gap between first- and third-party interpretations.
- [Section 4.3] The sentence listing metrics says 'measured by F1, recall, precision, and recall,' duplicating recall; one of these should be Cohen's kappa.
- [Section 4.3] The word 'persevered' in 'we persevered individual third-party annotations' should be 'preserved.'
- [Table 3] The table header mixes median and mean labels: the left panel is labeled 'In-group Median' but the right panel says 'Out-group Mean,' and both columns contain a statistic and a p-value. The presentation should be made consistent and clearly defined.
- [Section 5] The F1 thresholds for high- and low-alignment cases (0.6 and 0.2) are described as 'determined based on the distribution of F1-scores,' which is post hoc. A brief justification or a statement that the qualitative patterns are robust to threshold choice would strengthen the analysis.
- [Section 4.2 and Figure 1] The main F1 results are presented only in a figure at the first occurrence, with exact numeric macro-averages given in the text. Adding a table with per-emotion F1, precision, and recall for in-group, out-group, and LLMs would improve reproducibility and readability.
Circularity Check
No significant circularity: the central comparison is a direct empirical measurement of third-party annotations against independent first-party self-reports, with no fitted parameters, no predicted quantities, and no load-bearing self-citation.
full rationale
The paper's derivation chain is a direct empirical comparison rather than a circular reduction. First-party emotion labels were collected from authors, and third-party annotations were collected independently from human annotators and five LLMs; alignment was then measured with Cohen's kappa, F1, recall, and precision, with no parameters fitted to the third-party data in order to predict the first-party labels. The central claim that third-party annotations misalign with first-party labels is literally the measured outcome, not a consequence of defining one quantity in terms of another. The broader claim that this misalignment shows failure to represent authors' private states does rely on treating first-party self-reports as the operational gold standard, and the paper explicitly acknowledges the limits of that assumption in Section 7 ('whether first-party labels faithfully represent first-party internal emotion can not be externally verifiable') and Appendix D (9.2% of first-party labels were flagged as spurious yet retained). This is a validity or robustness limitation, not a circularity: the first-party labels are independent evidence obtained from authors, and the paper does not define 'faithful representation' as mere agreement with its own fitted values. The only self-citation is Venkit et al. (2023) in Section 1 as related work on sentiment analysis concerns, and it is not load-bearing for any conclusion. No equations are defined in terms of one another, no fitted quantity is renamed as a prediction, no uniqueness theorem is imported from the authors' prior work, and no ansatz is smuggled in via citation. The derivation is therefore self-contained with respect to circularity.
Assumptions & free parameters
free parameters (1)
- high/low alignment F1 thresholds =
F1 > 0.6 for high, F1 < 0.2 for low
assumptions (4)
- domain assumption First-party self-reports are valid gold standards for authors' private states.
- domain assumption The 28-category taxonomy and its written definitions convey the same emotion concepts to authors, crowd annotators, and LLMs.
- domain assumption Demographic similarity in age, gender, and race is a sufficient proxy for the shared cultural and interpretive context that improves emotion alignment.
- domain assumption Majority voting across annotators or LLMs yields a dependable group annotation.
Cite this review
Pith. "Pith review of Can Third-parties Read Our Emotions?." pith.science (2026). https://pith.science/paper/YCIL6LNV
@misc{pith2026250418673,
author = {Pith},
title = {Pith review of: Can Third-parties Read Our Emotions?},
year = {2026},
howpublished = {\url{https://pith.science/paper/YCIL6LNV}},
note = {Machine review of arXiv:2504.18673}
}
read the original abstract
Natural Language Processing tasks that aim to infer an author's private states, e.g., emotions and opinions, from their written text, typically rely on datasets annotated by third-party annotators. However, the assumption that third-party annotators can accurately capture authors' private states remains largely unexamined. In this study, we present human subjects experiments on emotion recognition tasks that directly compare third-party annotations with first-party (author-provided) emotion labels. Our findings reveal significant limitations in third-party annotations-whether provided by human annotators or large language models (LLMs)-in faithfully representing authors' private states. However, LLMs outperform human annotators nearly across the board. We further explore methods to improve third-party annotation quality. We find that demographic similarity between first-party authors and third-party human annotators enhances annotation performance. While incorporating first-party demographic information into prompts leads to a marginal but statistically significant improvement in LLMs' performance. We introduce a framework for evaluating the limitations of third-party annotations and call for refined annotation practices to accurately represent and model authors' private states.
Figures
Figures from the paper (13 more)
Reference graph
Works this paper leans on
-
[1]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Abeer AlDayel and Walid Magdy. 2021. Stance detection on social media: State of the art and trends. Information Processing & Management, 58(4):102597
2021
-
[4]
Nourah Alswaidan and Mohamed El Bachir Menai. 2020. A survey of state-of-the-art approaches for emotion recognition in text. Knowledge and Information Systems, 62(8):2937--2987
work page 2020
-
[5]
Saima Aman and Stan Szpakowicz. 2007. Identifying expressions of emotion in text. In International Conference on Text, Speech and Dialogue, pages 196--205. Springer
work page 2007
-
[6]
Anthropic. 2024. The claude 3 model family: Opus, sonnet, haiku. Technical report, Anthropic
work page 2024
-
[7]
Alexandra Balahur, Jes \'u s M Hermida, and Andr \'e s Montoyo. 2012. Detecting implicit expressions of emotion in text: A comparative analysis. Decision support systems, 53(4):742--753
work page 2012
-
[8]
David Bamman and Noah Smith. 2015. Contextualized sarcasm detection on twitter. In proceedings of the international AAAI conference on web and social media, volume 9, pages 574--577
work page 2015
Show all 64 references
-
[9]
Ann Banfield. 1982. https://api.semanticscholar.org/CorpusID:62165816 Unspeakable sentences : Narration and representation in the language of fiction
1982
-
[10]
Lisa Feldman Barrett. 2004. Feelings or words? understanding the content in self-report ratings of experienced emotion. Journal of personality and social psychology, 87(2):266
2004
-
[11]
Patricia Bauer, Leif Stennes, and Jennifer Haight. 2003. Representation of the inner self in autobiography: Women's and men's use of internal states language in personal narratives. Memory, 11(1):27--42
2003
-
[12]
Tilman Beck, Hendrik Schuff, Anne Lauscher, and Iryna Gurevych. 2024. Sensitivity, performance, robustness: Deconstructing the effect of sociodemographic prompting. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (...
2024
-
[13]
Eric Brill. 1994. Some advances in transformation-based part of speech tagging. arXiv preprint cmp-lg/9406010
1994 arXiv
-
[14]
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877--1901
2020
-
[15]
Sven Buechel and Udo Hahn. 2017. Readers vs. writers vs. texts: Coping with different perspectives of text understanding in emotion annotation. In Proceedings of the 11th linguistic annotation workshop, pages 1--12
2017
-
[16]
Patricia Chiril, Endang Wahyu Pamungkas, Farah Benamara, V \'e ronique Moriceau, and Viviana Patti. 2022. Emotionally informed hate speech detection: a multi-target perspective. Cognitive Computation, pages 1--31
2022
-
[17]
Thomas Davidson, Dana Warmsley, Michael Macy, and Ingmar Weber. 2017. Automated hate speech detection and the problem of offensive language. In Proceedings of the international AAAI conference on web and social media, volume 11, pages 512--515
2017
-
[18]
David C DeAndrea, Allison S Shaw, and Timothy R Levine. 2010. Online language: The role of culture in self-expression and self-construal on facebook. Journal of language and social psychology, 29(4):425--442
2010
-
[19]
Dorottya Demszky, Dana Movshovitz-Attias, Jeongwoo Ko, Alan Cowen, Gaurav Nemade, and Sujith Ravi. 2020. https://doi.org/10.18653/v1/2020.acl-main.372 G o E motions: A dataset of fine-grained emotions . In Proceedings of the 58th Annual Meeting of the Association for Computati...
2020 doi
-
[20]
Yi Ding, Jacob You, Tonja-Katrin Machulla, Jennifer Jacobs, Pradeep Sen, and Tobias H \"o llerer. 2022. Impact of annotator demographics on sentiment dataset labeling. Proceedings of the ACM on Human-Computer Interaction, 6(CSCW2):1--22
2022
-
[21]
Benjamin D Douglas, Patrick J Ewell, and Markus Brauer. 2023. Data quality in online human-subjects research: Comparisons between mturk, prolific, cloudresearch, qualtrics, and sona. Plos one, 18(3):e0279720
2023
-
[22]
Paul Ekman. 1992. An argument for basic emotions. Cognition & emotion, 6(3-4):169--200
1992
-
[23]
Hillary Anger Elfenbein and Nalini Ambady. 2002 a . Is there an in-group advantage in emotion recognition?
2002
-
[24]
Hillary Anger Elfenbein and Nalini Ambady. 2002 b . On the universality and cultural specificity of emotion recognition: a meta-analysis. Psychological bulletin, 128(2):203
2002
-
[25]
Peer Eyal, Rothschild David, Gordon Andrew, Evernden Zak, and Damer Ekaterina. 2021. Data quality of platforms and panels for online behavioral research. Behavior research methods, pages 1--20
2021
-
[26]
Maria Gendron, Debi Roberson, Jacoba Marietta van der Vyver, and Lisa Feldman Barrett. 2014. Perceptions of emotion from facial expressions are not culturally universal: evidence from a remote culture. Emotion, 14(2):251
2014
-
[27]
Fabrizio Gilardi, Meysam Alizadeh, and Ma \" e l Kubli. 2023. Chatgpt outperforms crowd-workers for text-annotation tasks. CoRR, abs/2303.15056
2023 arXiv
-
[28]
Maram Hasanain, Fatema Ahmad, and Firoj Alam. 2024. Large language models for propaganda span annotation. In EMNLP (Findings) , pages 14522--14532. Association for Computational Linguistics
2024
-
[29]
Dirk Hovy and Diyi Yang. 2021. https://doi.org/10.18653/v1/2021.naacl-main.49 The importance of modeling social factors of language: Theory and practice . In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Huma...
2021 doi
-
[30]
EunJeong Hwang, Bodhisattwa Majumder, and Niket Tandon. 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.393 Aligning language models to user opinions . In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 5906--5919, Singapore. Association for ...
2023 doi
-
[31]
Mohit Iyyer, Peter Enns, Jordan Boyd-Graber, and Philip Resnik. 2014. Political ideology detection using recursive neural networks. In Proceedings of the 52nd annual meeting of the Association for Computational Linguistics (volume 1: long papers), pages 1113--1122
2014
-
[32]
Kenneth Joseph, Sarah Shugars, Ryan Gallagher, Jon Green, Alexi Quintana Math \'e , Zijian An, and David Lazer. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.27 (mis)alignment between stance expressed in social media data and public opinion surveys . In Proceedings of the ...
2021 doi
-
[33]
Hannah Kim, Kushan Mitra, Rafael Li Chen, Sajjadur Rahman, and Dan Zhang. 2024. https://aclanthology.org/2024.eacl-demo.18/ MEGA nno+: A human- LLM collaborative annotation system . In Proceedings of the 18th Conference of the European Chapter of the Association for Computatio...
2024
-
[34]
Mao Li and Frederick Conrad. 2024. Advancing annotation of stance in social media posts: A comparative analysis of large language models and crowd sourcing. arXiv preprint arXiv:2406.07483
2024 arXiv
-
[35]
Minzhi Li, Taiwei Shi, Caleb Ziems, Min-Yen Kan, Nancy F Chen, Zhengyuan Liu, and Diyi Yang. 2023. Coannotating: Uncertainty-guided work allocation between human and large language models for data annotation. arXiv preprint arXiv:2310.15638
2023 arXiv
-
[36]
Yanran Li, Hui Su, Xiaoyu Shen, Wenjie Li, Ziqiang Cao, and Shuzi Niu. 2017. Dailydialog: A manually labelled multi-turn dialogue dataset. arXiv preprint arXiv:1710.03957
2017 arXiv
-
[37]
Nangyeon Lim. 2016. Cultural differences in emotion: differences in emotional arousal level between the east and the west. Integrative medicine research, 5(2):105--109
2016
-
[38]
Chen Liu, Muhammad Osama, and Anderson De Andrade. 2019. Dens: A dataset for multi-class emotion analysis. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP...
2019
-
[39]
Aire Mill, J \"u ri Allik, Anu Realo, and Raivo Valk. 2009. Age-related differences in emotion recognition ability: a cross-sectional study. Emotion, 9(5):619
2009
-
[40]
Saif Mohammad and Felipe Bravo-Marquez. 2017. https://doi.org/10.18653/v1/S17-1007 Emotion intensities in tweets . In Proceedings of the 6th Joint Conference on Lexical and Computational Semantics (* SEM 2017) , pages 65--77, Vancouver, Canada. Association for Computational Li...
2017 doi
-
[41]
Saif Mohammad, Felipe Bravo-Marquez, Mohammad Salameh, and Svetlana Kiritchenko. 2018. https://doi.org/10.18653/v1/S18-1001 S em E val-2018 task 1: Affect in tweets . In Proceedings of the 12th International Workshop on Semantic Evaluation, pages 1--17, New Orleans, Louisiana....
2018 doi
-
[42]
Sagnik Mukherjee, Muhammad Farid Adilazuarda, Sunayana Sitaram, Kalika Bali, Alham Fikri Aji, and Monojit Choudhury. 2024. https://doi.org/10.18653/v1/2024.emnlp-main.884 Cultural conditioning or placebo? on the effectiveness of socio-demographic prompting . In Proceedings of ...
2024 doi
-
[43]
David Nadeau and Satoshi Sekine. 2009. A survey of named entity recognition and classification. In Named Entities: Recognition, classification and use, pages 3--28. John Benjamins publishing company
2009
-
[44]
L \'a szl \'o Nemes and Attila Kiss. 2021. Social media sentiment analysis based on covid-19. Journal of Information and Telecommunication, 5(1):1--15
2021
-
[45]
Minxue Niu, Mimansa Jaiswal, and Emily Mower Provost. 2024. From text to emotion: Unveiling the emotion annotation capabilities of llms. CoRR, abs/2408.17026
2024 arXiv
-
[46]
OpenAI. 2023. GPT-4 technical report. CoRR, abs/2303.08774
2023 arXiv
-
[47]
Silviu Oprea and Walid Magdy. 2019. isarcasm: A dataset of intended sarcasm. arXiv preprint arXiv:1911.03123
2019 arXiv
-
[48]
John Pavlopoulos, Jeffrey Sorensen, Lucas Dixon, Nithum Thain, and Ion Androutsopoulos. 2020. Toxicity detection: Does context really matter? In ACL , pages 4296--4305. Association for Computational Linguistics
2020
-
[49]
Robert Plutchik. 2001. The nature of emotions: Human emotions have deep evolutionary roots, a fact that may explain their complexity and provide tools for clinical practice. American scientist, 89(4):344--350
2001
-
[50]
Randolph Quirk, Sidney Greenbaum, Geoffrey Leech, and Jan Svartvik. 1985. A Comprehensive Grammar of the English Language. Longman, New York
1985
-
[51]
Cyrus Rashtchian, Peter Young, Micah Hodosh, and Julia Hockenmaier. 2010. Collecting image annotations using amazon’s mechanical turk. In Proceedings of the NAACL HLT 2010 workshop on creating speech and language data with Amazon’s Mechanical Turk, pages 139--147
2010
-
[52]
Lillicrap, Jean - Baptiste Alayrac, Radu Soricut, Angeliki Lazaridou, Orhan Firat, Julian Schrittwieser, Ioannis Antonoglou, Rohan Anil, Sebastian Borgeaud, Andrew M
Machel Reid, Nikolay Savinov, Denis Teplyashin, Dmitry Lepikhin, Timothy P. Lillicrap, Jean - Baptiste Alayrac, Radu Soricut, Angeliki Lazaridou, Orhan Firat, Julian Schrittwieser, Ioannis Antonoglou, Rohan Anil, Sebastian Borgeaud, Andrew M. Dai, Katie Millican, Ethan Dyer, M...
2024 arXiv
-
[53]
Disa A Sauter, Frank Eisner, Paul Ekman, and Sophie K Scott. 2010. Cross-cultural recognition of basic emotions through nonverbal emotional vocalizations. Proceedings of the National Academy of Sciences, 107(6):2408--2412
2010
-
[54]
Phillip Shaver, Judith Schwartz, Donald Kirson, and Cary O'connor. 1987. Emotion knowledge: further exploration of a prototype approach. Journal of personality and social psychology, 52(6):1061
1987
-
[55]
Hannah Shaw and Minna Lyons. 2017. Lie detection accuracy—the role of age and the use of emotions as a reliable cue. Journal of Police and Criminal Psychology, 32:300--304
2017
-
[56]
Qinlan Shen and Carolyn Rose. 2021. https://doi.org/10.18653/v1/2021.eacl-main.152 What sounds right to me? experiential factors in the perception of political ideology . In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguis...
2021 doi
-
[57]
Rion Snow, Brendan O’connor, Dan Jurafsky, and Andrew Y Ng. 2008. Cheap and fast--but is it good? evaluating non-expert annotations for natural language tasks. In Proceedings of the 2008 conference on empirical methods in natural language processing, pages 254--263
2008
-
[58]
Huaman Sun, Jiaxin Pei, Minje Choi, and David Jurgens. 2023. Aligning with whom? large language models have gender and racial biases in subjective nlp tasks. arXiv preprint arXiv:2311.09730
2023 arXiv
-
[59]
Pranav Venkit, Mukund Srinath, Sanjana Gautam, Saranya Venkatraman, Vipul Gupta, Rebecca Passonneau, and Shomir Wilson. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.848 The sentiment problem: A critical survey towards deconstructing sentiment analysis . In Proceedings of ...
2023 doi
-
[60]
Sophie F Waterloo, Susanne E Baumgartner, Jochen Peter, and Patti M Valkenburg. 2018. Norms of online expressions of emotion: Comparing facebook, twitter, instagram, and whatsapp. New media & society, 20(5):1813--1831
2018
-
[61]
Janyce Wiebe, Theresa Wilson, Rebecca Bruce, Matthew Bell, and Melanie Martin. 2004. Learning subjective language. Computational linguistics, 30(3):277--308
2004
-
[62]
Janyce Wiebe, Theresa Wilson, and Claire Cardie. 2005. Annotating expressions of opinions and emotions in language. Language resources and evaluation, 39:165--210
2005
-
[63]
Tianlin Zhang, Kailai Yang, Shaoxiong Ji, and Sophia Ananiadou. 2023. Emotion fusion for mental illness detection from social media: A survey. Information Fusion, 92:231--246
2023
-
[64]
Artur Zygad o, Marek Koz owski, and Artur Janicki. 2021. Text-based emotion recognition in english and polish for therapeutic chatbot. Applied Sciences, 11(21):10146
2021
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.