REVIEW 3 major objections 5 minor 88 references
Representations of Fact, Fiction and Forecast in Large Language Models: Epistemics and Attitudes
T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Controlled tests show LLMs frequently fail to choose the correct epistemic modal or attitude verb, so their verbalized uncertainty is not always reliable.
desk verdict A useful, well-controlled behavioral study of LLMs' epistemic modal semantics, but the abstract's "generation" claim overreaches what the forced-choice experiments actually show. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is a typological map of epistemic meaning that separates evidence (direct vs. indirect) from commitment (a scale of full, partial, and neutral support). The paper operationalizes this map as two template families: hidden-object stories, adapted from child-language tasks, that vary whether evidence forces a unique conclusion (necessity, 'has to') or leaves alternatives open (possibility, 'may'), and theory-of-mind stories that contrast true facts, false beliefs, and uncertain beliefs, requiring 'know', 'believe', or 'doubt'. Each story is presented in three response formats (direct slot, indirect slot, indirect sentence) to test whether modal semantic knowledge is robust to surface format. The headline metric is forced-choice accuracy and paired accuracy (both members of a necessity/possibility or high/low-certainty pair correct), with logistic regressions estimating the effects of model size, story type, modal condition, and format.
What would settle it
Run the same story items (modal-auxiliary templates and attitude-verb statements) with adult native speakers; if human agreement with the paper's correct answers falls well below ceiling (for instance below 90%), especially on the counterfactual and belief items, then the accuracy figures are not measuring the appropriateness of epistemic usage and the central claim would need to be revisited.
Extended reading notes
Core claim
The paper's central discovery is that instruction-tuned LLMs have only partial and fragile semantic knowledge of epistemic modality. In Experiment 1, forced-choice selection between 'may/might' and 'must/have to' in stories that either rule out all but one location (necessity) or leave several locations open (possibility) yields accuracies from 51.2% to 95.3% across eight open-weight models, with small models near chance and even the best model missing simple necessity items and failing paired accuracy (both necessity and possibility correct in the same template) in a meaningful fraction of cases. In Experiment 2, selecting 'know', 'believe', or 'doubt' for fact and belief statements from theory-of-mind stories gives accuracies from 36.4% to 72.8%, with joint accuracy (all eight statements from one story correct) at or near zero. The paper reads these results as evidence that the linguistic knowledge required for generating expressions of uncertainty is not robustly encoded, and that this gap is a bottleneck for truthful verbalized uncertainty in LLMs.
Load-bearing premise
The paper assumes the experimenters' labels for which epistemic expression is 'correct' are the ground truth for appropriate use, but these labels are theory-based and were not validated by testing the same stories on adult human speakers.
Editorial extensions
If this is right
- If LLMs cannot reliably select the correct epistemic modal in simple controlled stories, then their verbalized confidence in open-ended generation is not a trustworthy signal without additional calibration or semantic enrichment.
- Scaling from small (7–8B) to medium (70–72B) parameters substantially improves modal selection, but the gap remains: even the largest models fail paired accuracy on a meaningful fraction of items and perform at or near chance on joint story accuracy.
- The systematic asymmetries (necessity better than possibility, facts better than beliefs, and high-certainty verbs better than 'doubt' in most models) identify specific weaknesses that targeted training data could address.
- Methods that prompt models to 'express uncertainty in words' inherit these gaps, so they should be complemented by work that enriches the semantic representation of epistemic modality itself.
Reading between the lines
- Because the experiments use forced-choice selection rather than open-ended generation, the paper's headline claim about 'generating' epistemic expressions is only partially supported; an open-generation test (e.g., measuring the distribution of modal auxiliaries produced in free continuation) could either confirm a generation-level deficit or reveal that models know more than selection tasks show.
- The correctness labels are theory-derived and not human-normed; if adult speakers disagreed with labels on counterfactual or belief items, the accuracy numbers would need reinterpretation. The same story templates could serve as a benchmark with human normative data as the gold standard.
- The paper tests only English; typological variation in how modalities are marked across languages (affixes, case, particles) means the conclusions about LLM knowledge may not transfer to multilingual settings, and the templates could be adapted cross-linguistically.
- A directly testable extension would use the story templates with logit-based confidence measures to see whether a model's internal uncertainty predicts its error patterns on 'may' versus 'has to', connecting the linguistic-knowledge gap to the calibration literature the paper critiques.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper uses controlled storyboards to probe whether eight open-weight instruction-tuned LLMs select semantically appropriate epistemic modal auxiliaries (may/might vs. must/have to, Experiment 1) and attitude verbs (know/believe/doubt, Experiment 2). Stimuli are built from a Boye-style typological taxonomy and from ToM templates; accuracy is measured in three forced-choice response formats (direct slot, indirect slot, indirect sentence). Logistic regression analyses show that 70-72B models outperform 7-8B models, necessity modals are selected more accurately than possibility modals, and fact statements are reported more accurately than belief statements in most model families. The authors conclude that LLMs' 'generating' epistemic expressions is limited and unreliable and that semantic enrichment of epistemic modality is needed.
Significance. If read as a study of forced-choice semantic selection, this is a solid, reproducible contribution: stimuli and code are released, model families and sizes are compared transparently, and the statistical analyses are described in detail (contrast coding, AIC-based selection, ROC curves). It also usefully brings typological semantic distinctions into LLM evaluation. The central reliability claim, however, is not warranted by the experimental design, and the validity of the gold labels remains the main threat to the accuracy numbers.
major comments (3)
- [Abstract; §7 Conclusion; §4.1 and §5.1 design] The headline claim that 'the performance of LLMs in generating epistemic expressions is limited and not robust' is not supported by the experiments. In both Experiment 1 and Experiment 2, every QA format is forced-choice: the model must choose between 'may' and 'has to' (or among 'know', 'believe', 'doubt'), or choose an index/sentence. No condition elicits an open-ended epistemic expression. Forced-choice accuracy is a metalinguistic selection measure; a model could fail it while generating appropriate modals in free text, or vice versa. The contribution list in §1 is carefully worded ('selecting appropriate epistemic expressions'), but the abstract and conclusion generalize to generation. Re-scope the claims to selection, or add an open-generation condition, before the broader reliability claim can be accepted.
- [§4.1; §5.1; Appendix E.1; Limitations] The gold labels come from the authors' theory-laden interpretation of Boye's semantic map and from the hidden-object paradigm, not from adult judgments on the actual stimuli. This matters most in the counterfactual condition: for the item in Table E.1 labeled as Necessity, the story states that Tom moved the peanut from the purple box to the black box, and the correct answer is given as 'has to' in the imagined counterfactual; but it is a substantive semantic assumption that the counterfactual world uniquely determines the peanut's location, and alternative reasonable judgments are possible. Since every accuracy number in Tables 1 and 2 is defined relative to these labels, the lack of same-stimulus adult norming is a load-bearing gap. The Limitations section acknowledges this, but the conclusion is still phrased as a general statement about LLM unreliability.
- [§4.2; §5.2; Appendix C] The 'number of parameters' effects are inferred from logistic regressions that treat trials as independent within each model family, with only two model instances per family. Any small-vs-medium difference is therefore confounded with model identity, tokenizer, training data, and alignment choices. The AIC-based procedure in Appendix C did not retain random effects, but that is not evidence against the confounding. The descriptive accuracy tables remain valid, but the causal-sounding claim that 'a large number of parameters leads to increased accuracy' should be softened to a correlational observation across two instances per family, or supported with additional checkpoints, seeds, or model families.
minor comments (5)
- [Appendix D.2, Figure 9] The caption of Figure 9 says 'modal condition' for Experiment 2, but the variable displayed is epistemic certainty (low vs. not low); the caption should be corrected for consistency with the text.
- [§6.1] There is a missing space in 'Epistemic modal verbsOzturk and Papafragou (2015)'; it should read 'Epistemic modal verbs. Ozturk and Papafragou (2015)' or similar.
- [References] The reference 'Htrotugu Akaike. 1973' should be 'Hirotugu Akaike. 1973'.
- [Figure 2] Figure 2 is dense and the annotations '①', '②', '③', etc. are not explained in the caption; a short legend would make the experimental flow much easier to follow.
- [§4.2 and §5.2] The text reports that 1-shot stories were expected to be easier than base stories, but the regression results show inconsistent and sometimes negative effects of story type; a brief discussion of this unexpected pattern would improve the paper.
Circularity Check
No significant circularity: the paper measures LLM forced-choice accuracy against externally defined semantic labels; no fitted parameter is renamed as a prediction, and no load-bearing self-citation chain appears.
full rationale
The paper's derivation chain is empirical rather than definitional. The authors define epistemic meaning dimensions using Boye's typological framework, construct controlled stories whose logical structure determines the target modal auxiliary or attitude verb, and then score LLM responses against those externally fixed labels. Model responses are never used to define the labels, so there is no self-definitional circularity: the 'may' versus 'has to' distinction in Experiment 1 follows from whether the story rules out all but one candidate or leaves multiple candidates open, and the 'know/believe/doubt' distinction in Experiment 2 follows from fact versus belief status and evidence strength. No parameter is fitted to a subset of data and then reported as a prediction; the logistic regressions are descriptive summaries of measured accuracy, not fitted predictors of the headline conclusion. There is also no load-bearing self-citation chain: the paper is by Meng Li, Michael Vrazitulis, and David Schlangen, and its citations to 'Li et al. 2024' refer to Kenneth Li et al.'s inference-time intervention work, not to the present authors. The child-development and adult-control comparisons come from external literature such as Ozturk and Papafragou (2015), which is independent evidence rather than self-citation. Two concerns raised in review are real but do not constitute circularity: the abstract and conclusion generalize from forced-choice selection to 'generating' epistemic expressions, which is an external-validity gap, and the correctness labels are theory-laden, which the Limitations section partially acknowledges by noting that human participants were not tested directly. Neither concern makes the reported accuracies equivalent to the paper's own inputs by construction. Therefore no specific circular step can be exhibited, and the appropriate score is 0.
Assumptions & free parameters
assumptions (2)
- domain assumption Forced-choice selection behavior reflects underlying linguistic knowledge of epistemic modality.
- domain assumption The experimenter-defined correct answers represent appropriate epistemic expressions for the story contexts.
Cite this review
Pith. "Pith review of Representations of Fact, Fiction and Forecast in Large Language Models: Epistemics and Attitudes." pith.science (2026). https://pith.science/paper/LZ2LF6HB
@misc{pith2026250601512,
author = {Pith},
title = {Pith review of: Representations of Fact, Fiction and Forecast in Large Language Models: Epistemics and Attitudes},
year = {2026},
howpublished = {\url{https://pith.science/paper/LZ2LF6HB}},
note = {Machine review of arXiv:2506.01512}
}
read the original abstract
Rational speakers are supposed to know what they know and what they do not know, and to generate expressions matching the strength of evidence. In contrast, it is still a challenge for current large language models to generate corresponding utterances based on the assessment of facts and confidence in an uncertain real-world environment. While it has recently become popular to estimate and calibrate confidence of LLMs with verbalized uncertainty, what is lacking is a careful examination of the linguistic knowledge of uncertainty encoded in the latent space of LLMs. In this paper, we draw on typological frameworks of epistemic expressions to evaluate LLMs' knowledge of epistemic modality, using controlled stories. Our experiments show that the performance of LLMs in generating epistemic expressions is limited and not robust, and hence the expressions of uncertainty generated by LLMs are not always reliable. To build uncertainty-aware LLMs, it is necessary to enrich semantic knowledge of epistemic modality in LLMs.
Figures
Figures from the paper (6 more)
Reference graph
Works this paper leans on
-
[1]
Alexandra Y Aikhenvald. 2004. Evidentiality. Oxford University Press
2004
-
[2]
AI@Meta. 2024. https://github.com/meta-llama/llama3/blob/main/MODEL_CARD.md Llama 3 model card
2024
-
[3]
Htrotugu Akaike. 1973. Maximum likelihood identification of gaussian autoregressive moving average models. Biometrika, 60(2):255--265
1973
-
[4]
Amos Azaria and Tom Mitchell. 2023. The internal state of an llm knows when it’s lying. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 967--976
2023
-
[5]
Misha Becker and Bruno Estigarribia. 2013. Harder words: Learning abstract verbs with opaque syntax. Language Learning and Development, 9(3):211--244
2013
-
[6]
Catarina G Bel \'e m, Markelle Kelly, Mark Steyvers, Sameer Singh, and Padhraic Smyth. 2024. https://doi.org/10.18653/v1/2024.emnlp-main.483 Perceptions of linguistic uncertainty by language models and humans . In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 8467--8502, Miami, Florida, USA. Association for ...
-
[7]
Paul Bloom. 2002. How children learn the meanings of words. MIT press
2002
-
[8]
Kasper Boye. 2012. Epistemic meaning: A crosslinguistic and functional-cognitive study, volume 43. Walter de Gruyter
2012
Show all 88 references
-
[9]
Collin Burns, Haotian Ye, Dan Klein, and Jacob Steinhardt. 2023. Discovering latent knowledge in language models without supervision. In The Eleventh International Conference on Learning Representations
2023
-
[10]
Strang Burton and Lisa Matthewson. 2015. Targeted construction storyboards in semantic fieldwork. In M. Ryan Bochnak and Lisa Matthewson, editors, Methodologies in Semantic Fieldwork, chapter 5, pages 135--156. Oxford University Press
2015
-
[11]
Joan Bybee, Revere Perkins, and William Pagliuca. 1994. The Evolution of Grammar: Tense, Aspect, and Modality in the Languages of the World. University of Chicago Press
1994
-
[12]
Joan L Bybee. 1985. Morphology: A study of the relation between meaning and form. John Benjamins Publishing Company
1985
-
[13]
Herbert H Clark and Catherine R Marshall. 1981. Definite knowledge and mutual knowledge. In Aravind K Joshi, Bonnie L Webber, and Ivan A Sag, editors, Elements of Discourse Understanding, pages 10--63. Cambridge University Press
1981
-
[14]
Jennifer Coates. 1983. The semantics of the modal auxiliaries. Croom Helm
1983
-
[15]
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. 2021. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168
2021 arXiv
-
[16]
Ferdinand de Haan. 2006. Typological approaches to modality. In Wolfgang Klein and Stephen Levinson, editors, The Expression of Modality, pages 27--69. Mouton de Gruyter
2006
-
[17]
Andreas Demetriou, Antigoni Mouyi, and George Spanoudis. 2010. The development of mental processing. In Richard M. Lerner., editor, The Handbook of Life-Span Development: Cognition, Biology, and Methods, page 36–55. Wiley
2010
-
[18]
Holger Diessel and Michael Tomasello. 2001. The acquisition of finite complement clauses in english: A corpus-based analysis. Cognitive Linguistics, 12(2)
2001
-
[19]
Jinhao Duan, Hao Cheng, Shiqi Wang, Alex Zavalny, Chenan Wang, Renjing Xu, Bhavya Kailkhura, and Kaidi Xu. 2024. Shifting attention to relevance: Towards the predictive uncertainty quantification of free-form large language models. In Proceedings of the 62nd Annual Meeting of ...
2024
-
[20]
Wade Fagen-Ulmschneider. 2019. https://waf.cs.illinois.edu/get-probability-words/ Perception of probability words. Accessed: Feb-01-2025
2019
-
[21]
Kanishk Gandhi, Jan-Philipp Fr \"a nken, Tobias Gerstenberg, and Noah Goodman. 2024. Understanding social reasoning in language models with language models. Advances in Neural Information Processing Systems, 36
2024
-
[22]
Jiahui Geng, Fengyu Cai, Yuxia Wang, Heinz Koeppl, Preslav Nakov, and Iryna Gurevych. 2024. A survey of confidence estimation and calibration in large language models. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Ling...
2024
-
[23]
Mor Geva, Daniel Khashabi, Elad Segal, Tushar Khot, Dan Roth, and Jonathan Berant. 2021. Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies. Transactions of the Association for Computational Linguistics, 9:346--361
2021
-
[24]
Anastasia Giannakidou and Alda Mari. 2021. Truth and veridicality in grammar and thought: Mood, modality, and propositional attitudes. University of Chicago Press
2021
-
[25]
Lila Gleitman. 1990. The structural sources of verb meanings. Language acquisition, 1(1):3--55
1990
-
[26]
Alison Gopnik and Janet W Astington. 1988. Children's understanding of representational change and its relation to the understanding of false belief and the appearance-reality distinction. Child development, pages 26--37
1988
-
[27]
Erin Grant, Aida Nematzadeh, and Thomas L Griffiths. 2017. How can memory-augmented neural networks pass a false-belief task? In 39th Annual Meeting of the Cognitive Science Society: Computational Foundations of Cognition, CogSci 2017, pages 427--432. The Cognitive Science Society
2017
-
[28]
Valentine Hacquard and Jeffrey Lidz. 2019. Children's attitude problems: Bootstrapping verb meaning from syntax and pragmatics. Mind & Language, 34(1):73--96
2019
-
[29]
Valentine Hacquard and Jeffrey Lidz. 2022. On the acquisition of attitude verbs. Annual Review of Linguistics, 8(1):193--212
2022
-
[30]
William Hirst and Joyce Weil. 1982. Acquisition of epistemic and deontic meaning of modals. Journal of child language, 9(3):659--666
1982
-
[31]
Wesley Holliday, Matthew Mandelkern, and Cedegao Zhang. 2024. Conditional and modal reasoning in large language models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 3800--3821
2024
-
[32]
Laurence R. Horn. 1989. A Natural History of Negation. University of Chicago Press
1989
-
[33]
Jennifer Hu and Roger Levy. 2023. Prompting is not a substitute for probability measurements in large language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 5040--5060
2023
-
[34]
Yuheng Huang, Jiayang Song, Zhijie Wang, Huaming Chen, and Lei Ma. 2023. Look before you leap: An exploratory study of uncertainty measurement for large language models. arXiv preprint arXiv:2307.10236
2023 arXiv
-
[35]
Michael Israel. 2008. Mental spaces and mental verbs in early child english. In Andrea Tyler, Yiyoung Kim, and Mari Takada, editors, Language in the context of use : discourse and cognitive approaches to language, pages 199--232. Mouton de Gruyter
2008
-
[36]
Zhengbao Jiang, Frank F Xu, Jun Araki, and Graham Neubig. 2020. How can we know what language models know? Transactions of the Association for Computational Linguistics, 8:423--438
2020
-
[37]
Emily Jin, Zhuoyi Huang, Jan-Philipp Fr \"a nken, Weiyu Liu, Hannah Cha, Erik Brockbank, Sarah A Wu, Ruohan Zhang, Jiajun Wu, and Tobias Gerstenberg. 2024. https://openreview.net/forum?id=nAFBHoMpQs MARPLE : A benchmark for long-horizon inference . In The Thirty-eight Conferen...
2024
-
[38]
Mandar Joshi, Eunsol Choi, Daniel S Weld, and Luke Zettlemoyer. 2017. Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p...
2017
-
[39]
Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, et al. 2022. Language models (mostly) know what they know. arXiv preprint arXiv:2207.05221
2022 arXiv
-
[40]
Mehran Kazemi, Quan Yuan, Deepti Bhatia, Najoung Kim, Xin Xu, Vaiva Imbrasaite, and Deepak Ramachandran. 2023. Boardgameqa: A dataset for natural language reasoning with contradictory information. Advances in Neural Information Processing Systems, 36:39052--39074
2023
-
[41]
Alex Kendall and Yarin Gal. 2017. What uncertainties do we need in bayesian deep learning for computer vision? Advances in neural information processing systems, 30
2017
-
[42]
Michal Kosinski. 2024. Evaluating large language models in theory of mind tasks. Proceedings of the National Academy of Sciences, 121(45):e2405460121
2024
-
[43]
Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar. 2023. Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation. In The Eleventh International Conference on Learning Representations
2023
-
[44]
Barbara Landau and Lila R Gleitman. 1985. Language and experience: Evidence from the blind child. Harvard University Press
1985
-
[45]
Matthew Le, Y-Lan Boureau, and Maximilian Nickel. 2019. Revisiting the evaluation of theory of mind through question answering. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Langu...
2019
-
[46]
Kenneth Li, Oam Patel, Fernanda Vi \'e gas, Hanspeter Pfister, and Martin Wattenberg. 2024. Inference-time intervention: Eliciting truthful answers from a language model. Advances in Neural Information Processing Systems, 36
2024
-
[47]
Stephanie Lin, Jacob Hilton, and Owain Evans. 2022. Teaching models to express their uncertainty in words. Transactions on Machine Learning Research
2022
-
[48]
Zhen Lin, Shubhendu Trivedi, and Jimeng Sun. 2024. Generating with confidence: Uncertainty quantification for black-box large language models. Transactions on Machine Learning Research
2024
-
[49]
Potsawee Manakul, Adian Liusie, and Mark Gales. 2023. Selfcheckgpt: Zero-resource black-box hallucination detection for generative large language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 9004--9017
2023
-
[50]
Sabrina J Mielke, Arthur Szlam, Emily Dinan, and Y-Lan Boureau. 2022. Reducing conversational agents’ overconfidence through linguistic calibration. Transactions of the Association for Computational Linguistics, 10:857--872
2022
-
[51]
Derek E Montgomery. 2002. Mental verbs and semantic development. Journal of Cognition and Development, 3(4):357--384
2002
-
[52]
Aida Nematzadeh, Kaylee Burns, Erin Grant, Alison Gopnik, and Tom Griffiths. 2018. Evaluating theory of mind in question answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2392--2400
2018
-
[53]
Ira A Noveck, Simon Ho, and Maria Sera. 1996. Children's understanding of epistemic modals. Journal of child language, 23(3):621--643
1996
-
[54]
Ozge Ozturk and Anna Papafragou. 2015. The acquisition of epistemic modality: From semantic meaning to pragmatic interpretation. Language learning and development, 11(3):191--214
2015
-
[55]
Frank Robert Palmer. 1986. Mood and modality. Cambridge University
1986
-
[56]
Anna Papafragou. 2002. Modality and theory of mind: Perspectives from language development and autism. In Modality and its Interaction with the Verbal System, pages 185--204. John Benjamins Publishing Company
2002
-
[57]
Anna Papafragou, Kimberly Cassidy, and Lila Gleitman. 2007. When we think about thinking: The acquisition of belief verbs. Cognition, 105(1):125--165
2007
-
[58]
Paul Portner. 2009. Modality. Oxford Surveys in Semantics and Pragmatics. Oxford University Press
2009
-
[59]
R Core Team . 2023. https://www.R-project.org/ R: A Language and Environment for Statistical Computing . R Foundation for Statistical Computing, Vienna, Austria
2023
-
[60]
David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R Bowman. 2023. Gpqa: A graduate-level google-proof q&a benchmark. arXiv preprint arXiv:2311.12022
2023 arXiv
-
[61]
Jie Ren, Jiaming Luo, Yao Zhao, Kundan Krishna, Mohammad Saleh, Balaji Lakshminarayanan, and Peter J Liu. 2022. Out-of-distribution detection and selective generation for conditional language models. In The Eleventh International Conference on Learning Representations
2022
-
[62]
Maarten Sap, Ronan Le Bras, Daniel Fried, and Yejin Choi. 2022. Neural theory-of-mind? on the limits of social intelligence in large lms. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 3762--3780
2022
-
[63]
Melanie Sclar, Sachin Kumar, Peter West, Alane Suhr, Yejin Choi, and Yulia Tsvetkov. 2023. Minding language models’(lack of) theory of mind: A plug-and-play multi-character belief tracker. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguisti...
2023
-
[64]
Natalie Shapira, Mosh Levy, Seyed Hossein Alavi, Xuhui Zhou, Yejin Choi, Yoav Goldberg, Maarten Sap, and Vered Shwartz. 2024. Clever hans or neural theory of mind? stress testing social reasoning in large language models. In Proceedings of the 18th Conference of the European C...
2024
-
[65]
Vaishnavi Shrivastava, Percy Liang, and Ananya Kumar. 2023. Llamas know what gpts don't show: Surrogate models for confidence estimation. arXiv preprint arXiv:2311.08877
2023 arXiv
-
[66]
Damien Sileo and Marie-francine Moens. 2023. https://doi.org/10.18653/v1/2023.starsem-1.41 Probing neural language models for understanding of words of estimative probability . In Proceedings of the 12th Joint Conference on Lexical and Computational Semantics (*SEM 2023), page...
2023 doi
-
[67]
James WA Strachan, Dalila Albergo, Giulia Borghini, Oriana Pansardi, Eugenio Scaliti, Saurabh Gupta, Krati Saxena, Alessandro Rufo, Stefano Panzeri, Guido Manzi, et al. 2024. Testing theory of mind in large language models and humans. Nature Human Behaviour, pages 1--11
2024
-
[68]
o rgy Szarvas, Veronika Vincze, Rich \'a rd Farkas, Gy \
Gy \"o rgy Szarvas, Veronika Vincze, Rich \'a rd Farkas, Gy \"o rgy M \'o ra, and Iryna Gurevych. 2012. Cross-genre and cross-domain detection of semantic uncertainty. Computational Linguistics, 38(2):335--367
2012
-
[69]
Qwen Team. 2024. https://qwenlm.github.io/blog/qwen2.5/ Qwen2.5: A party of foundation models
2024
-
[70]
Tomer Ullman. 2023. Large language models fail on trivial alterations to theory-of-mind tasks. arXiv preprint arXiv:2302.08399
2023 arXiv
-
[71]
Johan Van der Auwera and Vladimir A Plungian. 1998. Modality’s semantic map. Linguistic Typology
1998
-
[72]
Artem Vazhentsev, Akim Tsvigun, Roman Vashurin, Sergey Petrakov, Daniil Vasilev, Maxim Panov, Alexander Panchenko, and Artem Shelmanov. 2023. Efficient out-of-domain detection for sequence to sequence models. In Findings of the Association for Computational Linguistics: ACL 20...
2023
-
[73]
Cesko C. Voeten. 2023. https://CRAN.R-project.org/package=buildmer buildmer: Stepwise Elimination and Term Reordering for Mixed-Effects Regression . R package version 2.11
2023
-
[74]
Thomas S Wallsten, David V Budescu, Amnon Rapoport, Rami Zwick, and Barbara Forsyth. 1986. Measuring the vague meanings of probability terms. Journal of Experimental Psychology: General, 115(4):348
1986
-
[75]
Alexander Wan, Eric Wallace, and Dan Klein. 2024. https://doi.org/10.18653/v1/2024.acl-long.403 What evidence do language models find convincing? In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 7468--748...
2024 doi
-
[76]
Alex Warstadt, Alicia Parrish, Haokun Liu, Anhad Mohananey, Wei Peng, Sheng-Fu Wang, and Samuel R Bowman. 2020. Blimp: The benchmark of linguistic minimal pairs for english. Transactions of the Association for Computational Linguistics, 8:377--392
2020
-
[77]
Alex Wilf, Sihyun Shawn Lee, Paul Pu Liang, and Louis-Philippe Morency. 2023. Think twice: Perspective-taking improves large language models' theory-of-mind capabilities. arXiv preprint arXiv:2311.10227
2023 arXiv
-
[78]
Sanne JW Willems, Casper J Albers, and Ionica Smeets. 2019. Variability in the interpretation of dutch probability phrases-a risk for miscommunication. arXiv preprint arXiv:1901.09686
2019 arXiv
-
[79]
Sarah A Wu, Erik Brockbank, Hannah Cha, Jan-Philipp Fr \"a nken, Emily Jin, Zhuoyi Huang, Weiyu Liu, Ruohan Zhang, Jiajun Wu, and Tobias Gerstenberg. 2024. Whodunnit? inferring what happened from multimodal evidence. In Proceedings of the Annual Meeting of the Cognitive Scienc...
2024
-
[80]
Yufan Wu, Yinghui He, Yilin Jia, Rada Mihalcea, Yulong Chen, and Naihao Deng. 2023. Hi-tom: A benchmark for evaluating higher-order theory of mind reasoning in large language models. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 10691--10706
2023
-
[81]
Miao Xiong, Zhiyuan Hu, Xinyang Lu, YIFEI LI, Jie Fu, Junxian He, and Bryan Hooi. 2024. Can llms express their uncertainty? an empirical evaluation of confidence elicitation in llms. In The Twelfth International Conference on Learning Representations
2024
-
[82]
Hainiu Xu, Runcong Zhao, Lixing Zhu, Jinhua Du, and Yulan He. 2024. https://doi.org/10.18653/v1/2024.acl-long.466 O pen T o M : A comprehensive benchmark for evaluating theory-of-mind reasoning capabilities of large language models . In Proceedings of the 62nd Annual Meeting o...
2024 doi
-
[83]
An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, Guanting Dong, Haoran Wei, Huan Lin, Jialong Tang, Jialin Wang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Ma, Jin Xu, Jingren Zhou, Jinze Bai, Jinzheng...
2024 arXiv
-
[84]
Gal Yona, Roee Aharoni, and Mor Geva. 2024. https://doi.org/10.18653/v1/2024.emnlp-main.443 Can large language models faithfully express their intrinsic uncertainty in words? In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 7752-...
2024 doi
-
[85]
Kaitlyn Zhou, Jena Hwang, Xiang Ren, and Maarten Sap. 2024. https://doi.org/10.18653/v1/2024.acl-long.198 Relying on the unreliable: The impact of language models' reluctance to express uncertainty . In Proceedings of the 62nd Annual Meeting of the Association for Computationa...
2024 doi
-
[86]
Kaitlyn Zhou, Dan Jurafsky, and Tatsunori B Hashimoto. 2023. Navigating the grey area: How expressions of uncertainty and overconfidence affect language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 5506--5524
2023
-
[87]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...
-
[88]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.