REVIEW 4 major objections 6 minor 8 references
A Survey of Idiom Datasets for Psycholinguistic and Computational Research
T0 review · 4 major / 6 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read A survey of 53 idiom datasets finds no bridge between two research traditions.
desk verdict Useful reference map of 53 idiom datasets, but the central 'no relation' claim is asserted without evidence and undercut by the paper's own entries. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the survey's comparative framework: a two-table inventory of 53 idiom datasets organized by discipline (psycholinguistic norming studies vs computational corpora/benchmarks). The analytical work is done by grouping datasets along three axes — content (rating dimensions vs task labels), form (types vs tokens), and intended use (experimental control vs NLP evaluation). The type/token distinction is the load-bearing identity: it explains why annotation schemes do not transfer across the two traditions and why no unified benchmark exists.
What would settle it
Search the reference lists of the surveyed computational idiom papers for citations to psycholinguistic norming studies (or vice versa). Finding even one dataset built on both traditions — for example, a computational idiomaticity corpus whose labels were derived from psycholinguistic norming ratings, or a norming study that used a computational idiom corpus to select items — would weaken the paper's claim that the two fields have no relation.
Extended reading notes
Core claim
The central discovery is a gap, established by systematic comparison: the paper catalogs 53 idiom datasets and finds that psycholinguistic norming studies and computational idiom corpora operate in parallel, with no dataset, benchmark, or shared annotation scheme connecting them. Psycholinguistic resources aggregate Likert ratings of idiom properties for idioms treated as types; computational resources annotate instances of idioms in textual context, treating them as tokens. The authors identify the type/token distinction as the structural barrier: because the two traditions frame idioms at different granularities, their data cannot be directly compared or combined. They propose two possible
Load-bearing premise
The conclusion that psycholinguistic and computational idiom research are unrelated holds only if the 53 surveyed datasets are a fair and representative sample of the field, and the paper concedes some resources may have been missed.
Editorial extensions
If this is right
- Future idiom benchmarks should be built with explicit metadata about idiom types vs tokens, and rating dimensions should be defined consistently across norming studies.
- Psycholinguistic norms (familiarity, age of acquisition, valence) could be attached to computational instance-level datasets to test whether model errors track human-rated properties.
- Computational methods could be used to predict or explain human ratings of idiom properties, one of the convergence paths the paper names.
- Multilingual idiom datasets (LIdioms, IMIL, PETCI, SemEval-2022) are growing, but current resources lack semantically aligned idiom instances across languages, limiting cross-lingual bridge-building.
- Shared metadata conventions and interoperable formats would help integration across the two traditions.
Reading between the lines
- A direct test of the 'no relation' claim would be a citation analysis: count cross-citations between psycholinguistic idiom norming papers and computational idiom papers; if the paper is right, such cross-citations are rare.
- The type/token gap may be a special case of a broader phenomenon in multiword-expression research, where lexicographic/psycholinguistic resources describe expressions while NLP systems need surface instances; the same split appears in compounds and other multiword units.
- The survey's list suggests a concrete missing resource: no dataset yet provides both normed psycholinguistic ratings and per-instance contextual annotations for the same idiom set; building one would directly test the paper's central gap.
- If the paper's diagnosis is correct, embedding norming dimensions as auxiliary labels in token-level datasets could improve idiom representation in language models, though this remains untested.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper surveys 53 datasets for idiom research, split between psycholinguistic norming resources (Table 1) and computational-linguistics resources (Table 2), and summarizes their language coverage, annotation dimensions, task framings, sizes, and availability. It identifies heterogeneity in annotation schemes and argues that the two research traditions have remained largely unconnected, attributing this to a type/token divide: psycholinguistic datasets treat idioms as types, while computational datasets work with tokens in context. The paper also proposes possible future bridges, such as sentiment/valence norms and computational modeling of human ratings.
Significance. If its descriptive content is accurate, this survey is a useful reference for researchers seeking idiom datasets: the appendix tables consolidate a scattered literature, the GitHub resource list is a practical asset, and spot-checks of the reported statistics (e.g., MAGPIE, VNC-Tokens, RU Idioms, LIdioms, PANIG, Gavilán et al.) match the cited publications. The organizing distinction between psycholinguistic type-level norms and computational token-level resources is a suggestive framing. However, the paper's central synthesizing claim—that there is 'no relation yet' between the two fields—is asserted without operationalization or evidence and is in tension with several entries in the paper's own tables. This weakens the main conclusion, although the descriptive survey remains valuable as a reference work.
major comments (4)
- [§3.5] The central claim that 'there seems to be no relation yet between psycholinguistic and computational research on idioms' is a negative existential that is never operationalized and no evidence is provided for it: there is no citation-overlap analysis, dataset-reuse analysis, author-overlap check, or shared-task analysis. Moreover, the paper's own contents contradict the claim as stated. Section 2 explicitly notes that Peng et al. (2014) used psycholinguistic valence ratings of idiom component words for automated idiom detection. Table 2 includes Senaldi (2019), a thesis explicitly titled 'Working both sides of the street: computational and psycholinguistic investigations on idiomatic variability.' Several computational datasets in Table 2 (Reddy et al. 2011; Cordeiro et al. 2019; NCS; Swedish MWEs) rely on human compositionality ratings—a psycholinguistic construct. The claim needs eithe
- [§3.5] The type/token distinction offered as the explanation for the alleged lack of relation is not supported by Tables 1 and 2. Many computational datasets are explicitly type-level resources: IDIOMENT (580 idiom types), SLIDE (5,000 idiom types), LIdioms (815 types), CCT (7,395 types), CIKB (38K types), ChID (3,848 types), and IdiomKB. Conversely, psycholinguistic norming studies could in principle be applied to tokens, and some psycholinguistic studies use multiple instances or contexts. Thus the type/token dichotomy is not a clean structural barrier, and it cannot bear the weight of the no-relation conclusion.
- [§1, Limitations, Appendix] The survey gives no search protocol, inclusion/exclusion criteria, date range, or database/source list. Section 1 only says the search was 'extensive,' and the Limitations section concedes 'some resources might have been missed.' For a negative existential claim about the field, the representativeness of the 53 selected datasets is load-bearing: a single missed bridging dataset would weaken the conclusion. The authors should either describe a systematic, reproducible selection methodology or weaken the conclusion to a more observational statement about the datasets they actually surveyed.
- [Tables 1–2, scope] The survey includes nominal-compound compositionality resources (Reddy et al. 2011; Cordeiro et al. 2019; NCS; Swedish MWEs) as 'idiom datasets,' but this classification is not justified. Nominal compounds are multword expressions, not necessarily idioms; including them broadens the scope beyond the abstract's definition of idiomatic expressions whose meanings cannot be inferred from their parts. The paper should either defend this inclusion or exclude such resources, since the dataset count and the trends derived from it depend on this scope decision.
minor comments (6)
- [Table 2, Senaldi 2019 row] The row for Senaldi (2019) lists only '90 verb-noun and 24 adjective-noun expressions (types)' without indicating the task, rating dimensions, or the thesis's explicit computational+psycholinguistic design. Adding one or two descriptors would make the table more informative.
- [Table 2, RU Idioms row] The entry says '5.4K instances of 100 idiomatic expressions (3K literal, 2.4K idiomatic)'—should specify '100 idiom types' rather than 'expressions' for consistency with the rest of the table.
- [Table 2, ID10M row] Missing comma/punctuation: '10K idioms (types) 262781 sentences' should read '10K idioms (types), 262,781 sentences'.
- [§3.4] Typo: 'probe LLMs inference' should be 'probe LLMs' inference'.
- [Table 1, Beck 2020 row] The Beck (2020) row has no count; if the thesis includes norming data for a specific number of English idioms, that number should be reported, or the row should state that no norming count is specified.
- [§2, Morid and Sabourin] The availability column for Morid and Sabourin (2024) is marked '-' in Table 1; the text says 'most of those datasets are publicly available.' Clarify whether this resource is unavailable or whether the dash denotes something else.
Circularity Check
No circularity: survey-level descriptive claim, no derivation chain; minor self-citations not load-bearing.
full rationale
This paper is a survey, not a derivation. It makes no equations, fits no parameters, and derives no predictions from a model. Its central claim—'there seems to be no relation yet between psycholinguistic and computational research on idioms' (Section 3.5)—is an empirical generalization about the surveyed dataset landscape, supported by the tabulated content of 53 resources rather than by any reductive argument. The type/token explanation is an interpretive framing, not a definitional identity. Two entries in the survey involve the authors' own prior work (Peng et al. 2014; Aharodnik et al. 2018), and one of them (Peng et al.) actually describes a computational method built on a psycholinguistic rating dimension, which, if anything, cuts against the paper's no-relation conclusion rather than serving as an input that forces it. The survey's claims do not reduce to its own citations or to any fitted quantity, so there is no circularity. The appropriate concerns are about completeness and representativeness of the 53-dataset sample, which the Limitations section itself acknowledges ('some resources might have been missed'), but that is a correctness/coverage risk, not circularity.
Assumptions & free parameters
assumptions (3)
- domain assumption The 53 identified datasets constitute a representative sample of idiom research resources.
- domain assumption Psycholinguistic norming dimensions with shared names measure comparable constructs across studies.
- ad hoc to paper Nominal-compound compositionality resources belong under 'idiom datasets'.
Cite this review
Pith. "Pith review of A Survey of Idiom Datasets for Psycholinguistic and Computational Research." pith.science (2026). https://pith.science/paper/P6OZKSJF
@misc{pith2026250811828,
author = {Pith},
title = {Pith review of: A Survey of Idiom Datasets for Psycholinguistic and Computational Research},
year = {2026},
howpublished = {\url{https://pith.science/paper/P6OZKSJF}},
note = {Machine review of arXiv:2508.11828}
}
read the original abstract
Idioms are figurative expressions whose meanings often cannot be inferred from their individual words, making them difficult to process computationally and posing challenges for human experimental studies. This survey reviews datasets developed in psycholinguistics and computational linguistics for studying idioms, focusing on their content, form, and intended use. Psycholinguistic resources typically contain normed ratings along dimensions such as familiarity, transparency, and compositionality, while computational datasets support tasks like idiomaticity detection/classification, paraphrasing, and cross-lingual modeling. We present trends in annotation practices, coverage, and task framing across 53 datasets. Although recent efforts expanded language coverage and task diversity, there seems to be no relation yet between psycholinguistic and computational research on idioms.
Reference graph
Works this paper leans on
-
[3]
Examining the tip of the iceberg: A data set for idiom translation. In Proceedings of the Eleventh In- ternational Conference on Language Resources and Evaluation (LREC 2018), Miyazaki, Japan. European Language Resources Association (ELRA). Afsaneh Fazly, Paul Cook, and Suzanne Stevenson
work page 2018
-
[1968]
Journal of Experimental Psy- chology, 76(1):1–25
Concreteness, imagery, and meaningfulness values for 925 nouns. Journal of Experimental Psy- chology, 76(1):1–25. Jing Peng, Anna Feldman, and Ekaterina Vylomova
-
[1994]
Idioms. Language, 70(3):491–538. Charles E. Osgood, George J. Suci, and Percy H. Tan- nenbaum. 1957. The measurement of meaning. Uni- versity of Illinois Press. Irene Pagliai. 2023. Bridging the Gap: Creation of a Lexicon of 150 Pairs of English and Italian Idioms Including Normed Variables for the Exploration of Idiomatic Ambiguity. Journal of Open Human...
work page 1957
-
[2008]
The vnc-tokens dataset. In Proceedings of the LREC Workshop Towards a Shared Task for Mul- tiword Expressions (MWE 2008) , Marrakech, Mo- rocco. Silvio Cordeiro, Aline Villavicencio, Marco Idiart, and Carlos Ramisch. 2019. Unsupervised compositional- ity prediction of nominal compounds. Computational Linguistics, 45(1):1–57. Ricarda Dormeyer and Ingrid Fi...
work page 2008
-
[2009]
Comparative Study of Multilingual Idioms and Similes in Large Language Models
Unsupervised type and token identification of idiomatic expressions. Computational Linguistics, 35(1):61–103. Christiane Fellbaum and Alexander Geyken. 2005. Transforming a Corpus Into a Lexical Resource - the Berlin Idiom Project. Revue française de linguistique appliquée, X(2):49–62. Fraser. 1970. Idioms within a transformational grammar. Foundation of ...
work page Pith review arXiv 2005
-
[2011]
PETCI: A Parallel English Translation Dataset of Chinese Idioms
Descriptive norms for 245 Italian idiomatic ex- pressions. Behavior Research Methods, 43:110–123. Patrizia Tabossi, Rachele Fanari, and Kinou Wolf. 2008. Processing idiomatic expressions: Effects of se- mantic compositionality. Journal of Experimen- tal Psychology: Learning, Memory, and Cognition, 34(2):313–327. Kenan Tang. 2022. Petci: A parallel english...
work page Pith review arXiv 2008
-
[2014]
Classifying idiomatic and literal expressions using topic models and intensity of emotions. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2019–2027, Doha, Qatar. Association for Com- putational Linguistics. Maria Pershina, Yifan He, and Ralph Grishman. 2015. Idiom paraphrases: Seventh heaven vs cl...
work page 2014
-
[2018]
Designing a Russian idiom-annotated cor- pus. In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018), Miyazaki, Japan. European Language Resources Association (ELRA). Sara D. Beck. 2020. Native and Non-native Idiom Pro- cessing: Same Difference. Ph.D. thesis, University of Tübingen. Sara D. Beck and Andrea...
work page 2018
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.