Pith. sign in

REVIEW 4 major objections 6 minor 8 references

A Survey of Idiom Datasets for Psycholinguistic and Computational Research

T0 review · 4 major / 6 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read A survey of 53 idiom datasets finds no bridge between two research traditions.

desk verdict Useful reference map of 53 idiom datasets, but the central 'no relation' claim is asserted without evidence and undercut by the paper's own entries. read the letter →

arxiv 2508.11828 v1 pith:P6OZKSJF submitted 2025-08-15 cs.CL

classification cs.CL
keywords idiomdatasetspsycholinguisticnormscomputationallinguisticsfigurativelanguagemultiwordexpressionstype-tokendistinctionsurvey
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper surveys 53 idiom datasets used in psycholinguistics and computational linguistics, cataloging their content, form, and intended use. Psycholinguistic resources norm idioms as types, providing aggregated human ratings on dimensions like familiarity, transparency, and compositionality; computational datasets work with idioms as tokens, labeling individual instances in context for tasks like idiomaticity detection, disambiguation, paraphrasing, and multilingual modeling. The survey's central claim is that these two traditions have not yet connected: there is no dataset, benchmark, or shared annotation scheme linking psycholinguistic norms to computational idiom resources. The authors argue this gap is structural rather than technical, rooted in the type/token divide, and suggest two possible convergence points: valence/sentiment of idioms, and using computational methods to model human ratings. If correct, the paper implies that progress in idiom-aware NLP will depend on deliberately building bridges between the two fields.

What carries the argument

The central object is the survey's comparative framework: a two-table inventory of 53 idiom datasets organized by discipline (psycholinguistic norming studies vs computational corpora/benchmarks). The analytical work is done by grouping datasets along three axes — content (rating dimensions vs task labels), form (types vs tokens), and intended use (experimental control vs NLP evaluation). The type/token distinction is the load-bearing identity: it explains why annotation schemes do not transfer across the two traditions and why no unified benchmark exists.

What would settle it

Search the reference lists of the surveyed computational idiom papers for citations to psycholinguistic norming studies (or vice versa). Finding even one dataset built on both traditions — for example, a computational idiomaticity corpus whose labels were derived from psycholinguistic norming ratings, or a norming study that used a computational idiom corpus to select items — would weaken the paper's claim that the two fields have no relation.

Watch

Extended reading notes

Core claim

The central discovery is a gap, established by systematic comparison: the paper catalogs 53 idiom datasets and finds that psycholinguistic norming studies and computational idiom corpora operate in parallel, with no dataset, benchmark, or shared annotation scheme connecting them. Psycholinguistic resources aggregate Likert ratings of idiom properties for idioms treated as types; computational resources annotate instances of idioms in textual context, treating them as tokens. The authors identify the type/token distinction as the structural barrier: because the two traditions frame idioms at different granularities, their data cannot be directly compared or combined. They propose two possible

Load-bearing premise

The conclusion that psycholinguistic and computational idiom research are unrelated holds only if the 53 surveyed datasets are a fair and representative sample of the field, and the paper concedes some resources may have been missed.

Editorial extensions

If this is right

  • Future idiom benchmarks should be built with explicit metadata about idiom types vs tokens, and rating dimensions should be defined consistently across norming studies.
  • Psycholinguistic norms (familiarity, age of acquisition, valence) could be attached to computational instance-level datasets to test whether model errors track human-rated properties.
  • Computational methods could be used to predict or explain human ratings of idiom properties, one of the convergence paths the paper names.
  • Multilingual idiom datasets (LIdioms, IMIL, PETCI, SemEval-2022) are growing, but current resources lack semantically aligned idiom instances across languages, limiting cross-lingual bridge-building.
  • Shared metadata conventions and interoperable formats would help integration across the two traditions.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A direct test of the 'no relation' claim would be a citation analysis: count cross-citations between psycholinguistic idiom norming papers and computational idiom papers; if the paper is right, such cross-citations are rare.
  • The type/token gap may be a special case of a broader phenomenon in multiword-expression research, where lexicographic/psycholinguistic resources describe expressions while NLP systems need surface instances; the same split appears in compounds and other multiword units.
  • The survey's list suggests a concrete missing resource: no dataset yet provides both normed psycholinguistic ratings and per-instance contextual annotations for the same idiom set; building one would directly test the paper's central gap.
  • If the paper's diagnosis is correct, embedding norming dimensions as auxiliary labels in token-level datasets could improve idiom representation in language models, though this remains untested.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper surveys 53 datasets for idiom research, split between psycholinguistic norming resources (Table 1) and computational-linguistics resources (Table 2), and summarizes their language coverage, annotation dimensions, task framings, sizes, and availability. It identifies heterogeneity in annotation schemes and argues that the two research traditions have remained largely unconnected, attributing this to a type/token divide: psycholinguistic datasets treat idioms as types, while computational datasets work with tokens in context. The paper also proposes possible future bridges, such as sentiment/valence norms and computational modeling of human ratings.

Significance. If its descriptive content is accurate, this survey is a useful reference for researchers seeking idiom datasets: the appendix tables consolidate a scattered literature, the GitHub resource list is a practical asset, and spot-checks of the reported statistics (e.g., MAGPIE, VNC-Tokens, RU Idioms, LIdioms, PANIG, Gavilán et al.) match the cited publications. The organizing distinction between psycholinguistic type-level norms and computational token-level resources is a suggestive framing. However, the paper's central synthesizing claim—that there is 'no relation yet' between the two fields—is asserted without operationalization or evidence and is in tension with several entries in the paper's own tables. This weakens the main conclusion, although the descriptive survey remains valuable as a reference work.

major comments (4)
  1. [§3.5] The central claim that 'there seems to be no relation yet between psycholinguistic and computational research on idioms' is a negative existential that is never operationalized and no evidence is provided for it: there is no citation-overlap analysis, dataset-reuse analysis, author-overlap check, or shared-task analysis. Moreover, the paper's own contents contradict the claim as stated. Section 2 explicitly notes that Peng et al. (2014) used psycholinguistic valence ratings of idiom component words for automated idiom detection. Table 2 includes Senaldi (2019), a thesis explicitly titled 'Working both sides of the street: computational and psycholinguistic investigations on idiomatic variability.' Several computational datasets in Table 2 (Reddy et al. 2011; Cordeiro et al. 2019; NCS; Swedish MWEs) rely on human compositionality ratings—a psycholinguistic construct. The claim needs eithe
  2. [§3.5] The type/token distinction offered as the explanation for the alleged lack of relation is not supported by Tables 1 and 2. Many computational datasets are explicitly type-level resources: IDIOMENT (580 idiom types), SLIDE (5,000 idiom types), LIdioms (815 types), CCT (7,395 types), CIKB (38K types), ChID (3,848 types), and IdiomKB. Conversely, psycholinguistic norming studies could in principle be applied to tokens, and some psycholinguistic studies use multiple instances or contexts. Thus the type/token dichotomy is not a clean structural barrier, and it cannot bear the weight of the no-relation conclusion.
  3. [§1, Limitations, Appendix] The survey gives no search protocol, inclusion/exclusion criteria, date range, or database/source list. Section 1 only says the search was 'extensive,' and the Limitations section concedes 'some resources might have been missed.' For a negative existential claim about the field, the representativeness of the 53 selected datasets is load-bearing: a single missed bridging dataset would weaken the conclusion. The authors should either describe a systematic, reproducible selection methodology or weaken the conclusion to a more observational statement about the datasets they actually surveyed.
  4. [Tables 1–2, scope] The survey includes nominal-compound compositionality resources (Reddy et al. 2011; Cordeiro et al. 2019; NCS; Swedish MWEs) as 'idiom datasets,' but this classification is not justified. Nominal compounds are multword expressions, not necessarily idioms; including them broadens the scope beyond the abstract's definition of idiomatic expressions whose meanings cannot be inferred from their parts. The paper should either defend this inclusion or exclude such resources, since the dataset count and the trends derived from it depend on this scope decision.
minor comments (6)
  1. [Table 2, Senaldi 2019 row] The row for Senaldi (2019) lists only '90 verb-noun and 24 adjective-noun expressions (types)' without indicating the task, rating dimensions, or the thesis's explicit computational+psycholinguistic design. Adding one or two descriptors would make the table more informative.
  2. [Table 2, RU Idioms row] The entry says '5.4K instances of 100 idiomatic expressions (3K literal, 2.4K idiomatic)'—should specify '100 idiom types' rather than 'expressions' for consistency with the rest of the table.
  3. [Table 2, ID10M row] Missing comma/punctuation: '10K idioms (types) 262781 sentences' should read '10K idioms (types), 262,781 sentences'.
  4. [§3.4] Typo: 'probe LLMs inference' should be 'probe LLMs' inference'.
  5. [Table 1, Beck 2020 row] The Beck (2020) row has no count; if the thesis includes norming data for a specific number of English idioms, that number should be reported, or the row should state that no norming count is specified.
  6. [§2, Morid and Sabourin] The availability column for Morid and Sabourin (2024) is marked '-' in Table 1; the text says 'most of those datasets are publicly available.' Clarify whether this resource is unavailable or whether the dash denotes something else.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: survey-level descriptive claim, no derivation chain; minor self-citations not load-bearing.

full rationale

This paper is a survey, not a derivation. It makes no equations, fits no parameters, and derives no predictions from a model. Its central claim—'there seems to be no relation yet between psycholinguistic and computational research on idioms' (Section 3.5)—is an empirical generalization about the surveyed dataset landscape, supported by the tabulated content of 53 resources rather than by any reductive argument. The type/token explanation is an interpretive framing, not a definitional identity. Two entries in the survey involve the authors' own prior work (Peng et al. 2014; Aharodnik et al. 2018), and one of them (Peng et al.) actually describes a computational method built on a psycholinguistic rating dimension, which, if anything, cuts against the paper's no-relation conclusion rather than serving as an input that forces it. The survey's claims do not reduce to its own citations or to any fitted quantity, so there is no circularity. The appropriate concerns are about completeness and representativeness of the 53-dataset sample, which the Limitations section itself acknowledges ('some resources might have been missed'), but that is a correctness/coverage risk, not circularity.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

No free parameters, no equations, and no invented entities. The survey's claims rest on literature-search completeness and on the comparability of rating constructs across studies, both assumed rather than demonstrated.

assumptions (3)
  • domain assumption The 53 identified datasets constitute a representative sample of idiom research resources.
    The survey's trend and gap conclusions in Sections 3.4 and 3.5 assume the 'extensive search' (Section 1) captured the relevant literature; the Limitations section concedes resources may have been missed, and no search protocol is reported.
  • domain assumption Psycholinguistic norming dimensions with shared names measure comparable constructs across studies.
    Section 2 notes 'the constructs are not always the same across studies' (familiarity, compositionality), yet the survey aggregates these ratings into a single narrative about the psycholinguistic tradition.
  • ad hoc to paper Nominal-compound compositionality resources belong under 'idiom datasets'.
    Reddy et al. 2011, Cordeiro et al. 2019, and Garcia et al. 2021 concern compound nouns, not idioms; their inclusion widens the 53-dataset count without a stated definition of 'idiom dataset'.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Survey of Idiom Datasets for Psycholinguistic and Computational Research." pith.science (2026). https://pith.science/paper/P6OZKSJF

@misc{pith2026250811828,
  author       = {Pith},
  title        = {Pith review of: A Survey of Idiom Datasets for Psycholinguistic and Computational Research},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/P6OZKSJF}},
  note         = {Machine review of arXiv:2508.11828}
}
read the original abstract

Idioms are figurative expressions whose meanings often cannot be inferred from their individual words, making them difficult to process computationally and posing challenges for human experimental studies. This survey reviews datasets developed in psycholinguistics and computational linguistics for studying idioms, focusing on their content, form, and intended use. Psycholinguistic resources typically contain normed ratings along dimensions such as familiarity, transparency, and compositionality, while computational datasets support tasks like idiomaticity detection/classification, paraphrasing, and cross-lingual modeling. We present trends in annotation practices, coverage, and task framing across 53 datasets. Although recent efforts expanded language coverage and task diversity, there seems to be no relation yet between psycholinguistic and computational research on idioms.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

8 extracted references · 8 canonical work pages

  1. [3]

    In Proceedings of the Eleventh In- ternational Conference on Language Resources and Evaluation (LREC 2018), Miyazaki, Japan

    Examining the tip of the iceberg: A data set for idiom translation. In Proceedings of the Eleventh In- ternational Conference on Language Resources and Evaluation (LREC 2018), Miyazaki, Japan. European Language Resources Association (ELRA). Afsaneh Fazly, Paul Cook, and Suzanne Stevenson

  2. [1968]

    Journal of Experimental Psy- chology, 76(1):1–25

    Concreteness, imagery, and meaningfulness values for 925 nouns. Journal of Experimental Psy- chology, 76(1):1–25. Jing Peng, Anna Feldman, and Ekaterina Vylomova

  3. [1994]

    Language, 70(3):491–538

    Idioms. Language, 70(3):491–538. Charles E. Osgood, George J. Suci, and Percy H. Tan- nenbaum. 1957. The measurement of meaning. Uni- versity of Illinois Press. Irene Pagliai. 2023. Bridging the Gap: Creation of a Lexicon of 150 Pairs of English and Italian Idioms Including Normed Variables for the Exploration of Idiomatic Ambiguity. Journal of Open Human...

  4. [2008]

    In Proceedings of the LREC Workshop Towards a Shared Task for Mul- tiword Expressions (MWE 2008) , Marrakech, Mo- rocco

    The vnc-tokens dataset. In Proceedings of the LREC Workshop Towards a Shared Task for Mul- tiword Expressions (MWE 2008) , Marrakech, Mo- rocco. Silvio Cordeiro, Aline Villavicencio, Marco Idiart, and Carlos Ramisch. 2019. Unsupervised compositional- ity prediction of nominal compounds. Computational Linguistics, 45(1):1–57. Ricarda Dormeyer and Ingrid Fi...

  5. [2009]

    Comparative Study of Multilingual Idioms and Similes in Large Language Models

    Unsupervised type and token identification of idiomatic expressions. Computational Linguistics, 35(1):61–103. Christiane Fellbaum and Alexander Geyken. 2005. Transforming a Corpus Into a Lexical Resource - the Berlin Idiom Project. Revue française de linguistique appliquée, X(2):49–62. Fraser. 1970. Idioms within a transformational grammar. Foundation of ...

  6. [2011]

    PETCI: A Parallel English Translation Dataset of Chinese Idioms

    Descriptive norms for 245 Italian idiomatic ex- pressions. Behavior Research Methods, 43:110–123. Patrizia Tabossi, Rachele Fanari, and Kinou Wolf. 2008. Processing idiomatic expressions: Effects of se- mantic compositionality. Journal of Experimen- tal Psychology: Learning, Memory, and Cognition, 34(2):313–327. Kenan Tang. 2022. Petci: A parallel english...

  7. [2014]

    In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2019–2027, Doha, Qatar

    Classifying idiomatic and literal expressions using topic models and intensity of emotions. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2019–2027, Doha, Qatar. Association for Com- putational Linguistics. Maria Pershina, Yifan He, and Ralph Grishman. 2015. Idiom paraphrases: Seventh heaven vs cl...

  8. [2018]

    Going to town

    Designing a Russian idiom-annotated cor- pus. In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018), Miyazaki, Japan. European Language Resources Association (ELRA). Sara D. Beck. 2020. Native and Non-native Idiom Pro- cessing: Same Difference. Ph.D. thesis, University of Tübingen. Sara D. Beck and Andrea...

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.