REVIEW 3 major objections 5 minor 85 references
SenWiCh: Sense-Annotation of Low-Resource Languages for WiC using Hybrid Methods
T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read This paper releases sense-annotated corpora for ten low-resource languages and a semi-automatic annotation method that speeds up finding rare word senses.
desk verdict A genuinely useful ten-language WiC/WSD resource, but the abstract's confidence outruns the evidence: no direct annotation-quality validation and an efficiency claim built on a tiny, non-transparent sample. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is a two-stage pipeline. First, a contextual embedding model (mBERT, XLM-R, XL-LEXEME, or a language-specific model) encodes every sentence containing a target polysemous word; K-Means or agglomerative clustering groups the encodings, and UMAP or MDS projects them into a 2D plot that annotators can inspect. Annotators select dispersed sentences, prioritising those the embedding space treats as distinct, and assign sense labels from dictionary inventories. The efficiency claim is carried by the Lift metric, defined as Precision(sense)/Prior(sense), which quantifies how much faster the guided selection finds a sense than random sampling. On the evaluation side, the datasets are converted to Word-in-Context (WiC) form, where the task is to decide whether a target word has the same sense in two sentences, using a word-level 70/15/15 split with sentence redistribution and balanced same/different pairing.
What would settle it
Draw a stratified random sample of WiC sentence pairs from each language's test set, have independent native speakers annotate same versus different sense without seeing the released labels, and measure agreement with the dataset. If any language's human-model agreement is near chance or substantially below the model's own accuracy, the claim that these are high-quality evaluation datasets for that language is falsified.
Extended reading notes
Core claim
The paper's own contribution is the resource release plus the method that produced it: sense-annotated corpora for Azerbaijani, Kannada, Korean, Marathi, Polish, Punjabi, Swahili, Telugu, Urdu, and Vietnamese, formatted both as WSD annotations and as WiC sentence-pair evaluation sets. The release is accompanied by a hybrid annotation tool in which transformer embeddings of sentences are clustered and projected into an interactive 2D space, letting native-speaker annotators select sentences that cover each sense rather than scanning random samples. The paper argues the method is demonstrably beneficial by measuring Lift, the ratio of selection precision to a sense's corpus prior; for subordinate senses this reaches 900% in Kannada and 1414% in Urdu, corresponding to roughly an eightfold reduction in sentences that must be reviewed. Experiments with XLM-R show the datasets are usable and informative: zero-shot English transfer outperforms full-shot target training in eight of ten languages, mixed training improves most languages, and transfer accuracy varies from 62.4% to 82.2%, which the paper takes as evidence that per-language evaluation is necessary rather than optional.
Load-bearing premise
The load-bearing premise is that the released sense labels are correct, but the paper reports no inter-annotator agreement or human-accuracy figures for the new datasets; the experiments use downstream transfer accuracy as an indirect quality signal, which cannot distinguish correct annotations from annotations that are merely learnable by a model.
Editorial extensions
If this is right
- Researchers can now evaluate WiC-style polysemy disambiguation in ten low-resource languages that previously lacked such benchmarks.
- The reported Lift scores imply that the semi-automatic method lowers the cost of building similar sense-annotated datasets, making broader language coverage more feasible.
- The finding that zero-shot transfer from English often beats full-shot training on small data warns that tiny target-language training sets can be worse than none for polysemy.
- Mixed training, where English fine-tuning is followed by target-language fine-tuning, gives the best results for most languages, suggesting small amounts of in-language data still matter.
- The variation in zero-shot accuracy across languages shows that conclusions about cross-lingual transfer drawn from one or two target languages may not generalise.
Reading between the lines
- A direct check the paper does not carry out: independent native-speaker agreement on a sample of the released WiC pairs would convert the indirect quality signal into a direct one, and would also reveal which languages have noisier labels.
- The zero-shot-beats-full-shot pattern suggests a crossover point in training-set size below which target-language fine-tuning adds noise; a testable extension would vary training data per language to find that threshold.
- Because the pipeline's gains depend on the embedding model, Lift should vary with model coverage; comparing Lift across languages with different encoder choices could identify when the hybrid method needs language-specific models.
- The dataset's design, spanning five language families and several scripts, invites studies of whether polysemy transfer failures track typological or orthographic distance from English.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces SenWiCh, a semi-automatic hybrid annotation pipeline that combines embedding-based clustering with human verification, and releases sense-annotated corpora and Word-in-Context (WiC) evaluation datasets for ten low-resource languages. The authors evaluate their datasets by fine-tuning XLM-R bi-encoders under full-shot, zero-shot, and mixed training conditions, and they report a Lift metric intended to quantify the efficiency of their semi-automatic sentence-selection method. The paper's two central claims are that it releases new sense-annotated datasets and that it presents a 'demonstrably beneficial' annotation method.
Significance. If the sense labels are reliable, the released datasets would fill a genuine gap in evaluation resources for low-resource languages, covering diverse language families and scripts and enabling cross-lingual transfer research in WiC/WSD. The released code and annotation tool are also potentially useful for future resource creation. However, the paper's evidence for label reliability is indirect: the WiC transfer results show that the labels are learnable, not that they are correct, and the Lift metric is computed on a small and non-transparent sample of words. The significance of the contribution therefore depends on validation that is currently missing.
major comments (3)
- [§4.1 (Kannada), §4, §8] The central resource claim rests on sense labels whose correctness is never directly demonstrated. No inter-annotator agreement (e.g., kappa) or human-accuracy figure is reported for any of the ten languages. For Kannada, the text says an 'independent reviewer verified all annotations' but gives no agreement rate; for the other languages, the number of annotators and the disagreement resolution process are unspecified. The Limitations section (Section 8) acknowledges 'potential inconsistencies in ... annotation quality' but does not bound this risk. Because the WiC pairs are derived from these sense labels, any systematic labeling error propagates into the released benchmark and affects all experimental conclusions.
- [§6, Table 3] The claim that the transfer results 'demonstrate that we were able to produce high-quality datasets' is not entailed by the evidence. Full-shot accuracy on the released WiC pairs shows that the labels are learnable by an XLM-R model, but learnability does not imply correctness; a model can also learn systematic misannotations. In addition, the full-shot versus zero-shot differences are reported without error bars or significance tests, and several test sets are small (e.g., 656 pairs for Korean and 1,397 for Vietnamese, per Appendix C), so the claimed '8 out of 10' zero-shot advantage may be partly due to noise.
- [§3.3, Table 1, Appendix B] The Lift metric that supports the 'demonstrably beneficial' annotation method is computed on only 4 of the 10 languages and on 20 words, with no error bars, confidence intervals, or significance tests. The text does not state how Precision(sense) was measured (e.g., against which gold standard, by how many annotators, on how many sentences per sense). Moreover, Appendix B reports an infinite Lift value for one word (Punjabi 'Defeat' with prior = 0), which indicates an undefined division by zero and shows that the metric is not handled in a statistically sound way. This evidence does not substantiate the abstract's claim of a 'demonstrably beneficial' method.
minor comments (5)
- [§4] Swahili is described as 'Afro-semitic'; Swahili is a Bantu language (Niger-Congo family), not Afro-Asiatic. This factual misclassification should be corrected.
- [§3.3] The term 'adjusted Lift' is introduced but never defined. If the metric differs from standard Lift, the adjustment should be specified; otherwise, the word 'adjusted' should be removed.
- [§4.1] Several references in the language-specific sections are incomplete: 'Kishiyev et al.' and 'Ramesh et al.' and 'Khan et al.' lack publication years, and 'Google AI, 2018' is not a standard citation. Please add full citation details.
- [§6] The inference that low full-shot performance in Korean, Swahili, and Vietnamese is 'most likely due to their smaller training size rather than quality issues' is not warranted by the evidence, because zero-shot and full-shot involve different training setups and the test sets are small. A more cautious interpretation is needed.
- [Abstract] The phrase 'demonstrably beneficial annotation method' overstates the evidence in the current manuscript. Given the limitations of the Lift evaluation, the abstract should say 'potentially beneficial' or the authors should provide more comprehensive validation.
Circularity Check
No circularity found: the WiC evaluation and the Lift efficiency metric are measured against held-out manual sense labels, not against the tool's own outputs or a fitted quantity.
full rationale
The paper's derivation chain is self-contained: sense inventories come from external dictionaries; sentence sampling and clustering are only a selection aid; human annotators assign sense labels; WiC pairs are then constructed from these labels and split so no sentence appears in more than one split (§4.2). The supervised XLM-R bi-encoder is trained on the training pairs, its decision threshold is tuned on the validation set, and Table 3 reports accuracy on held-out test pairs—so the reported numbers are generalization measurements, not reconstructions of the training signal. The Lift metric in §3.3 is defined as Precision(sense)/Prior(sense), where the prior comes from a random 100-sentence sample and precision is the proportion of automatically selected sentences that human annotators assigned to the sense; this is an enrichment measure against manual labels, not a self-prediction. The only overlapping-author citation (Dairkee and Dubossarsky 2024) is used as contrastive prior work and is contradicted by the paper's own zero-shot results, so it is not load-bearing. The Limitations section's acknowledgment of 'potential inconsistencies in ... annotation quality' and the absence of inter-annotator agreement are legitimate data-quality and validity concerns, but they concern label correctness, not circularity of the derivation.
Assumptions & free parameters
free parameters (2)
- clustering parameters (number of clusters, linkage) =
not reported
- WiC threshold per model =
not reported
assumptions (4)
- domain assumption The sense inventories from the source dictionaries are appropriate and complete for the WiC task in each language.
- domain assumption The embedding-based clustering surfaces the same sense distinctions that human annotators would make.
- domain assumption The 100-sentence random sample used to estimate sense priors is representative of the corpus sense distribution.
- domain assumption The WiC pairing algorithm preserves the original sense distribution across train/dev/test splits.
Cite this review
Pith. "Pith review of SenWiCh: Sense-Annotation of Low-Resource Languages for WiC using Hybrid Methods." pith.science (2026). https://pith.science/paper/5DYNK52S
@misc{pith2026250523714,
author = {Pith},
title = {Pith review of: SenWiCh: Sense-Annotation of Low-Resource Languages for WiC using Hybrid Methods},
year = {2026},
howpublished = {\url{https://pith.science/paper/5DYNK52S}},
note = {Machine review of arXiv:2505.23714}
}
read the original abstract
This paper addresses the critical need for high-quality evaluation datasets in low-resource languages to advance cross-lingual transfer. While cross-lingual transfer offers a key strategy for leveraging multilingual pretraining to expand language technologies to understudied and typologically diverse languages, its effectiveness is dependent on quality and suitable benchmarks. We release new sense-annotated datasets of sentences containing polysemous words, spanning ten low-resource languages across diverse language families and scripts. To facilitate dataset creation, the paper presents a demonstrably beneficial semi-automatic annotation method. The utility of the datasets is demonstrated through Word-in-Context (WiC) formatted experiments that evaluate transfer on these low-resource languages. Results highlight the importance of targeted dataset creation and evaluation for effective polysemy disambiguation in low-resource settings and transfer studies. The released datasets and code aim to support further research into fair, robust, and truly multilingual NLP.
Figures
Reference graph
Works this paper leans on
-
[1]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Kabir Ahuja, Harshita Diddee, Rishav Hada, Millicent Ochieng, Krithika Ramesh, Prachi Jain, Akshay Nambi, Tanuja Ganu, Sameer Segal, Mohamed Ahmed, Kalika Bali, and Sunayana Sitaram. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.258 MEGA : Multilingual evaluation of generative AI . In Proceedings of the 2023 Conference on Empirical Methods in Natural ...
-
[4]
Google AI. 2018. https://huggingface.co/bert-base-multilingual-cased Multilingual bert (mbert) . Accessed: August 2024
2018
-
[5]
Anthropic. 2024. https://www.anthropic.com/news/claude-3-5-sonnet Claude 3.5 sonnet . Available at https://www.anthropic.com/news/claude-3-5-sonnet
2024
-
[6]
Michele Bevilacqua, Tommaso Pasini, Alessandro Raganato, and Roberto Navigli. 2021. Recent trends in word sense disambiguation: A survey. In International Joint Conference on Artificial Intelligence, pages 4330--4338. International Joint Conference on Artificial Intelligence, Inc
2021
-
[7]
Terra Blevins, Mandar Joshi, and Luke Zettlemoyer. 2021. https://aclanthology.org/2021.eacl-main.219/ FEWS : Large-scale, low-shot word sense disambiguation with the dictionary . In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 2554--2560. Association for Computational Linguistics
work page 2021
-
[8]
Anna Breit, Artem Revenko, Kiamehr Rezaee, Mohammad Taher Pilehvar, and Jose Camacho-Collados. 2021. https://doi.org/10.18653/v1/2021.eacl-main.140 WiC-TSV : An evaluation benchmark for target sense verification of words in context . In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume...
Show all 85 references
-
[9]
Singh Brothers. 2006. https://www.exoticindiaart.com/book/details/punjabi-english-dictionary-idk610/ Punjabi-English Dictionary . Singh Brothers, Amritsar. Accessed: August 2024
2006
-
[10]
Gregory Maxwell Bruce. 2021. https://www.degruyter.com/document/doi/10.1515/9781474467216/html Urdu Vocabulary: A Workbook for Intermediate and Advanced Students . Edinburgh University Press, Edinburgh. Accessed: August 2024
2021 doi
-
[11]
Agostina Calabrese, Michele Bevilacqua, Roberto Navigli, et al. 2020. Fatality killed the cat or: Babelpic, a multimodal dataset for non-concrete concepts. In Proceedings of the 58th annual meeting of the association for computational linguistics, pages 4680--4686. Association...
2020
-
[12]
Pierluigi Cassotti, Lucia Siciliani, Marco DeGemmis, Giovanni Semeraro, and Pierpaolo Basile. 2023. Xl-lexeme: Wic pretrained model for cross-lingual lexical semantic change. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: ...
2023
-
[13]
Nadezhda Chirkova and Vassilina Nikoulina. 2024. https://aclanthology.org/2024.inlg-main.53/ Zero-shot cross-lingual transfer in instruction tuning of large language models . In Proceedings of the 17th International Natural Language Generation Conference, pages 695--708, Tokyo...
2024
-
[14]
Chuo Kikuu cha Dar es Salaam, Taasisi ya Taaluma za Kiswahili . 2013. https://books.google.com/books/about/Kamusi_ya_Kiswahili_sanifu.html?id=Lk2CAQAACAAJ Kamusi ya Kiswahili Sanifu . Oxford University Press, East Africa Limited, Nairobi, Kenya. Accessed: August 2024
2013
-
[15]
Leipzig Corpora Collection. 2018. pol\_news\_2018: Polish news corpus. https://corpora.wortschatz-leipzig.de/ic?corpusId=pol_news_2018. Accessed: August 2024
2018
-
[16]
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzm \'a n, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. https://doi.org/10.18653/v1/2020.acl-main.747 Unsupervised cross-lingual representation learning ...
2020 doi
-
[17]
Farheen Dairkee and Haim Dubossarsky. 2024. Strengthening the wic: New polysemy dataset in hindi and lack of cross lingual transfer. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), pa...
2024
-
[18]
DBMDZ . 2025. https://huggingface.co/dbmdz/bert-base-turkish-128k-uncased BERT-Base Turkish 128K Uncased . Available at Hugging Face, Accessed: August 2024
2025
-
[19]
ukasz Deg \'o rski and Adam Przepi \'o rkowski. 2012. Recznie znakowany milionowy podkorpus nkjp. Przepi \'o rkowski et al.[17] , pages 51--58
2012
-
[20]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. https://arxiv.org/abs/1810.04805 BERT : Pre-training of deep bidirectional transformers for language understanding . arXiv preprint, arXiv:1810.04805
2019 arXiv
-
[21]
B a \.z ej Dolicki and Gerasimos Spanakis. 2021. Analysing the impact of linguistic features on cross-lingual transfer. arXiv preprint arXiv:2105.05975
2021 arXiv
-
[22]
Philip Edmonds and Scott Cotton. 2001. https://doi.org/10.3115/1117755.1117756 Senseval-2: Overview . In Proceedings of the Second International Workshop on Evaluating Word Sense Disambiguation Systems, pages 1--5. Association for Computational Linguistics
2001
-
[23]
Phan Van Giuong. 2014. https://books.google.com/books/about/Tuttle_Concise_Vietnamese_Dictionary.html?id=zbxGCgAAQBAJ Tuttle Concise Vietnamese Dictionary: Vietnamese-English, English-Vietnamese . Tuttle Publishing, North Clarendon, VT. Accessed: August 2024
2014
-
[24]
K. K. Goswami. 2000. https://www.amazon.com/Punjabi-English-English-Punjabi-Dictionary-Hippocrene/dp/0781807166 Punjabi-English/English-Punjabi Dictionary . Hippocrene Books, New York. Accessed: August 2024
2000
-
[25]
J. Ham, Y. J. Choe, K. Park, I. Choi, and H. Soh. 2020. https://huggingface.co/jhgan/ko-sroberta-multitask Ko-sroberta multitask model . Available at Hugging Face, Accessed: August 2024
2020
-
[26]
Junxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg-Kirkpatrick, and Graham Neubig. 2021. Towards a unified view of parameter-efficient transfer learning. arXiv preprint arXiv:2110.04366
2021 arXiv
-
[27]
Hesret Hesenov. 2007. Azerbaycan dilinin omonimler lugeti. Serq Qerb nesriyyati. Baki
2007
-
[28]
Eduard Hovy, Mitchell Marcus, Martha Palmer, Lance Ramshaw, and Ralph Weischedel. 2006. https://doi.org/10.3115/1614049.1614064 Ontonotes: The 90\ In Proceedings of the Human Language Technology Conference of the NAACL, Companion Volume: Short Papers, pages 57--60. Association...
2006
-
[29]
Bushra Jawaid, Amir Kamran, and Ond r ej Bojar. 2014. https://aclanthology.org/L14-1449/ A tagged corpus and a tagger for urdu . In Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC'14), pages 2938--2943, Reykjavik, Iceland. European ...
2014
-
[30]
Raviraj Joshi. 2022. L3cube-hindbert and devbert: Pre-trained bert transformer models for devanagari based hindi and marathi languages. arXiv preprint arXiv:2211.11418. https://arxiv.org/abs/2211.11418
2022 arXiv
-
[31]
Raviraj Joshi et al. 2022. https://arxiv.org/abs/2205.14728 L3Cube-MahaNLP : Marathi natural language processing datasets, models, and library . Accessed: August 2024
2022 arXiv
-
[32]
S. S. Joshi. 2009. https://www.amazon.com/Punjabi-English-Dictionary-Yuniwarasiti-Panjabi-angarezi/dp/8173800960 Punjabi-English Dictionary: Panjabi Yuniwarasiti Panjabi-Angarezi Kosha . Punjabi University, Patiala. Accessed: August 2024
2009
-
[34]
Khapra, and Pratyush Kumar
Divyanshu Kakwani, Anoop Kunchukuttan, Satish Golla, Gokul N.C., Avik Bhattacharyya, Mitesh M. Khapra, and Pratyush Kumar. 2020 b . https://doi.org/10.18653/v1/2020.findings-emnlp.445 IndicNLPSuite: Monolingual Corpora, Evaluation Benchmarks and Pre-trained Multilingual Langua...
2020 doi
-
[35]
Mohammed Safi Ur Rahman Khan, Priyam Mehta, Ananth Sankar, Umashankar Kumaravelan, Sumanth Doddapaneni, Suriyaprasaad G, Varun Balan G, Sparsh Jain, Anoop Kunchukuttan, Pratyush Kumar, Raj Dabre, and Mitesh M. Khapra. https://arxiv.org/abs/2403.06350 IndicLLMSuite : A Blueprin...
-
[36]
Simran Khanuja, Diksha Bansal, Sarvesh Mehtani, Savya Khosla, Atreyee Dey, Balaji Gopalan, Dilip Kumar Margam, Pooja Aggarwal, Rajiv Teja Nagipogu, Shachi Dave, Shruti Gupta, Subhash Chandra Bose Gali, Vish Subramanian, and Partha Talukdar. 2021. https://arxiv.org/abs/2103.107...
2021 arXiv
-
[37]
Hyunjoong Kim. 2020. https://github.com/lovit/kowikitext Kowikitext: A wikitext format korean corpus . Accessed: August 2024
2020
-
[38]
https://huggingface.co/datasets/azcorpus/azcorpus_v0 azcorpus: The largest open-source nlp corpus for azerbaijani (1.9m documents, 18m sentences)
Huseyn Kishiyev, Jafar Isbarov, Kanan Suleymanli, Khazar Heydarli, Leyla Eminova, and Nijat Zeynalov. https://huggingface.co/datasets/azcorpus/azcorpus_v0 azcorpus: The largest open-source nlp corpus for azerbaijani (1.9m documents, 18m sentences) . Accessed: August 2024
2024
-
[39]
Anoop Kumar et al. 2023. https://arxiv.org/abs/2305.15019 Sangraha: A high-quality, multilingual dataset for indic language pretraining . Accessed: August 2024
2023 arXiv
-
[40]
Khapra, and Pratyush Kumar
Anoop Kunchukuttan, Divyanshu Kakwani, Satish Golla, Gokul N.C., Avik Bhattacharyya, Mitesh M. Khapra, and Pratyush Kumar. 2020. https://arxiv.org/abs/2005.00085 AI4Bharat-IndicNLP Corpus: Monolingual Corpora and Word Embeddings for Indic Languages . arXiv preprint arXiv:2005....
2020 arXiv
-
[41]
Darek Kłeczek. 2020. https://huggingface.co/dkleczek/bert-base-polish-uncased-v1 Polbert: Polish bert language model . Accessed: August 2024
2020
-
[42]
Anne Lauscher, Vinit Ravishankar, Ivan Vuli \'c , and Goran Glava s . 2020. https://doi.org/10.18653/v1/2020.emnlp-main.363 From zero to hero: O n the limitations of zero-shot language transfer with multilingual T ransformers . In Proceedings of the 2020 Conference on Empirica...
2020 doi
-
[43]
Dmitrii Lebedev. 2023. https://www.kaggle.com/datasets/dmitriilebedev/polish-corpus Polish classic literature text corpus . Accessed: August 2024
2023
-
[44]
Chang W. Lee. 2024. https://huggingface.co/datasets/lcw99/wikipedia-korean-20240501 Wikipedia korean dataset (2024-05-01) . Accessed: August 2024
2024
-
[45]
Leipzig Corpora Collection . 2017. https://corpora.uni-leipzig.de?corpusId=tel_community_2017 Telugu community corpus (2017) . Accessed: August 2024
2017
-
[46]
Qianchu Liu, Edoardo Maria Ponti, Diana McCarthy, Ivan Vuli \'c , and Anna Korhonen. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.571 AM 2i C o: Evaluating word meaning in context across low-resource languages with adversarial examples . In Proceedings of the 2021 Confere...
2021 doi
-
[47]
Daniel Loureiro, Kiamehr Rezaee, Mohammad Taher Pilehvar, and Jose Camacho-Collados. 2021. Analysis and evaluation of language models for word sense disambiguation. Computational Linguistics, 47(2):387--443
2021
-
[48]
Federico Martelli, Najla Kalach, Gabriele Tola, Roberto Navigli, et al. 2021. Semeval-2021 task 2: Multilingual and cross-lingual word-in-context disambiguation (mcl-wic). In Proceedings of the 15th International Workshop on Semantic Evaluation (SemEval-2021), pages 24--36
2021
-
[49]
Gati Martin, Medard Edmund Mswahili, Young-Seob Jeong, and Jiyoung Woo. 2022. https://aclanthology.org/2022.naacl-main.23 S wah BERT : Language model of S wahili . In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguist...
2022
-
[50]
Bernard Masua and Noel Masasi. 2024. https://doi.org/10.1016/j.dib.2024.110751 In the heart of swahili: An exploration of data collection methods and corpus curation for natural language processing . Data in Brief, 55:110751
2024
-
[51]
George A. Miller. 1995. https://doi.org/10.1145/219717.219748 WordNet : A lexical database for english . Communications of the ACM, 38(11):39--41
1995
-
[52]
Miller, Claudia Leacock, Randee Tengi, and Ross Bunker
George A. Miller, Claudia Leacock, Randee Tengi, and Ross Bunker. 1993. https://doi.org/10.3115/1075671.1075742 A semantic concordance . In Proceedings of the Workshop on Human Language Technology, pages 303--308. Association for Computational Linguistics
1993
-
[53]
James Thomas Molesworth. 1857. https://dsal.uchicago.edu/dictionaries/molesworth/ A Dictionary, Marathi and English , 2 edition. Printed for government at the Bombay Education Society's Press, Bombay. Accessed: August 2024
2024
-
[54]
National Institute of Korean Language (NIKL) . 2025. https://huggingface.co/datasets/binjang/NIKL-korean-english-dictionary Nikl korean-english dictionary . Dataset available on Hugging Face. Accessed: August 2024
2025
-
[55]
Roberto Navigli. 2009. Word sense disambiguation: A survey. ACM computing surveys (CSUR), 41(2):1--69
2009
-
[56]
Roberto Navigli, David Jurgens, and Daniele Vannella. 2013. https://aclanthology.org/S13-2040/ Semeval-2013 task 12: Multilingual word sense disambiguation . In Second Joint Conference on Lexical and Computational Semantics (*SEM), Volume 2: Proceedings of the Seventh Internat...
2013
-
[57]
Roberto Navigli and Simone Paolo Ponzetto. 2012. https://doi.org/10.1016/j.artint.2012.07.001 BabelNet : The automatic construction, evaluation and application of a wide-coverage multilingual semantic network . Artificial Intelligence, 193:217--250
2012 doi
-
[58]
Roberto Navigli et al. 2023. BabelNet 2023. https://babelnet.org/publications
2023
-
[59]
Quoc Hung Ngo, Werner Winiwarter, and Bartholom \"a us Wloka. 2013. https://aclanthology.org/I13-1001 EVBCorpus - a multi-layer english-vietnamese bilingual corpus for studying tasks in comparative linguistics . In Proceedings of the 11th Workshop on Asian Language Resources (...
2013
-
[60]
Dat Quoc Nguyen and Anh Tuan Nguyen. 2020. https://doi.org/10.18653/v1/2020.findings-emnlp.92 P ho BERT : Pre-trained language models for V ietnamese . In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 1037--1042, Online. Association for Computati...
2020 doi
-
[61]
Nha Van Nguyen. 2025. https://huggingface.co/NlpHUST/ner-vietnamese-electra-base Nlphust/ner-vietnamese-electra-base . Accessed: August 2024
2025
-
[62]
Fred Philippy, Siwen Guo, and Shohreh Haddadan. 2023. https://doi.org/10.18653/v1/2023.acl-long.323 Towards a common understanding of contributing factors for cross-lingual transfer in multilingual language models: A review . In Proceedings of the 61st Annual Meeting of the As...
2023 doi
-
[64]
Mohammad Taher Pilehvar and Jose Camacho-Collados. 2019 b . https://doi.org/10.18653/v1/N19-1128 WiC : the word-in-context dataset for evaluating context-sensitive meaning representations . In Proceedings of the 2019 Conference of the North American Chapter of the Association ...
2019 doi
-
[65]
Telmo Pires, Eva Schlinger, and Dan Garrette. 2019. https://doi.org/10.18653/v1/P19-1493 How multilingual is multilingual BERT ? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4996--5001, Florence, Italy. Association for Compu...
2019 doi
-
[66]
Edoardo Maria Ponti, Ivan Vuli \'c , Goran Glava s , Nikola Mrk s i \'c , and Anna Korhonen. 2018. https://doi.org/10.18653/v1/D18-1026 Adversarial propagation and zero-shot cross-lingual transfer of word vector specialization . In Proceedings of the 2018 Conference on Empiric...
2018 doi
-
[67]
Sameer Pradhan, Edward Loper, Dmitriy Dligach, and Martha Palmer. 2007. https://aclanthology.org/S07-1016/ Semeval-2007 task-17: English lexical sample, srl and all words . In Proceedings of the Fourth International Workshop on Semantic Evaluations (SemEval-2007), pages 87--92...
2007
-
[68]
Alessandro Raganato, Tommaso Pasini, Jose Camacho-Collados, and Mohammad Taher Pilehvar. 2020. Xl-wic: A multilingual benchmark for evaluating semantic contextualization. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 7...
2020
-
[69]
https://doi.org/10.1162/tacl_a_00452 Samanantar: The Largest Publicly Available Parallel Corpora Collection for 11 Indic Languages
Gowtham Ramesh, Sumanth Doddapaneni, Aravinth Bheemaraj, Mayank Jobanputra, Raghavan AK, Ajitesh Sharma, Sujit Sahoo, Harshita Diddee, Mahalakshmi J, Divyanshu Kakwani, Navneet Kumar, Aswin Pradeep, Srihari Nagaraj, Kumar Deepak, Vivek Raghavan, Anoop Kunchukuttan, Pratyush Ku...
-
[70]
Mortensen, and Graham Neubig
Nathaniel Robinson, Perez Ogayo, David R. Mortensen, and Graham Neubig. 2023. https://doi.org/10.18653/v1/2023.wmt-1.40 C hat GPT MT : Competitive for high- (but not low-) resource languages . In Proceedings of the Eighth Conference on Machine Translation, pages 392--418, Sing...
2023 doi
-
[71]
Christoph Rzymski, Tiago Tresoldi, Simon J Greenhill, Mei-Shin Wu, Nathanael E Schweikhard, Maria Koptjevskaja-Tamm, Volker Gast, Timotheus A Bodt, Abbie Hantgan, Gereon A Kaiping, et al. 2020. The database of cross-linguistic colexifications, reproducible analysis of cross-li...
2020
-
[72]
Ali Saeed, Rao Muhammad Adeel Nawab, Mark Stevenson, and Paul Rayson. 2019 a . https://doi.org/10.1145/3314940 A sense annotated corpus for all-words urdu word sense disambiguation . ACM Transactions on Asian and Low-Resource Language Information Processing, 18(4):1--14. Acces...
2019 doi
-
[73]
Ali Saeed, Rao Muhammad Adeel Nawab, Mark Stevenson, and Paul Rayson. 2019 b . https://doi.org/10.1007/s10579-018-9438-7 A word sense disambiguation corpus for urdu . Language Resources and Evaluation, 53(3):397--418. Accessed: August 2024
2019 doi
-
[74]
Uri Shaham, Jonathan Herzig, Roee Aharoni, Idan Szpektor, Reut Tsarfaty, and Matan Eyal. 2024. https://doi.org/10.18653/v1/2024.findings-acl.136 Multilingual instruction tuning with just a pinch of multilinguality . In Findings of the Association for Computational Linguistics:...
2024 doi
-
[75]
Jagbir Singh and Iqbal Singh. 2015. https://doi.org/10.5120/ijca2015906938 Word sense disambiguation: Enhanced lesk approach in punjabi language . International Journal of Computer Applications, 129(6):23--27. Accessed: August 2024
2015 doi
-
[76]
Varinder Pal Singh and Parteek Kumar. 2018. https://doi.org/10.22452/mjcs.vol31no3.2 Naive bayes classifier for word sense disambiguation of punjabi language . Malaysian Journal of Computer Science, 31(3):188--199. Accessed: August 2024
2018 doi
-
[77]
Varinder Pal Singh and Parteek Kumar. 2019. https://doi.org/10.1007/s12046-019-1206-x Sense disambiguation for punjabi language using supervised machine learning techniques . Sādhanā, 44(11):226. Accessed: August 2024
2019 doi
-
[78]
Varinder Pal Singh and Parteek Kumar. 2020. https://doi.org/10.1007/s00521-019-04581-3 Word sense disambiguation for punjabi language using deep learning techniques . Neural Computing and Applications, 32(8):2963--2973. Accessed: August 2024
2020 doi
-
[79]
Anirudh Srinivasan, Sunayana Sitaram, Tanuja Ganu, Sandipan Dandapat, Kalika Bali, and Monojit Choudhury. 2021. Predicting the performance of multilingual nlp models. arXiv preprint arXiv:2110.08875
2021 arXiv
-
[80]
Asahi Ushio, Luis Espinosa Anke, Steven Schockaert, and Jose Camacho-Collados. 2021. https://doi.org/10.18653/v1/2021.acl-long.280 BERT is to NLP what A lex N et is to CV : Can pre-trained language models identify analogies? In Proceedings of the 59th Annual Meeting of the Ass...
2021 doi
-
[81]
Venkatasubbaiah, L
G. Venkatasubbaiah, L. S. Sheshagiri Rao, and H. K. Ramachandra. 1981. https://baraha.com/kannada/browse.php Kannada-Kannada-English Dictionary . Ibh Prakashana, Bangalore, India. Accessed: August 2024
1981
-
[82]
Jason Wei, Maarten Bosma, Vincent Y Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le. 2021. Finetuned language models are zero-shot learners. arXiv preprint arXiv:2109.01652
2021 arXiv
-
[83]
Wikimedia Foundation . 2024. https://dumps.wikimedia.org/tewiki/20240501/ Telugu wikipedia dump . Accessed: August 2024
2024
-
[84]
Wikimedia Foundation . 2025. https://en.wikipedia.org/wiki/List_of_Wikipedias List of wikipedias . Accessed: August 2024
2025
-
[85]
Wydawnictwo Naukowe PWN . 2025. https://sjp.pwn.pl/ Słownik języka polskiego pwn . Accessed: August 2024
2025
-
[86]
Tatu Ylonen. 2022. https://aclanthology.org/2022.lrec-1.140/ Wiktextract: W iktionary as machine-readable structured data . In Proceedings of the Thirteenth Language Resources and Evaluation Conference, pages 1317--1325, Marseille, France. European Language Resources Associati...
2022
-
[87]
Micha Ziemski, Marcin Junczys-Dowmunt, and Bruno Pouliquen. 2016. https://aclanthology.org/L16-1561/ The U nited N ations parallel corpus v1.0 . In Proceedings of the Tenth International Conference on Language Resources and Evaluation ( LREC `16) , pages 3530--3534, Portoro z ...
2016
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.