Pith. sign in

REVIEW 5 major objections 5 minor 52 references

Unsupervised Bilingual Lexicon Induction for Low Resource Languages

T0 review · 5 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read The paper claims that stacking CSCBLI on a linear-transformed UVecMap yields the best unsupervised bilingual lexicon induction for three low-resource language pairs.

desk verdict The combination ranking is undercut by tuning on the test dictionaries, but the new En-Si and En-Pa evaluation sets are worth having. read the letter →

arxiv 2412.16894 v1 pith:IBZRBAM4 submitted 2024-12-22 cs.CL

classification cs.CL
keywords bilinguallexiconinductionunsupervisedBLIlow-resourcelanguagesVecMapwordembeddingscontextualSinhalaTamil
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper asks whether improvements to structure-based unsupervised bilingual lexicon induction (UBLI), developed and tested separately, still help when stacked onto the same framework. On three low-resource pairs – English–Sinhala, English–Tamil, and English–Punjabi – the authors extend the unsupervised VecMap framework (UVecMap) with embedding-creation, pre-processing, and initialization upgrades, including their own PCA-based dimensionality reduction. They report that the best overall accuracy comes from CSCBLI, a method that adds XLM-R contextual representations on top of static embeddings, combined with a linear transformation and UVecMap (combination M17), for both Word2Vec and FastText embeddings. The paper also releases human-curated evaluation dictionaries for English–Sinhala and English–Punjabi, and it treats the question of whether simultaneous use produces equal gains as an empirical one answered by their experiments.

What carries the argument

The central object is the UVecMap framework, an unsupervised self-learning method that iterates two steps: an optimal orthogonal mapping between the source and target embedding spaces, obtained via SVD of $X^T D Z$, and a new dictionary induced over the mapped similarities using CSLS retrieval. Onto this base, the paper stacks three families of extensions: a linear transformation of the input monolingual spaces, based on $n$-th order similarity matrices $M_n(X) = (XX^T)^n$ parameterized by $\alpha$; CSCBLI, which builds a unified word representation by adding contextual offsets from XLM-R through a spring network and then interpolates similarity with a weight $\lambda$; and PCA-based dimensionality reduction applied either as pre-processing or iteratively during initialization. The best recipe, M17, is the specific pipeline where CSCBLI runs on top of a linear-transformed UVecMap alignment.

What would settle it

Build a held-out, human-verified test set for English–Sinhala, English–Tamil, and English–Punjabi that is never touched during alpha or frequency tuning, then re-run M17 and the M1 baseline. If M17's advantage over M1 disappears or reverses on the held-out set, or if a different combination wins there, the paper's central claim that M17 is the best recipe would be falsified.

Watch

Extended reading notes

Core claim

The paper's central claim is that combining multiple independently proposed extensions inside one structure-based UBLI pipeline gives a recipe that outperforms the baseline UVecMap and most individual extensions. In their experiments, the best configurations were M17 (CSCBLI + Linear Transformation + UVecMap) for both Word2Vec and FastText, with M19 and M20 close behind; the one exception was English–Sinhala with Word2Vec, where M3 (linear transformation alone) won. The gains over the baseline are positive but modest, ranging roughly from 0.1 to 3.6 absolute points of precision@1. The paper also reports that embedding fusion consistently hurt performance, and that several combinations on English–Tamil FastText collapsed to near-zero accuracy, showing that stacked extensions can destabilize alignment rather than improve it.

Load-bearing premise

The entire ranking of combinations is computed on evaluation dictionaries that were also used to pick frequency thresholds and alpha values, and for English–Sinhala that dictionary was produced by machine translation with back-translation filtering; if those lists are biased or noisy, the reported best combination may not be the best on genuinely held-out data.

Editorial extensions

If this is right

  • Practitioners working on low-resource language pairs can adopt M17 as a default unsupervised BLI recipe when static Word2Vec or FastText embeddings and XLM-R contextual embeddings are available.
  • The released human-curated dictionaries for English–Sinhala and English–Punjabi give the community evaluation sets for two low-resource languages that previously relied on automatically built lexicons.
  • The frequent collapse to zero accuracy on English–Tamil FastText shows that stacking extensions is not universally safe, so each combination should be validated before use.
  • Because the margins over the baseline are small and the hyperparameters were selected on the same evaluation dictionaries, the reported ranking should be re-checked against a genuinely held-out human dictionary.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The gains of M17 over the baseline are small enough that tuning on the same test dictionaries may be inflating the recipe's apparent advantage; a held-out human-verified test set could reveal whether the stacked pipeline truly generalizes.
  • Since the English–Sinhala evaluation dictionary was created by machine translation with back-translation filtering, its quality ceiling is uncertain, and part of the reported accuracy on that pair may reflect quirks of the translation tool rather than real lexical correspondence.
  • The method's dependence on manually tuned alpha values per language pair suggests a natural extension: learn alpha from intrinsic monolingual statistics, such as the spectral properties of the embedding spaces, to make the recipe applicable to new low-resource languages without additional tuning effort.
  • The three tested pairs all involve English and South Asian languages; testing M17 on typologically more distant or script-divergent pairs would clarify whether the stacked combination is broadly useful or pair-specific.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 5 minor

Summary. The paper studies unsupervised bilingual lexicon induction (UBLI) for three low-resource language pairs (English-Sinhala, English-Tamil, English-Punjabi) by extending the unsupervised VecMap (UVecMap) framework with techniques from prior work: linear transformation, embedding fusion, iterative and effective dimensionality reduction, and CSCBLI (combining static and contextual embeddings). It reports an extensive grid of 20 method combinations and evaluates them with precision@1 on newly created evaluation dictionaries, concluding that CSCBLI with linear transformation over UVecMap (M17) generally performs best. The paper also releases bilingual dictionaries for English-Sinhala and English-Punjabi.

Significance. If the evaluation were sound, the paper would make a useful practical contribution: it systematically compares combinations of previously independently tested UBLI extensions on genuinely low-resource language pairs, and it releases new evaluation dictionaries for English-Sinhala and English-Punjabi, which are valuable resources. The detailed hyperparameter tables are also a useful record. However, the central claim about the best combination is currently not established because the evaluation protocol selects hyperparameters on the same dictionaries used for final scoring, reports no uncertainty estimates, and contains internal inconsistencies in the baseline numbers. The contribution of the dictionaries is real, but the paper's headline conclusion needs a stronger evaluation design.

major comments (5)
  1. [§4.4, §4.4.1, §5, Table 5] The final ranking in Table 5 is in-sample, not an independent test. The alpha values in Table 3 were selected by running UVecMap/CSCBLI on the evaluation dictionaries built in §4.2 (stated in §4.4.1: 'These embeddings were evaluated against the evaluation dictionaries we created'), and the Min_Freq thresholds in §4.4 were likewise chosen through an ablation study whose accuracy columns (Tables 12–14) report pr@1 on those same dictionaries. Consequently, every non-baseline configuration in Table 5 has been tuned to the test set, and the margins over the baseline are not out-of-sample evidence. The problem is compounded by the absence of significance tests or error bars: for example, M17 exceeds M1 by only +0.11 on EnTa Word2Vec and +0.45 on EnPa Word2Vec, while M17 ties M19/M20 on EnPa Word2Vec and EnPa FastText. A held-out split or nested validation, together with variance estimates, is needed before the 'best combination' claim can be supported.
  2. [§5, Table 5] The claim in §5 that 'CSCBLI with Linear Transformation and UVecMap (M17) yielded the best results for both Word2Vec and Fasttext' is contradicted by the paper's own Table 5. For EnSi Word2Vec, M3 (Linear Transformation + UVecMap) achieves 33.18, which is higher than M17's 32.84; the same table also shows M17 tied with M19 on EnPa Word2Vec and with M20 on EnPa FastText. The sentence immediately following acknowledges the EnSi exception, but the generalized conclusion 'best results for both Word2Vec and Fasttext' is therefore not supported. The recommendation should be rephrased to identify a family of competitive configurations rather than a single winner.
  3. [§4.2.1 and Contributions (Abstract, §1)] The paper states as a contribution that it releases 'human-curated bilingual lexicons' for English-Sinhala and English-Punjabi, but §4.2.1 describes the EnSi dictionary as generated by machine translation (Google Translate) followed by back-translation filtering and automatic consistency checks, with no human verification step. Since this dictionary is used both for hyperparameter selection and for the final accuracy numbers, the contradiction matters: machine-translated evaluation data can carry systematic errors that bias the ranking. The authors should either add genuine human curation and document it, or explicitly describe the EnSi dictionary as machine-translated with automatic filtering and discuss the quality ceiling this imposes.
  4. [Appendix B, Tables 12–14; Table 5] The frequency-threshold ablation tables are inconsistent with the baseline results in Table 5. For EnSi, Table 14 with Min_Freq=8 reports Word2Vec=27.48 and FastText=31.49, while Table 5 M1 reports Word2Vec=31.49 and FastText=27.48 — the values are swapped. For EnTa, Table 12 with Min_Freq=6 reports Word2Vec=15.39 and FastText=12.81, while Table 5 M1 reports 16.74 and 11.69. These discrepancies make it impossible to verify which threshold configuration produced the reported baselines, and they undermine the claim that the thresholds were 'carefully selected' as described. Please align the appendix with the final experimental configuration and state explicitly which threshold was used for each language pair and embedding type.
  5. [§5, Table 5] Several combinations collapse to near-zero accuracy on EnTa FastText (e.g., M2=0.11, M4=0, M7=14.16, M10=0, M18=0), and the paper offers only a qualitative explanation that 'the mapping relies heavily on a few key embeddings.' This instability is itself a threat to the reliability of the ranking: small changes in preprocessing can produce total mapping failure, so differences of a few points between surviving configurations may reflect initialization or numerical sensitivity rather than genuine method quality. A robustness analysis (e.g., multiple random seeds or multiple training runs) is needed to establish that the reported ordering is stable.
minor comments (5)
  1. [§4.4.2] The text says dimensionality reduction 'halving the original dimensions' but gives no ablation or justification for choosing 150 (Word2Vec/FastText) and 512 (XLM-R) rather than other target dimensions; a brief explanation or reference would be helpful.
  2. [Table 9] The line 'Received 20 February 2007; revised 12 March 2009; accepted 5 June 2009' appears to be a leftover template artifact and should be removed.
  3. [§3.3] The notation for the spring network weights γ1 and γ2 is introduced as vectors initialized to zero, but the subsequent equations use scalar-looking multiplication; clarifying whether the multiplication is elementwise would improve readability.
  4. [Throughout] The spelling 'Fasttext' and 'FastText' is used inconsistently; please standardize.
  5. [§5] The phrase 'best result is in boldface, second best is in italics' is not consistently applied in Table 5: for EnSi Word2Vec the best is M3 (33.18), and M17 (32.84) is the second best, so it should be italicized; please check all columns.

Circularity Check

2 steps flagged · score 5.0 of 10

Best-recipe claim is partly in-sample: alpha and frequency thresholds are tuned on the same evaluation dictionaries later used to report Table 5 accuracies.

  1. fitted input called prediction [Section 4.4.1 (Linear Transformation Method), Table 3, Appendix A Tables 6-11; final ranking in Section 5 Table 5]
    "Subsequently, we ran the UVecMap model for each pair of αs and αt values to generate cross-lingual word embeddings. These embeddings were evaluated against the evaluation dictionaries we created. ... It appears that CSCBLI with Linear Transformation and UVecMap (M17) yielded the best results for both Word2Vec and Fasttext."

    The same Section 4.2 dictionaries are used for alpha selection and for the final pr@1 numbers in Table 5. Every linear-transformation method (including M17) uses the αs/αt values that were chosen by maximizing pr@1 on those dictionaries, so the comparison between M17 and the baselines is an in-sample selection result rather than an out-of-sample prediction. The winning combination is therefore partly fitted to the evaluation set that is later reported as the result.

  2. fitted input called prediction [Section 4.1 and Section 4.4 (frequency threshold ablation), Appendix B Tables 12-14]
    "We carried out an ablation study to select threshold values for minimum frequencies, which resulted in the minimum frequencies being 8 for EnSi, and 6 for both EnTa and EnPa. The detailed results of the ablation study can be found in Tables 12-14 of Appendix B."

    Tables 12-14 show pr@1 accuracy of UVecMap at each candidate Min_Freq, computed on the same Section 4.2 evaluation dictionaries. The chosen thresholds fix the vocabulary for all subsequent models, and the very same dictionaries are then used to compute the final pr@1 values in Table 5. The frequency threshold is thus a fitted input to the reported results, not an independent evaluation condition.

full rationale

The core unsupervised BLI pipeline is not self-definitional: the mappings are learned from monolingual embeddings, and the evaluation dictionaries are not used during training. There is no load-bearing self-citation chain or imported uniqueness theorem; references to the authors' earlier work on language categorization and corpora are contextual, not argumentative. However, the model-selection loop is circular in a statistical sense. Section 4.4.1 tunes the linear-transformation hyperparameters αs and αt by evaluating UVecMap and CSCBLI against the Section 4.2 dictionaries, and Section 4.4 selects Min_Freq thresholds by an ablation study whose accuracy columns are pr@1 on the same dictionaries. Table 5 then reports pr@1 on those identical dictionaries and selects M17 as the best combination. The final ranking therefore contains no held-out component: the winning configuration has been selected on the test set. This is pattern 2 (fitted input called prediction) in partial form, because the winning method is chosen, not derived, by the tuning procedure. The concern is amplified by the absence of significance tests or variance estimates and by the fact that the EnSi dictionary is machine-translated with back-translation filtering (Section 4.2.1) despite the contribution statement calling the released lexicons human-curated. These evaluation-quality issues mean the small margins in Table 5 are not established as external predictions. The paper is not fully circular, so the score is moderate rather than high.

Assumptions & free parameters 5 free parameters · 4 assumptions · 0 invented entities

No new theoretical entities are introduced; the paper assembles existing components and adds evaluation resources.

free parameters (5)
  • alpha_source and alpha_target (linear transformation) = Table 3, e.g., (0.15, 0.25) for EnSi W2V
    Selected by grid search over {-0.5,...,0.5} on the evaluation dictionaries, maximizing pr@1.
  • Min_Freq threshold = 8 for EnSi, 6 for EnTa and EnPa
    Chosen through an ablation study (Tables 12-14) that measures pr@1 on the same evaluation sets.
  • PCA target dimension = 150 for W2V/FastText, 512 for XLM-R
    Halving of the original dimensions; chosen by hand for computational efficiency.
  • lambda (CSCBLI interpolation weight) = not reported
    Appears in the interpolation formula in Section 3.3 but no value or tuning procedure is given.
  • gamma1/gamma2 spring network weights = learned, not reported
    Initialized as zero vectors and updated by backpropagation; final values not provided.
assumptions (4)
  • domain assumption Monolingual embedding spaces of the source and target languages are approximately isometric
    Invoked in Section 2.1 (initialization via sorted similarity matrices) and Section 3.2 (fusion derivation assumes X = PZO).
  • domain assumption The monolingual corpora (SiTa for EnSi/EnTa, BPCC for EnPa) are clean and representative enough to train useful embeddings
    Section 4.1 relies on these corpora without an independent quality check.
  • domain assumption The evaluation dictionaries are accurate translations (Google Translate filtered output / IndoWordNet filtered pairs)
    Section 4.2 constructs the dictionaries with automatic tools and filtering; their correctness is assumed when scoring all methods.
  • ad hoc to paper Combining the extensions adds improvements without harmful interactions beyond what the grid search reveals
    The paper assumes the 20 combinations are the right space to search; interactions are not modeled or predicted a priori.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Unsupervised Bilingual Lexicon Induction for Low Resource Languages." pith.science (2026). https://pith.science/paper/IBZRBAM4

@misc{pith2026241216894,
  author       = {Pith},
  title        = {Pith review of: Unsupervised Bilingual Lexicon Induction for Low Resource Languages},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/IBZRBAM4}},
  note         = {Machine review of arXiv:2412.16894}
}
read the original abstract

Bilingual lexicons play a crucial role in various Natural Language Processing tasks. However, many low-resource languages (LRLs) do not have such lexicons, and due to the same reason, cannot benefit from the supervised Bilingual Lexicon Induction (BLI) techniques. To address this, unsupervised BLI (UBLI) techniques were introduced. A prominent technique in this line is structure-based UBLI. It is an iterative method, where a seed lexicon, which is initially learned from monolingual embeddings is iteratively improved. There have been numerous improvements to this core idea, however they have been experimented with independently of each other. In this paper, we investigate whether using these techniques simultaneously would lead to equal gains. We use the unsupervised version of VecMap, a commonly used structure-based UBLI framework, and carry out a comprehensive set of experiments using the LRL pairs, English-Sinhala, English-Tamil, and English-Punjabi. These experiments helped us to identify the best combination of the extensions. We also release bilingual dictionaries for English-Sinhala and English-Punjabi.

Figures

Figures reproduced from arXiv: 2412.16894 by the authors.

Figure 1
Figure 1. Schematic diagram of the UVecMap framework [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Improved UVecMap Framework. LN-Length Normalization, MC-Mean Center [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

52 extracted references · 48 canonical work pages

  1. [1]

    Hanan Aldarmaki, Mahesh Mohan, and Mona Diab. 2018. Unsupervised word mapping using structural similarities in monolingual embeddings. Transactions of the Association for Computational Linguistics 6 (2018), 185–196

  2. [2]

    Mikel Artetxe, Gorka Labaka, and Eneko Agirre. 2017. Learning bilingual word embeddings with (almost) no bilingual data. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . 451–462

  3. [3]

    Mikel Artetxe, Gorka Labaka, and Eneko Agirre. 2018. A robust self-learning method for fully unsupervised cross- lingual mappings of word embeddings. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . 789–798

  4. [4]

    Mikel Artetxe, Gorka Labaka, Eneko Agirre, and Kyunghyun Cho. 2018. Unsupervised neural machine translation. In 6th International Conference on Learning Representations, ICLR 2018

  5. [5]

    Mikel Artetxe, Gorka Labaka, In‘igo Lopez-Gazpio, and Eneko Agirre. 2018. Uncovering divergent linguistic information in word embeddings with lessons for intrinsic and extrinsic evaluation. In Proceedings of the 22nd Conference on Computational Natural Language Learning . 282–291

  6. [6]

    Niyati Bafna, Cristina España-Bonet, Josef van Genabith, Benoît Sagot, and Rachel Bawden. 2023. A Simple Method for Unsupervised Bilingual Lexicon Induction for Data-Imbalanced, Closely Related Language Pairs. CoRR (2023)

  7. [7]

    Timothy Baldwin, Jonathan Pool, and Susan Colowick. 2010. PanLex and LEXTRACT: Translating all words of all languages of the world. In Coling 2010: Demonstrations. 37–40

  8. [8]

    Laurent Besacier, Etienne Barnard, Alexey Karpov, and Tanja Schultz. 2014. Automatic speech recognition for under- resourced languages: A survey. Speech communication 56 (2014), 85–100

Show all 52 references
  1. [9]

    Hailong Cao, Liguo Li, Conghui Zhu, Muyun Yang, and Tiejun Zhao. 2023. Dual Word Embedding for Robust Unsupervised Bilingual Lexicon Induction. IEEE/ACM Transactions on Audio, Speech, and Language Processing (2023)

  2. [10]

    Hailong Cao and Tiejun Zhao. 2021. Word Embedding Transformation for Robust Unsupervised Bilingual Lexicon Induction. arXiv:2105.12297v1 (2021)

  3. [11]

    Hailong Cao, Tiejun Zhao, Weixuan Wang, and Wei Peng. 2023. Bilingual word embedding fusion for robust unsuper- vised bilingual lexicon induction. Information Fusion 97 (2023), 101818

  4. [12]

    Hailong Cao, Tiejun Zhao, Shu Zhang, and Yao Meng. 2016. A distribution-based model to learn bilingual word embeddings. In Proceedings of COLING 2016, the 26th International Conference on Computational Linguistics: Technical , Vol. 1, No. 1, Article . Publication date: Decembe...

  5. [13]

    Aditi Chaudhary, Karthik Raman, Krishna Srinivasan, and Jiecao Chen. 2020. Dict-mlm: Improved multilingual pre-training using bilingual dictionaries. arXiv preprint arXiv:2010.12566 (2020)

  6. [14]

    Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzman´, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. Unsupervised Cross-lingual Representation Learning at Scale. In Proceedings of the 58th Annual Meeti...

  7. [15]

    Fathima Farhath, Surangika Ranathunga, Sanath Jayasena, and Gihan Dias. 2018. Integration of bilingual lists for domain-specific statistical machine translation for sinhala-tamil. In 2018 Moratuwa Engineering Research Conference (MERCon). IEEE, 538–543

  8. [16]

    Zihao Feng, Hailong Cao, Tiejun Zhao, Weixuan Wang, and Wei Peng. 2022. Cross-lingual Feature Extraction from Monolingual Corpora for Low-resource Unsupervised Bilingual Lexicon Induction. In Proceedings of the 29th International Conference on Computational Linguistics . 5278–5287

  9. [17]

    Aloka Fernando, Surangika Ranathunga, and Gihan Dias. 2020. Data augmentation and terminology integration for domain-specific sinhala-english-tamil statistical machine translation. arXiv preprint arXiv:2011.02821 (2020)

  10. [18]

    Aloka Fernando, Surangika Ranathunga, Dilan Sachintha, Lakmali Piyarathna, and Charith Rajitha. 2023. Exploiting bilingual lexicons to improve multilingual embedding-based document and sentence alignment for low-resource languages. Knowledge and Information Systems 65, 2 (2023...

  11. [19]

    Jay Gala, Pranjal A.Chitale, Raghavan AK, Varun Gumma, Sumanth Doddapaneni, Aswanth Kumar, Janki Nawale, Anupama Sujatha, Ratish Puduppully, Vivek Raghavan, Pratyush Kumar, Mitesh M.Khapra, Raj Dabre, and Anoop Kunchuhuttan. 2023. IndicTrans2: Towards High-Quality and Accessib...

  12. [20]

    Goran Glavaš, Robert Litschko, Sebastian Ruder, and Ivan Vulić. 2019. How to (Properly) Evaluate Cross-Lingual Word Embeddings: On Strong Baselines, Comparative Analyses, and Some Misconceptions. In Proceedings of the 57th Annual Meeting of the Association for Computational Li...

  13. [21]

    Aria Haghighi, Percy Liang, Taylor Berg-Kirkpatrick, and Dan Klein. 2008. Learning bilingual lexicons from monolingual corpora. In Proceedings of ACL-08: Hlt . 771–779

  14. [22]

    Yedid Hoshen and Lior Wolf. 2018. Non-Adversarial Unsupervised Word Translation. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing . 469–478

  15. [23]

    Pratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali, and Monojit Choudhury. 2020. The State and Fate of Linguistic Diversity and Inclusion in the NLP World. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 6282–6293

  16. [24]

    Armand Joulin, Edouard Grave, Piotr Bojanowski, and Mikolov Tomas. 2017. Bag of Tricks for Efficient Text Classifica- tion. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics . 427–431

  17. [25]

    Alek Keersmaekers, Wouter Mercelis, and Toon Van Hal. 2023. Word Sense Disambiguation for Ancient Greek: Sourcing a training corpus through translation alignment. In Proceedings of the Ancient Language Processing Workshop . 148–159

  18. [26]

    Guillaume Lample and Alexis Conneau. 2019. Cross-lingual language model pretraining. In Proceedings of the Advances in Neural Information Processing Systems . 7059–7069

  19. [27]

    Guillaume Lample, Alexis Conneau, Marc’Aurelio Ranzato, Ludovic Denoyer, and Hervé Jégou. 2018. Word translation without parallel data. In International conference on learning representations

  20. [28]

    Yanyang Li, Yingfeng Luo, Ye Lin, Quan Du, Huizhen Wang, Shujian Huang, Tong Xiao, and Jingbo Zhu. 2020. A Simple and Effective Approach to Robust Unsupervised Bilingual Dictionary Induction. In Proceedings of the 28th International Conference on Computational Linguistics . 5990–6001

  21. [29]

    Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer

  22. [30]

    Anushika Liyanage, Surangika Ranathunga, and Sanath Jayasena. 2021. Bilingual lexical induction for sinhala-english using cross lingual embedding spaces. In 2021 Moratuwa Engineering Research Conference (MERCon) . IEEE, 579–584

  23. [31]

    Stephen Mayhew, Chen-Tse Tsai, and Dan Roth. 2017. Cheap translation for cross-lingual named entity recognition. In Proceedings of the 2017 conference on empirical methods in natural language processing . 2536–2545

  24. [32]

    Zhongtao Miao, Qiyu Wu, Kaiyan Zhao, Zilong Wu, and Yoshimasa Tsuruoka. 2024. Enhancing Cross-lingual Sentence Embedding for Low-resource Languages with Word Alignment. In Findings of the Association for Computational Linguistics: NAACL 2024. 3225–3236

  25. [33]

    Tomas Mikolov, Chen Kai, Greg Corrado, and Jeffrey Dean. 2013. Efficient Estimation of Word Representations in Vector Space. arXiv:1301.3781v3 (2013). , Vol. 1, No. 1, Article . Publication date: December 2024. Unsupervised Bilingual Lexicon Induction for Low Resource Languages 15

  26. [34]

    Tomas Mikolov, Quoc V Le, and Ilya Sutskever. 2013. Exploiting similarities among languages for machine translation. arXiv preprint arXiv:1309.4168 (2013)

  27. [35]

    Idi Mohammed and Rajesh Prasad. 2023. Building lexicon-based sentiment analysis model for low-resource languages. Methods X 11 (2023)

  28. [36]

    Sosuke Nishikawa, Ryokan Ri, and Yoshimasa Tsuruoka. 2021. Data Augmentation with Unsupervised Machine Translation Improves the Structural Similarity of Cross-lingual Word Embeddings. ACL-IJCNLP 2021 (2021), 163

  29. [37]

    Aitor Ormazabal, Mikel Artetxe, Aitor Soroa, Gorka Labaka, and Eneko Agirre. 2021. Beyond Offline Mapping: Learning Cross-lingual Word Embeddings through Context Anchoring. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th ...

  30. [38]

    Ellie Pavlick, Matt Post, Ann Irvine, Dmitry Kachaev, and Chris Callison-Burch. 2014. The language demographics of amazon mechanical turk. Transactions of the Association for Computational Linguistics 2 (2014), 79–92

  31. [39]

    Surangika Ranathunga and Nisansa De Silva. 2022. Some Languages are More Equal than Others: Probing Deeper into the Linguistic Disparity in the NLP World. In Proceedings of the 2nd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the ...

  32. [40]

    Surangika Ranathunga, Nisansa de Silva, Dilith Jayakody, and Aloka Fernando. 2024. Shoulders of Giants: A Look at the Degree and Utility of Openness in NLP Research. arXiv preprint arXiv:2406.06021 (2024)

  33. [41]

    Surangika Ranathunga, Nisansa De Silva, Velayuthan Menan, Aloka Fernando, and Charitha Rathnayake. 2024. Quality Does Matter: A Detailed Look at the Quality and Utility of Web-Mined Parallel Corpora. In Proceedings of the 18th Conference of the European Chapter of the Associat...

  34. [42]

    Shuo Ren, Shujie Liu, Ming Zhou, and Shuai Ma. 2020. A graph-based coarse-to-fine method for unsupervised bilingual lexicon induction. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 3476–3485

  35. [43]

    Haoyue Shi, Luke Zettlemoyer, and Sida I Wang. 2021. Bilingual Lexicon Induction via Unsupervised Bitext Construction and Word Alignment. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on N...

  36. [44]

    Anders Søgaard, Sebastian Ruder, and Ivan Vulić. 2018. On the Limitations of Unsupervised Bilingual Dictionary Induction. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 778–788

  37. [45]

    Ivan Vulić, Anna Korhonen, and Goran Glavaˆs. 2020. Improving bilingual lexicon induction with unsupervised post processing of monolingual word vector spaces. In Proceedings of the 5th Workshop on Representation Learning for NLP . 45–54

  38. [46]

    Ivan Vulić and Marie-Francine Moens. 2015. Monolingual and cross-lingual information retrieval models based on (bilingual) word embeddings. In Proceedings of the 38th international ACM SIGIR conference on research and development in information retrieval. 363–372

  39. [47]

    Kasun Wickramasinghe and Nisansa De Silva. 2023. Sinhala-English Parallel Word Dictionary Dataset. In 2023 IEEE 17th International Conference on Industrial and Information Systems (ICIIS) . IEEE, 61–66

  40. [48]

    Pengcheng Yang, Fuli Luo, Peng Chen, Tianyu Liu, and Xu Sun. 2019. MAAM: A morphology-aware alignment model for unsupervised bilingual lexicon induction. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 3190–3196

  41. [49]

    Zheng-Xin Yong, Cristina Menghini, and Stephen H Bach. 2024. LexC-Gen: Generating Data for Extremely Low- Resource Languages with Large Language Models and Bilingual Lexicons. arXiv preprint arXiv:2402.14086 (2024)

  42. [50]

    Jinpeng Zhang, Baijun Ji, Nini Xiao, Xiangyu Duan, Min Zhang, Yangbin Shi, and Weihua Luo. 2021. Combining Static Word Embeddings and Contextual Representations for Bilingual Lexicon Induction. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021. 2943–2955

  43. [51]

    Meng Zhang, Yang Liu, Huanbo Luan, and Maosong Sun. 2017. Adversarial Training for Unsupervised Bilingual Lexicon Induction. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics . 1959–1970. A Linear Transformation Method Tables 6 - 11 pre...

  44. [2020]

    In Transactions of the Association for Computational Linguistics

    Multilingual Denoising Pre-training for Neural Machine Translation. In Transactions of the Association for Computational Linguistics. 726–742

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.