REVIEW 1 major objections 3 minor 87 references
Indigenous Languages Spoken in Argentina: A Survey of NLP and Speech Resources
T0 review · 1 major / 3 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Argentina's Indigenous languages: seven families, few computational tools
desk verdict A useful country-specific resource survey whose demographic figure overstates what the census can support; worth publishing after revision. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the seven-family classification of Argentina's Indigenous languages, built on prior linguistic surveys (Censabella 1999; Ciccone 2010). The argument runs through a mapping from census community data to ISO 639-3 language codes: for each family, the authors take the self-identified Indigenous population and the percentage who say they speak their community's language, then multiply to get an estimated speaking population. The survey machinery is a protocol that collects papers from major NLP venues and organizes them by family and task. The demographic mapping is what carries the claim about how many speakers each language has, and the family classification is what carries the resource survey.
What would settle it
A community-by-community linguistic survey of, say, the Qom (Toba), Wichí, or Mapuche regions that asks people directly which languages they use would settle whether the census-derived speaking populations are accurate; if the survey found substantially different numbers from the INDEC-based estimates, the paper's demographic claims would need revision.
Extended reading notes
Core claim
The paper's central claim is that Argentina's Indigenous linguistic diversity can be usefully organized into seven language families, each with a characteristic demographic footprint and a specific pattern of computational neglect. For each family it reports a speaking population derived from INDEC's 2022 census, and it argues that these figures are systematically undercounted because the census question design excludes Indigenous language speakers who do not self-identify as members of an Indigenous community, as well as those who speak a language associated with a different community. The survey of NLP and speech resources shows that only a handful of languages—notably Mapudungún, Paraguayan Guaraní, and southern Quechua varieties—have any substantial corpus or model, and almost none of these resources were built for Argentine varieties such as Santiago del Estero Quichua or Corrientes Guaraní.
Load-bearing premise
The speaker estimates assume that the census's self-identification percentages can be mapped directly to specific languages, even though the census never asks which language a person speaks and excludes Indigenous-language speakers who do not self-identify as Indigenous.
Editorial extensions
If this is right
- If the seven-family systematization is adopted, future work has a common reference for which languages are in scope and which Argentine varieties need dedicated attention.
- The finding that most computational resources target non-Argentine varieties means existing models for Quechua or Guaraní may not transfer well to Argentine dialects such as Santiago del Estero Quichua.
- The census undercount argument implies that official figures of roughly 1.3 million Indigenous descendants and 29.3% speakers understate the actual number of Indigenous language speakers, so language-planning decisions based on those numbers may be misinformed.
- For NLP, the survey shows that speech data is especially scarce: only Mapudungún and some Quechua varieties have transcribed audio, while most families have no public corpora at all.
- The authors' call for involving Indigenous communities in evaluation suggests that future resource development should include community-defined usability criteria, not just technical benchmarks.
Reading between the lines
- A direct testable extension would be to change the next national census question to ask which specific Indigenous language a person speaks, which would resolve the mapping ambiguity the paper identifies; this is an inference from the paper's critique, not a claim the authors make.
- The scarcity of resources for Argentine varieties suggests that dialect-identification or cross-variety transfer methods could be a productive research direction, since existing corpora are dominated by Peruvian Quechua and Paraguayan Guaraní.
- Because the paper includes resources for neighboring-country varieties, its coverage hints that regional collaboration—sharing corpora across countries—could fill gaps faster than each country building resources alone, but the authors do not explicitly recommend this.
- The emphasis on written-language techniques while Indigenous languages are predominantly oral implies that speech-based technologies (ASR, TTS, spoken dialogue) may have higher community impact than text NLP; this is an inference from the paper's discussion of oral tradition.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper surveys Indigenous languages spoken in Argentina, organizing them into seven language families (Mapuche, Tupí-Guaraní, Guaycurú, Quechua, Mataco-Mataguaya, Aymara, and Chon), presenting demographic estimates based on the 2022 Argentine census (INDEC), and cataloguing NLP and speech resources for these languages. The demographic portion rests on mapping census self-identification data to specific ISO 639-3 language varieties, while the resource survey follows a described protocol in the appendix and covers both Argentine varieties and related varieties from neighboring countries. The paper concludes that most computational resources target Quechua, Tupí-Guaraní, and Mapuche, often for non-Argentine varieties.
Significance. If the demographic estimates were reliable, the paper would be a valuable reference for researchers working on Indigenous language technology in Argentina. The survey of computational resources is useful and appears carefully compiled: the appendix documents the collection protocol, the tables list corpora and tasks with public-availability markers, and the discussion of challenges (data scarcity, orality, community inclusion) is sensible. The authors are transparent about the census limitation and about the cultural situatedness of the work. However, the demographic contribution is the weaker half of the paper: the central Figure 1 presents exact speaker counts for named language varieties that the census data cannot support in the way they are derived, as detailed below. The resource survey alone would justify a revised version of the paper, but the demographic claims as currently presented are not reliable.
major comments (1)
- [Section 2, Figure 1, Appendix B] The 'Kolla Quechua' row in Figure 1 lists a population of 69,121 and a speaking population of 28,339 for a variety that the paper itself says has no established linguistic distinctiveness: 'We are not aware of any work addressing whether the Kolla Quechua shows peculiarities that justify treating it as a different variety' (Section 2). Appendix B further notes that there is no special ISO code for the variety spoken by the Kolla community. Presenting exact speaker numbers for this category is internally inconsistent with the paper's own caveats and should be corrected, either by removing the category or by clearly labeling it as a census community category rather than a language variety.
minor comments (3)
- [Section 3, first paragraph] The final line of Appendix B contains a stray 'F' character that appears to be a typographical artifact.
- [References] Two references listed as Oncevay (2021a) and Oncevay (2021b) have identical titles ('Peru is multilingual, its machine translation should be too?'); if they are the same work, one should be removed or disambiguated.
- [Figure 1] The left part of the figure shows the distribution of Indigenous communities, but the caption does not explain how the map symbols or the placement of communities correspond to the census data; a brief note on the geographic-coordinate source and the handling of multiple communities per polygon would improve reproducibility.
Circularity Check
No circular derivation: the survey compiles external census and literature data; the speaking-population figures are direct arithmetic on INDEC columns, not a fitted prediction.
full rationale
This paper is a literature and data survey. Its taxonomic claims are attributed to Censabella, Ciccone, Nercesian, and Ethnologue/Glottolog; its demographic numbers are taken from the 2022 INDEC census and multiplied by census-reported speaker percentages to form an estimate explicitly described as such. There are no equations, fitted parameters, or predictive claims whose outputs are defined by their own inputs. The Kolla/Chané/Ava Guaraní mappings and the vertical-centering of undifferentiated community data are validity choices acknowledged in Appendix B ('it is important to mention that not all information ... can be mapped to ISO languages straightforwardly'), and the census self-identification issue is discussed in Section 2; those are data-quality and interpretive limitations, not circularity. No load-bearing self-citation occurs: none of the authors' prior works are invoked to justify the framework. The survey's contribution is organizational and its claims rest on external, citable sources, so the circularity score is zero.
Assumptions & free parameters
assumptions (4)
- domain assumption The 2022 INDEC census provides valid estimates of the number of Indigenous community members who speak their community's language.
- domain assumption The classification of Argentine Indigenous languages into seven families by Censabella (1999) and Ciccone (2010) is correct and comprehensive.
- domain assumption Ethnologue vitality ratings and ISO 639-3 codes reliably represent the status and identity of each language.
- domain assumption The search protocol over selected NLP venues and snowballing captures the majority of relevant computational resources.
Cite this review
Pith. "Pith review of Indigenous Languages Spoken in Argentina: A Survey of NLP and Speech Resources." pith.science (2026). https://pith.science/paper/NC6K6Y5D
@misc{pith2026250109943,
author = {Pith},
title = {Pith review of: Indigenous Languages Spoken in Argentina: A Survey of NLP and Speech Resources},
year = {2026},
howpublished = {\url{https://pith.science/paper/NC6K6Y5D}},
note = {Machine review of arXiv:2501.09943}
}
read the original abstract
Argentina has a large yet little-known Indigenous linguistic diversity, encompassing at least 40 different languages. The majority of these languages are at risk of disappearing, resulting in a significant loss of world heritage and cultural knowledge. Currently, unified information on speakers and computational tools is lacking for these languages. In this work, we present a systematization of the Indigenous languages spoken in Argentina, classifying them into seven language families: Mapuche, Tup\'i-Guaran\'i, Guaycur\'u, Quechua, Mataco-Mataguaya, Aymara, and Chon. For each one, we present an estimation of the national Indigenous population size, based on the most recent Argentinian census. We discuss potential reasons why the census questionnaire design may underestimate the actual number of speakers. We also provide a concise survey of computational resources available for these languages, whether or not they were specifically developed for Argentinian varieties.
Figures
Reference graph
Works this paper leans on
-
[1]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Ezequiel Agustin Adamovsky. 2012. https://doi.org/10.7767/jbla.2012.49.1.343 El color de la nación argentina: conflictos y negociaciones por la definición de un ethnos nacional, de la crisis al bicentenario . Jahrbuch für Geschichte Lateinamerikas
-
[4]
Ife Adebara and Muhammad Abdul-Mageed. 2022. https://aclanthology.org/2022.acl-long.265.pdf Towards Afrocentric NLP for African Languages . Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics, pages 3814--3847
2022
-
[5]
Willem Adelaar. 2010. South america. In Christopher Moseley, editor, Atlas of the world's languages in danger, pages 86--94. UNESCO
2010
-
[6]
Willem Adelaar. 2012. Historical overview: Descriptive and comparative research on south american indian languages. In Ver\'onica Grondona and Lyle Campbell, editors, The Indigenous languages of South America: A comprehensive guide. De Gruyter
2012
-
[7]
u \' sticas del quichua de Santiago del Estero . Actas de las segundas jornadas de ling \
Willem FH Adelaar. 1995. Ra \' ces ling \"u \' sticas del quichua de Santiago del Estero . Actas de las segundas jornadas de ling \"u \' stica aborigen , 15:25--50
1995
-
[8]
Željko Agić and Ivan Vulić. 2019. https://doi.org/10.18653/v1/P19-1310 JW300 : A Wide-Coverage Parallel Corpus for Low-Resource Languages . In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , pages 3204--3210. Association for Computational Linguistics
Show all 87 references
-
[9]
Marvin Agüero-Torales, David Vilares, and Antonio López-Herrera. 2021. https://doi.org/10.18653/v1/2021.calcs-1.12 On the logistical difficulties and findings of Jopara Sentiment Analysis . In Proceedings of the Fifth Workshop on Computational Approaches to Linguistic Code-Swi...
2021 doi
-
[10]
Nouman Ahmed, Natalia Flechas Manrique, and Antonije Petrovi \'c . 2023. https://doi.org/10.18653/v1/2023.americasnlp-1.16 Enhancing S panish- Q uechua machine translation with pre-trained models and diverse data sources: LCT - EHU at A mericas NLP shared task . In Proceedings...
2023 doi
-
[11]
Cristian Ahumada, Claudio Gutierrez, and Antonios Anastasopoulos. 2022. Educational tools for mapuzugun. arXiv preprint arXiv:2205.10411
2022 arXiv
-
[12]
Alham Fikri Aji, Genta Indra Winata, Fajri Koto, Samuel Cahyawijaya, Ade Romadhony, Rahmad Mahendra, Kemal Kurniawan, David Moeljadi, Radityo Eko Prasojo, Timothy Baldwin, et al. 2022. One country, 700+ languages: NLP challenges for underrepresented languages and dialects in I...
2022 arXiv
-
[13]
Honorio Apaza Alanoca, Brisayda Aruhuanca Chahuares, Kewin Aroquipa Caceres, and Josimar Chire Saire. 2023. https://doi.org/10.1007/978-3-031-16075-2_19 Neural Machine Translation for Aymara to Spanish . In Intelligent Systems and Applications , pages 290--298, Cham. Springer ...
2023 doi
-
[14]
Abraham Alvarez-Crespo, Diego Miranda-Salazar, and Willy Ugarte. 2023. Model for R eal- T ime S ubtitling from S panish to Q uechua B ased on C ascade S peech translation. In Proceedings of the 15th International Conference on Agents and Artificial Intelligence (ICAART 2023), ...
2023
-
[15]
Kenneth Beesley. 2003. Finite-State Morphological Analysis and Generation for Aymara . Association for Computational Linguistics
2003
-
[16]
Steven Bird. 2020. https://doi.org/10.18653/v1/2020.coling-main.313 Decolonising Speech and Language Technology . In Proceedings of the 28th International Conference on Computational Linguistics , pages 3504--3519. International Committee on Computational Linguistics
2020 doi
-
[17]
Steven Bird. 2024. https://aclanthology.org/2024.acl-long.797 Must NLP be extractive? In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 14915--14929, Bangkok, Thailand. Association for Computational Linguistics
2024
-
[18]
Verena Blaschke, Hinrich Sch \"u tze, and Barbara Plank. 2023. A survey of corpora for Germanic low-resource languages and dialects . arXiv preprint arXiv:2304.09805
2023 arXiv
-
[19]
Ronald Cardenas, Rodolfo Zevallos, Reynaldo Baquerizo, and Luis Camacho. 2018. Siminchik: A speech corpus for preservation of southern quechua. In In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC'18)
2018
-
[20]
Cintia Carri\'o. 2014. Lenguas en argentina. notas sobre algunos desaf\'ios. In Laura Kornfeld, editor, De lenguas, ficciones y patrias, pages 149--184. Universidad Nacional de General Sarmiento, Buenos Aires
2014
-
[21]
Paulo Cavalin, Pedro Domingues, Julio Nogima, and Claudio Pinhanez. 2023. https://doi.org/10.18653/v1/2023.americasnlp-1.3 Understanding Native Language Identification for Brazilian Indigenous Languages . In Proceedings of the Workshop on Natural Language Processing for Indige...
2023 doi
-
[22]
Marisa Censabella. 1999. https://cir.nii.ac.jp/crid/1130282272168231680 Las Lenguas Indígenas de La Argentina : Una Mirada Actual . Eudeba, Buenos Aires
1999
-
[23]
Marisa In \'e s Censabella. 2009. Chaco ampliado. In Inge Sichra, editor, Atlas socioling\"u\'istico de pueblos ind\'igenas en A m\'erica L atina , pages 143--237. UNICEF
2009
-
[24]
Leonardo Cerno. 2010. Evidencias de diferenciaci \'o n dialectal del guaran \' correntino. Cadernos de Etnoling \"u \' stica , 2(3)
2010
-
[25]
Leonardo Cerno. 2013. El Guaran \' Correntino. Fonolog \' a, Gram \'a tica, Textos . Peter Lang, Frankfurt
2013
-
[26]
Andrés Chandía. 2022. https://aclanthology.org/2022.lrec-1.702 A Mapudüngun FST Morphological Analyser and its Web Interface . In Proceedings of the Thirteenth Language Resources and Evaluation Conference , pages 6540--6547. European Language Resources Association
2022
- [27]
-
[28]
Luis Chiruzzo, Pedro Amarilla, Adolfo Ríos, and Gustavo Giménez Lugo. 2020. https://aclanthology.org/2020.lrec-1.320 Development of a Guarani - Spanish Parallel Corpus . In Proceedings of the Twelfth Language Resources and Evaluation Conference , pages 2629--2633. European Lan...
2020
-
[29]
Luis Chiruzzo, Santiago Góngora, Aldo Alvarez, Gustavo Giménez-Lugo, Marvin Agüero-Torales, and Yliana Rodríguez. 2022. https://aclanthology.org/2022.lrec-1.226 Jojajovai: A Parallel Guarani-Spanish Corpus for MT Benchmarking . In Proceedings of the Thirteenth Language Resourc...
2022
-
[30]
Florencia Ciccone. 2010. Aportes al conocimiento de las lenguas ind\'igenas en A rgentina y su tratamiento desde la E I B . Material elaborated for the Ministerio de Educaci\'on de la Naci\'on
2010
-
[31]
Johanna Cordova, Capucine Boidin, César Itier, Marie-Anne Moreaux, and Damien Nouvel. 2019. https://doi.org/10.1007/978-3-030-11680-4_20 Processing Quechua and Guarani Historical Texts Query Expansion at Character and Word Level for Information Retrieval . In Information Manag...
2019 doi
-
[32]
Ona De Gibert, Ra \'u l V \'a zquez, Mikko Aulamo, Yves Scherrer, Sami Virpioja, and J \"o rg Tiedemann. 2023. https://doi.org/10.18653/v1/2023.americasnlp-1.20 Four Approaches to Low-Resource Multilingual NMT : The Helsinki Submission to the AmericasNLP 2023 Shared Task . In ...
2023 doi
-
[33]
Antonio D \' az-Fern \'a ndez. 2006. Glos \'o nimos aplicados a la lengua mapuche/glossonyms applied to the mapuche language. Anclajes, 10(10):95--111
2006
-
[34]
Wolf Dietrich. 1986. El idioma chiriguano: gram \'a tica, textos, vocabulario . Ediciones de cultura Hisp\'anica, Madrid
1986
-
[35]
Javier Domingo and Dora Manchado. 2018. http://hdl.handle.net/2196/00-0000-0000-0011-F549-B Usos cotidianos del tehuelche (aonekko ‘a’ien) – homenaje a dora manchado. endangered languages archive
2018
-
[36]
Mingjun Duan, Carlos Fasola, Sai Krishna Rallabandi, Rodolfo Vega, Antonios Anastasopoulos, Lori Levin, and Alan W Black. 2020. https://aclanthology.org/2020.lrec-1.350 A Resource for Computational Experiments on Mapudungun . In Proceedings of the Twelfth Language Resources an...
2020
-
[37]
David M Eberhard, Gary Francis Simons, and Charles D Fenning. 2023. Ethnologue: Languages of the world
2023
-
[38]
Ortega, Rolando Coto-solano, Hilaria Cruz, Alexis Palmer, and Katharina Kann
Abteen Ebrahimi, Manuel Mager, Shruti Rijhwani, Enora Rice, Arturo Oncevay, Claudia Baltazar, María Cortés, Cynthia Montaño, John E. Ortega, Rolando Coto-solano, Hilaria Cruz, Alexis Palmer, and Katharina Kann. 2023. https://doi.org/10.18653/v1/2023.americasnlp-1.23 Findings o...
2023 doi
-
[39]
Bruno Estigarribia. 2015. Guaraní-Spanish Jopara mixing in a Paraguayan novel: Does it reflect a third language, a language variety, or true codeswitching? Journal of Language Contact, 8(2):183--222
2015
-
[40]
Leandro Mart \' n Garber. 2022. https://bibliotecadigital.exactas.uba.ar/download/tesis/tesis_n7374_Garber.pdf Sistema de identificaci \'o n de idioma (LID) para grabaciones de entornos naturales biling \"u es en comunidades qom . Ph.D. thesis, Universidad de Buenos Aires
2022
-
[41]
Nat Gillin and Brian Gummibaerhausen. 2023. https://doi.org/10.18653/v1/2023.americasnlp-1.18 Few-shot Spanish-Aymara Machine Translation Using English-Aymara Lexicon . In Proceedings of the Workshop on Natural Language Processing for Indigenous Languages of the Americas ( Ame...
2023 doi
-
[42]
Santiago Góngora, Nicolás Giossa, and Luis Chiruzzo. 2021. https://doi.org/10.18653/v1/2021.americasnlp-1.16 Experiments on a Guarani Corpus of News and Social Media . In Proceedings of the First Workshop on Natural Language Processing for Indigenous Languages of the Americas ...
2021 doi
-
[43]
Santiago Góngora, Nicolás Giossa, and Luis Chiruzzo. 2022. https://doi.org/10.18653/v1/2022.computel-1.16 Can We Use Word Embeddings for Enhancing Guarani-Spanish Machine Translation ? In Proceedings of the Fifth Workshop on the Use of Computational Methods in the Study of End...
2022 doi
-
[44]
Harald Hammarström, Robert Forkel, Martin Haspelmath, and Sebastian Bank. 2023. https://doi.org/10.5281/zenodo.8131084 glottolog/glottolog: Glottolog database 4.8
2023 doi
-
[45]
Marcelo Yuji Himoro and Antonio Pareja-Lora. 2022. https://aclanthology.org/2022.lrec-1.584 Preliminary Results on the Evaluation of Computational Tools for the Analysis of Quechua and Aymara . In Proceedings of the Thirteenth Language Resources and Evaluation Conference , pag...
2022
-
[46]
Instituto Nacional de Estadística y Censos . 2024. https://www.indec.gob.ar/ftp/cuadros/poblacion/censo2022_poblacion_indigena.pdf Censo Nacional de Población, Hogares y Viviendas 2022: población indígena o descendiente de pueblos indígenas u originarios . Buenos Aires, Argentina
2024
-
[47]
Jauhiainen, and Krister Lindén
Tommi Jauhiainen, H. Jauhiainen, and Krister Lindén. 2023. https://www.semanticscholar.org/paper/Tuning-HeLI-OTS-for-Guarani-Spanish-Code-Switching-Jauhiainen-Jauhiainen/452da3a35c5e855c1018a2214d800933a1b389aa Tuning HeLI-OTS for Guarani-Spanish Code Switching Analysis
2023
-
[48]
Mayra Juanatey. 2020. Relaciones entre eventos y referencialidad en quichua santiague\ no: de la gram\'atica al discurso. Ph.D. thesis, Universidad de Buenos Aires, Buenos Aires
2020
-
[49]
Amir Hossein Kargaran, Ayyoob Imani, Fran c ois Yvon, and Hinrich Sch \"u tze. 2023. Glotlid: Language identification for low-resource languages. arXiv preprint arXiv:2310.16248
2023 arXiv
-
[50]
Lorraine Levin, Rodolfo M Vega, Jaime G Carbonell, Ralf D Brown, Alon Lavie, Eliseo Ca \ n ulef, and Carolina Huenchullan. 2002. Data collection and language technologies for mapudungun. In International Workshop on Resources and Tools in Field Linguistics, Las Palmas, Spain
2002
-
[51]
Alexandra Espichán Linares and Arturo Oncevay. 2017. A L ow- R esourced P eruvian L anguage I dentification M odel. In CEUR Workshop Proceedings. CEUR-WS, pages 57--63
2017
-
[52]
Zoey Liu, Crystal Richardson, Richard Hatcher, and Emily Prud'hommeaux. 2022. https://doi.org/10.18653/v1/2022.acl-long.272 Not always about you: Prioritizing community needs when developing endangered language technology . In Proceedings of the 60th Annual Meeting of the Asso...
2022 doi
-
[53]
Ariadna Font Llitj \'o s, Roberto Aranovich, and Lori Levin. 2005. Building M achine T ranslation systems for I ndigenous languages. In Second Conference on the Indigenous Languages of Latin America (CILLA II), Texas, USA
2005
-
[54]
Manuel Mager, Ximena Gutierrez-Vasques, Gerardo Sierra, and Ivan Meza-Ruiz. 2018. https://aclanthology.org/C18-1006 Challenges of language technologies for the indigenous languages of the A mericas . In Proceedings of the 27th International Conference on Computational Linguist...
2018
-
[55]
Manuel Mager, Elisabeth Mager, Katharina Kann, and Ngoc Thang Vu. 2023. https://doi.org/10.18653/v1/2023.acl-long.268 Ethical Considerations for Machine Translation of Indigenous Languages : Giving a Voice to the Speakers . In Proceedings of the 61st Annual Meeting of the Asso...
2023 doi
-
[56]
Lorena Martín Rodríguez, Tatiana Merzhevich, Wellington Silva, Tiago Tresoldi, Carolina Aragon, and Fabrício Gerardi. 2022. Tupian Language Resources : Data , Tools , Analyses
2022
-
[57]
Nelsi Melgarejo, Rodolfo Zevallos, Hector Gomez, and John E. Ortega. 2022. https://aclanthology.org/2022.coling-1.390 WordNet-QU : Development of a Lexical Database for Quechua Varieties . In Proceedings of the 29th International Conference on Computational Linguistics , pages...
2022
-
[58]
Christian Monson, Lori Levin, Rodolfo Vega, Ralf Brown, Ariadna Font Llitjos, Alon Lavie, Jaime G Carbonell, Eliseo Ca \ n ulef, and Rosendo Huisca. 2004. D ata C ollection and A nalysis of M apudungun M orphology for S pelling C orrection
2004
-
[59]
Christian Monson, Ariadna Font Llitj \'o s, Vamshi Ambati, Lorraine Levin, Alon Lavie, Alison Alvarez, Roberto Aranovitch, Jaime G Carbonell, Robert Frederking, Erik Peterson, et al. 2008. Linguistic structure and bilingual informants help induce machine translation of lesser-...
2008
-
[60]
Christian Monson, Ariadna Font Llitj \'o s, Roberto Aranovich, Lori Levin, Ralf Brown, Eric Peterson, Jaime Carbonell, and Alon Lavie. 2006. Building N L P systems for two R esource- S carce I ndigenous L anguages: M apudungun and Q uechua. In Strategies for developing machine...
2006
-
[61]
Christopher Moseley, editor. 2010. Atlas of the World’s Languages in Danger. UNESCO, Paris. 3rd edition
2010
-
[62]
Ver\'onica Nercesian. 2021. Las lenguas del mundo. In La ling\"u\'istica. Una introducci\'on a sus principales preguntas, pages 77--106. Eudeba, Ciudad de Buenos Aires
2021
-
[63]
Arturo Oncevay. 2021 a . Peru is multilingual, its machine translation should be too? In Proceedings of the First Workshop on Natural Language Processing for Indigenous Languages of the Americas, pages 194--201. Association for Computational Linguistics
2021
-
[64]
Arturo Oncevay. 2021 b . https://doi.org/10.18653/v1/2021.americasnlp-1.22 Peru is Multilingual , Its Machine Translation Should Be Too ? In Proceedings of the First Workshop on Natural Language Processing for Indigenous Languages of the Americas , pages 194--201. Association ...
2021 doi
-
[65]
John Ortega and Krishnan Pillaipakkamnatt. 2018. https://aclanthology.org/W18-2201 Using Morphemes from Agglutinative Languages like Quechua and Finnish to Aid in Low-Resource Translation . In Proceedings of the AMTA 2018 Workshop on Technologies for MT of Low Resource Languag...
2018
-
[66]
Ortega, Richard Castro Mamani, and Kyunghyun Cho
John E. Ortega, Richard Castro Mamani, and Kyunghyun Cho. 2020. https://doi.org/10.1007/s10590-020-09255-9 Neural machine translation with a polysynthetic low resource language
2020 doi
-
[67]
Rosa YG Paccotacya-Yanque, Candy A Huanca-Anquise, Judith Escalante-Calcina, Wilber R Ramos-Lov \'o n, and \'A lvaro E Cuno-Parari. 2022. https://doi.org/10.6084/m9.figshare.20292516.v4 A speech corpus of quechua collao for automatic dimensional emotion recognition . Scientifi...
2022 doi
-
[68]
Begoña Pendas, Andres Carvallo, and Carlos Aspillaga. 2023. https://doi.org/10.18653/v1/2023.americasnlp-1.2 Neural Machine Translation through Active Learning on low-resource languages: The case of Spanish to Mapudungun . In Proceedings of the Workshop on Natural Language Pro...
2023 doi
-
[69]
Andr\'es Osvaldo Porta. 2010 a . Un parser para la morfolog \'i a del quichua santiague \ n o con P C - K I M M O . In V \'i ctor M. Castel and y Liliana Cubo de Severino, editors, La renovaci\'on de la palabra en el bicentenario de la Argentina. Los colores de la mirada ling\...
2010
-
[70]
Andr \'e s Osvaldo Porta. 2010 b . The use of formal language models in the typology of the morphology of amerindian languages. In Proceedings of the ACL 2010 Student Research Workshop, pages 109--114
2010
-
[71]
Mónica Quijada. 2004. https://fhaycs-uader.edu.ar/files/2019/catedra_latinoamericano/quijada_-_de_mitos_nacionales_pdf.pdf De Mitos Nacionales, Definiciones Cívicas y Clasificaciones Grupales. Los Índígenas en la Construcción Nacional Argentina, siglos XIX y XX. Calidoscopio l...
2004
-
[72]
Alan Ramponi. 2024. Language varieties of italy: Technology challenges and opportunities. Transactions of the Association for Computational Linguistics, 12:19--38
2024
-
[73]
Andrea Rodrigo, Maximiliano Duran, and María Yanina Nalli. 2021. https://doi.org/10.1007/978-3-030-92861-2_12 Approach to the Automatic Treatment of Gerunds in Spanish and Quechua : A Pedagogical Application . In Formalizing Natural Languages : Applications to Natural Language...
2021 doi
-
[74]
Ríos, Pedro J
Adolfo A. Ríos, Pedro J. Amarilla, and Gustavo A. Giménez Lugo. 2014. https://doi.org/10.1109/BRACIS.2014.18 Sentiment Categorization on a Creole Language with Lexicon-Based and Machine Learning Techniques . In 2014 Brazilian Conference on Intelligent Systems , pages 37--43
2014 doi
-
[75]
Lane Schwartz. 2022. https://doi.org/10.18653/v1/2022.acl-short.82 Primum Non Nocere : Before working with Indigenous data, the ACL must confront ongoing colonialism . In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics ( Volume 2: Short ...
2022 doi
-
[76]
Inge Sichra, editor. 2009. Atlas socioling\"u\'istico de pueblos ind\'igenas en A m\'erica L atina . UNICEF
2009
-
[77]
Nllb Team, Marta R. Costa-juss \`a , James Cross, Onur C elebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Loic Barrault, Gabriel Mejia Gonzalez, ...
2022 arXiv
-
[78]
The Langas Project. 2012. Langas (2012-), corpora diacrónicos de lenguas generales en línea (xvi-xix). Last visited 2024-15-01
2012
-
[79]
Guillaume Thomas. 2019. https://doi.org/10.18653/v1/W19-8008 Universal Dependencies for Mbyá Guaraní . In Proceedings of the Third Workshop on Universal Dependencies ( UDW , SyntaxFest 2019) , pages 70--77. Association for Computational Linguistics
2019 doi
-
[80]
Jörg Tiedemann. 2012. http://www.lrec-conf.org/proceedings/lrec2012/pdf/463_Paper.pdf Parallel Data , Tools and Interfaces in OPUS . In Proceedings of the Eighth International Conference on Language Resources and Evaluation ( LREC '12) , pages 2214--2218. European Language Res...
2012
-
[81]
Atnafu Lambebo Tonja, Fazlourrahman Balouchzahi, Sabur Butt, Olga Kolesnikova, Hector Ceballos, Alexander Gelbukh, and Thamar Solorio. 2024. NLP Progress in Indigenous Latin American Languages . In Proceedings of the 2024 Conference of the North American Chapter of the Associa...
2024
-
[82]
Agüero Torales and Marvin Matías. 2022. https://digibug.ugr.es/handle/10481/72863 Machine Learning Approaches for Topic and Sentiment Analysis in Multilingual Opinions and Low-Resource Languages: From English to Guarani . Universidad de Granada
2022
-
[83]
Ra \'u l V \'a zquez, Yves Scherrer, Sami Virpioja, and J \"o rg Tiedemann. 2021. The helsinki submission to the americasnlp shared task. In Workshop on Natural Language Processing for Indigenous Languages of the Americas, pages 255--264. The Association for Computational Linguistics
2021
-
[84]
Pedro Viegas Barros
J. Pedro Viegas Barros. 1999. Aspectos fon \'e ticos y fonol \'o gicos de la dialectolog \'i a del mapudungun en la A rgentina. In Actas de las III Jornadas de Etnoling\"u\'itica, pages 141--149, Rosario. Universidad Nacional de Rosario, Facultad de Humanidades y Artes, Escuel...
1999
-
[85]
Rodolfo Zevallos, Nuria Bel, Guillermo C \'a mbara, Mireia Farr \'u s, and Jordi Luque. 2022 a . Data augmentation for low-resource quechua asr improvement. arXiv preprint arXiv:2207.06872
2022 arXiv
-
[86]
Rodolfo Zevallos, Luis Camacho, and Nelsi Melgarejo. 2022 b . https://aclanthology.org/2022.lrec-1.537 Huqariq: A Multilingual Speech Corpus of Native Languages of Peru forSpeech Recognition . In Proceedings of the Thirteenth Language Resources and Evaluation Conference , page...
2022
-
[87]
Rodolfo Zevallos, John Ortega, William Chen, Richard Castro, Núria Bel, Cesar Toshio, Renzo Venturas, Aradiel, and Hilario Nelsi Melgarejo. 2022 c . https://doi.org/10.18653/v1/2022.deeplo-1.1 Introducing QuBERT : A Large Monolingual Corpus and BERT Model for Southern Quechua ...
2022 doi
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.