Pith. sign in

REVIEW 3 major objections 6 minor 294 references

Observatorio Lazaro: A self-populating database of anglicism usage in the Spanish press

T0 review · 3 major / 6 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read Observatorio Lázaro is a self-populating, openly queryable monitor of anglicisms in the Spanish press, holding over two million borrowings with quantified precision.

desk verdict A genuinely useful, honestly documented resource that deserves a serious referee; the main caveat is that the type-level statistics apply the current detector's precision profile to the earlier CRF-era data without hedging. read the letter →

arxiv 2608.00713 v1 pith:HD4J5ZSQ submitted 2026-08-01 cs.CL

classification cs.CL
keywords anglicismslexicalborrowingSpanishpressdiachroniccorpusneologymonitoringsequencelabelinglanguageresourcedetection
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper claims that Observatorio Lázaro is a working, continuously updated monitor of unassimilated English borrowings in the Spanish press, and that six years of its output form a reliable diachronic database. It reports more than two million borrowing tokens in 1.88 million articles, with a rate near two anglicisms per thousand tokens that remains stable across the comparable period. The significance is that borrowing dictionaries are static and one-off corpora cannot follow a volatile phenomenon, so a self-populating, openly queryable record would let linguists watch anglicisms enter, spread, and fade in real time. The claim is backed by a held-out F1 of 0.86, inter-annotator agreement of 0.91, and a 1,000-span precision audit whose token-weighted precision is 0.87. The paper also gives a first statistical profile: a steeply skewed, open vocabulary dominated by rare types, concentrated in fashion, technology, and lifestyle sections.

What carries the argument

The load-bearing mechanism is an end-to-end daily pipeline: article retrieval, text cleaning and tokenization, span-level sequence labeling for unassimilated borrowings, lemmatization, and storage in a database served through a public website and API. The central object is the unassimilated lexical borrowing span, a foreign-origin word or multiword expression used in otherwise monolingual Spanish text and not yet adapted to Spanish spelling or morphology. The evaluation rests on a frequency-tiered precision audit: because borrowings follow a Zipfian distribution, the paper samples 1,000 distinct lemmas across five frequency tiers and weights each tier's measured precision by its actual share of types and tokens, which is what yields the contrast between token-weighted precision of 0.87 and type-weighted precision of 0.59. That correction, together with the detector's test-set recall, converts raw detections into population estimates such as the density of 2.14 per thousand tokens.

What would settle it

Take a random sample of complete articles from the stable core period of 2023-2025, annotate every unassimilated borrowing by hand, and compare with what the pipeline stored: if the wild recall is materially below the lab figure of 0.82, the corrected density of 2.14 per thousand tokens understates the true anglicism rate and the flat trend could hide a real increase; if precision on the pre-2022 detector period differs from the audit, the type-level statistics need recomputation.

Watch

Extended reading notes

Core claim

On its own terms, the paper's discovery is that a neural borrowing detector, run daily on a fixed set of news outlets, can produce a valid longitudinal record rather than only a benchmark. The record contains 2,007,647 borrowing tokens across 1,880,377 articles and 993.4 million running tokens from 2020 to 2026. Evaluated on held-out text the detector reaches a span-level F1 of 0.86, and a manual audit of 1,000 frequency-stratified spans gives a token-weighted precision of 0.87 and a type-weighted precision of 0.59. After discounting false positives and scaling by the detector's test-set recall, the paper estimates a true anglicism density of 2.14 per thousand tokens, about one in every 500, and treats it as stable over the single-model window starting in 2023. It further claims that the borrowing vocabulary is an open and growing class: after precision correction, 53.6% of types are attested once, the Heaps exponent is 0.611, and borrowing density varies by more than an order of magnitude across newspaper sections.

Load-bearing premise

Everything rests on assuming the detector makes the same kinds of mistakes out in the wild as it did on the 1,000 hand-checked spans and the lab test set, even though the system only stores sentences where it found something and never measures what it missed.

Editorial extensions

If this is right

  • Diachronic studies of anglicism birth and spread in Spanish can now run on open, continually updated data instead of static dictionaries or hand-annotated snapshots.
  • Frequency and trend analyses can treat the data as nearly benchmark-grade, while rare-type and neology studies must apply the paper's per-tier precision corrections.
  • Longitudinal comparisons should be restricted to the stable core outlets and the single-model window from 2023 onward, because the 2022 detector and outlet changes create a comparability boundary.
  • Lexicographers and language planners gain a candidate-detection feed of new anglicisms with first-attestation dates, contexts, and outlet and section distributions.
  • The measured density of about two unassimilated anglicisms per thousand tokens provides a current baseline for the Spanish press, replacing dated estimates from earlier decades.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A direct consequence of the paper's own recall limitation, worth making explicit: the stable 2023-2025 density is an upper bound on any real decline and a lower bound on any real rise, because a static detector will tend to miss the newest borrowings, so true contemporary usage could be higher than two per thousand.
  • A testable extension of the paper's reasoning is that the sharp section gradient, from about 10.5 borrowings per thousand tokens in fashion to under 0.7 in politics, points to register and domain rather than global language contact as the main driver, which the database's outlet and section fields make directly testable.
  • The same pipeline could be transplanted to other recipient languages or donor languages, but the paper's own finding that non-English borrowings are heavily under-detected suggests such a transplant would need rebalanced training data before its non-English counts could be trusted.
  • Because type-weighted precision is only 0.59, any study of lexical innovation built on this resource will be very sensitive to the precision correction; re-auditing the nonce tier on a larger sample could move the corrected hapax share of 53.6% substantially.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper presents Observatorio Lázaro, a continuously updated automatic monitor of unassimilated lexical borrowings (predominantly anglicisms) in Spanish digital news. It documents the acquisition pipeline, the BiLSTM-CRF detector, the database of 2,007,647 borrowing tokens in 1,880,377 articles over 993.4 million running tokens (2020–2026), and the public web/API access layer. The evaluation includes a held-out span-level F1 of 0.86, an inter-annotator agreement of κ=0.91, and a manual audit of 1,000 stratified spans that yields a token-weighted precision of 0.87, a type-weighted precision of 0.59, and a precision- and recall-corrected density of about 2.14 anglicisms per thousand tokens. Section 6 reports hapax shares (58.7% raw, 53.6% corrected), productivity constants, section-level density contrasts, and temporal stability over the single-model window 2023–2025. The paper explicitly acknowledges that deployed recall is unmeasured and that a mid-2022 detector/outlet change creates a discontinuity.

Significance. If the corrected statistics are supported, this is a valuable resource paper: it offers the first open, daily-updated, diachronic database of anglicism usage in the Spanish press, with a documented pipeline, a public API, a downloadable snapshot, and a rare attempt to quantify deployed precision rather than only lab performance. The paper's frank treatment of unmeasured recall, the 2022 discontinuity, and lemmatization uncertainty is a genuine strength, as is the release of the detector through a pip-installable library and the Zenodo snapshot. However, the headline statistical claims—corrected hapax share, corrected productivity exponents, and the true density estimate—currently rest on extrapolating the precision audit across a detector change and on assuming that test-set recall transfers to deployment. These are fixable but load-bearing issues, which is why I recommend major revision rather than acceptance at this stage.

major comments (3)
  1. [§5.3 / §6.2] Section 5.3 states that the deployed precision audit characterizes only the current BiLSTM-CRF detector (August 2022 onward), covering 79.5% of occurrences, and that the precision of the superseded CRF model for the remaining 20.5% has not been measured. Section 6.2 then applies Table 9's per-tier precisions to the full 2020–2026 database to produce the corrected hapax share of 53.6%, P=0.012, C=0.738, and β=0.611. The CRF was trained on a smaller, headline-only corpus and its error profile may differ, especially in the nonce tier where the current model already has strict precision 0.54. Because these corrected statistics appear in the abstract and in Section 6, the transfer is load-bearing. The authors should either audit a sample of CRF-era spans and recompute the corrections, or restrict all corrected type-level and productivity statistics to the period covered by the audit and report raw counts for the pre-August-2022 portion.
  2. [§5.3, Tables 7–9] The token-weighted precision of 0.87 is computed by weighting per-tier precision values by the tier's share of the token stream, but the audit selected 1,000 spans belonging to 1,000 distinct lemmas, i.e., one occurrence per type. The per-tier precision is therefore a type-level estimate; treating it as an occurrence-level estimate assumes correctness is constant within a lemma across all its occurrences. That assumption is questionable for ambiguous surface forms such as 'horror' or 'look', which are native Spanish words in some contexts and borrowings in others. Since the token-weighted figure is the basis for the corrected density of 2.14 per thousand tokens, the paper should either re-estimate token-weighted precision from an occurrence-level sample or provide evidence that within-lemma variation is negligible.
  3. [§5.3, §4.2, Abstract] Section 5.3 derives the headline estimate of 2.14 anglicisms per thousand tokens by scaling the precision-corrected detection rate by the test-set recall of 0.82, while acknowledging that deployed recall is unmeasured; Section 8 repeats that recall cannot be measured directly on the deployed data. The abstract and Section 4.2 nevertheless present the density as a stable point value without this condition. Because the detector is static and miss rates on borrowings that entered Spanish after training are a plausible source of downward bias, the density should be reported in the abstract and elsewhere as an order-of-magnitude estimate conditional on recall transfer, or as an explicitly labeled lower or upper bound, consistently with the hedging used in Section 6.4.
minor comments (6)
  1. [§1 vs. §4.2/§5.3] Section 1 cites an earlier estimate of around 2% of the vocabulary in El País in 1991, while the paper's own result is about two per thousand tokens (0.2%); please clarify whether the historical figure is 2% of tokens, 2% of types, or 2 per thousand, since the current juxtaposition appears to imply a tenfold decline where the text elsewhere suggests the modern figure is higher.
  2. [Abstract vs. Appendix A] The abstract says monitoring started in April 2020, but Appendix A lists first-seen dates in January and February 2020 for several core outlets; please make the dates consistent.
  3. [§6.2] The sentence 'All four constants fall' immediately follows the reporting of only P, C, and β; if a fourth constant, such as the fitted intercept of the Heaps curve, is intended, it should be named, otherwise the sentence should read 'all three.'
  4. [§4.2 vs. §6.4] Section 4.2 reports an overall density of approximately 2,020 per million tokens, while Section 6.4 reports 1,817 per million on the composition-stable core outlets; because the denominators differ, the paper should state explicitly that the latter is a raw, core-outlet rate so readers do not read the two numbers as inconsistent.
  5. [Table 9] The per-tier precisions in Table 9 are rounded to two decimals, but at the sample sizes in Table 7 they do not always correspond to integer true-positive counts (e.g., 0.63 × 250 = 157.5); please report the raw numerators or add a rounding note for reproducibility.
  6. [§6.2] The corrected values P=0.012, C=0.738, and β=0.611 are reported without the exact procedure by which per-tier precisions were applied to lemma counts and to the token stream; since the correction is described as changing the shape of the accumulation curve rather than merely rescaling it, a worked formula or the analysis code should be provided.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the resource statistics are independent measurements with a manual audit, and prior-work citations are explicit inputs rather than derived conclusions.

full rationale

The paper's central claims are about a deployed database and its measured properties, not about deriving those properties from the detector's design. The precision audit in Section 5.3 is an independent manual review of 1,000 spans sampled from the stored database, and the per-tier precision values are measured on that sample; the corrected density, hapax share, and productivity constants in Sections 4.2 and 6.2 are arithmetic applications of those measured tier precisions to the tier shares of the database, not quantities defined in terms of the conclusions they support. The detector and the coalas corpus are explicitly presented as inputs from prior work (Section 3.4), and the paper states plainly that its contribution is the operational system and the accumulated resource, not the model itself. The cited F1=0.86 and kappa=0.91 come from a separately published, publicly available annotated corpus and model; they are externally evaluable and are not constructed from the resource's own output, so they do not make the reasoning circular even though the authors overlap. The main weaknesses flagged by the paper itself, namely unmeasured deployed recall and the unmeasured precision of the superseded CRF model covering 20.5% of occurrences, are external-validity and measurement-error concerns that the paper explicitly discloses in Sections 5.3 and 8; they are not instances of a prediction being equivalent to its input by definition. No step in the derivation chain reduces an equation or fitted parameter to the claim it is used to support, so the appropriate finding is no circularity.

Assumptions & free parameters 4 free parameters · 4 assumptions · 0 invented entities

The paper does not postulate new physical or linguistic entities. Its free parameters are primarily statistical corrections derived from the audit and fitted distributional exponents, all clearly disclosed. The main implicit assumptions are the transfer of test-set recall to deployment, the representativeness of the 1,000-span audit, and the validity of the unassimilated-borrowing definition.

free parameters (4)
  • Detector recall on deployed data = 0.82 (test-set recall transferred to deployed data)
    Used to correct the raw density of 2.02 to 2.14 per thousand tokens. The paper states recall is unmeasured in deployment and can only be estimated on the curated test set, so applying it to the full database is a fitted assumption.
  • Token-weighted and type-weighted precision corrections = 0.87 token-weighted, 0.59 type-weighted, derived from a 1,000-span stratified sample
    These are used to transform raw counts into corrected counts for hapax share (58.7% to 53.6%), productivity P (0.020 to 0.012) and Heaps exponent beta (0.655 to 0.611). The correction is a statistical estimate, not a fitted parameter in the usual sense, but it depends on the assumption that the stratified sample represents the whole database.
  • Heaps-Herdan exponent beta = 0.655 (V = 5.06 N^0.655, R^2 = 0.999)
    Fitted to cumulative type-token data, used to support the open-class claim. Its value changes to 0.611 after precision correction.
  • Zipfian exponent = -1.365 over ranks 10-5000
    Fitted to the rank-frequency distribution to support heavy-tailedness. The fitted range explicitly excludes the first nine ranks and ranks above 5,000.
assumptions (4)
  • domain assumption The 1,000-span stratified sample is representative of the 68,424-type inventory and the token stream.
    The per-tier precisions are applied to the whole database; if the sampling or tier assignment is biased, all corrected statistics shift.
  • domain assumption Test-set recall (0.82) transfers to deployed newswire.
    Explicitly stated as unverifiable in Section 5.3 and Section 8. The density estimate of 2.14 per thousand tokens depends on it.
  • domain assumption The coalas annotation guidelines and Cohen's kappa of 0.91 define a valid gold standard for 'unassimilated borrowing'.
    All detection, precision audit and derived statistics rest on this operational definition of the phenomenon, which is a contested boundary in contact linguistics.
  • domain assumption Lemmatization by Pattern in English mode groups surface forms correctly enough for type-level statistics.
    The paper states lemmatization has not been validated against human judgment and affects around a tenth of lemma groups.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Observatorio Lazaro: A self-populating database of anglicism usage in the Spanish press." pith.science (2026). https://pith.science/paper/HD4J5ZSQ

@misc{pith2026260800713,
  author       = {Pith},
  title        = {Pith review of: Observatorio Lazaro: A self-populating database of anglicism usage in the Spanish press},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/HD4J5ZSQ}},
  note         = {Machine review of arXiv:2608.00713}
}
read the original abstract

This paper describes Observatorio L\'azaro, a language resource that monitors unassimilated lexical borrowings (predominantly English lexical borrowings or anglicisms) in the Spanish digital press. Since April 2020 the system has automatically processed the daily output of a collection of news outlets, detected borrowings with a neural sequence-labeling model, and made the results available through a public web interface and API. The result is a continuously updated diachronic database which, at the time of writing, records more than two million borrowings across 1.88 million articles and 993 million running tokens of text (2020-2026). The paper documents the resource: we describe the end-to-end pipeline (acquisition, detection, post-processing, storage and access), the data model and the terms of availability; we evaluate the resource through the detector's held-out performance (span-level F1=0.86 for the borrowing class), inter-annotator agreement on the training corpus (Cohen's kappa=0.91) and a manual precision audit of 1,000 spans from the deployed data; and we situate it with respect to Spanish borrowing lexicography, annotated borrowing corpora and neology-monitoring observatories. The data shows that unassimilated anglicisms are used in the Spanish press at a frequency of approximately two anglicisms per thousand tokens, and that this rate remains stable. Our statistical analysis over six years reveals that the anglicism vocabulary in Spanish behaves as an open and growing class, with 58.7% of its types attested only once (53.6% after correcting for detection precision), and that its density is highest in the fashion, technology and lifestyle sections and lowest in political and institutional news. The resource is intended to complement static borrowing dictionaries and one-off annotated corpora by providing a continuously updated record of borrowing in the Spanish press.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

294 extracted references · 70 canonical work pages

  1. [1]

    Semantically

    Ribeiro, Marco Tulio and Singh, Sameer and Guestrin, Carlos , editor =. Semantically. Proceedings of the 56th. 2018 , keywords =. doi:10.18653/v1/P18-1079 , abstract =

  2. [2]

    Proceedings of the 2024

    Ok, Hyunjong and Kil, Taeho and Seo, Sukmin and Lee, Jaeho , editor =. Proceedings of the 2024. 2024 , keywords =. doi:10.18653/v1/2024.naacl-long.427 , abstract =

  3. [3]

    Multicultural

    Loessberg-Zahl, Alexandra , month = jan, year =. Multicultural

  4. [4]

    IEEE Access , author =

    How. IEEE Access , author =. 2022 , note =. doi:10.1109/ACCESS.2022.3157854 , abstract =

  5. [5]

    NLPers , author =

    Doing. NLPers , author =. 2006 , note =

  6. [6]

    Evaluation , url =

    Van Rijsbergen, Cornelis Joost , year =. Evaluation , url =. Information retrieval , publisher =

  7. [7]

    Proceedings of the

    Grossman, Eitan and Eisen, Elad and Nikolaev, Dmitry and Moran, Steven , editor =. Proceedings of the. 2020 , keywords =

  8. [8]

    It takes two to borrow: a donor and a recipient

    Dinu, Liviu and Uban, Ana and Dinu, Anca and Iordache, Ioan-Bogdan and Georgescu, Simona and Zoicas, Laurentiu , editor =. It takes two to borrow: a donor and a recipient. Findings of the. 2024 , keywords =

Show all 294 references
  1. [9]

    Detecting

    Ali, Felermino Dario Mario and Lopes Cardoso, Henrique and Sousa-Silva, Rui , editor =. Detecting. Proceedings of the 2024. 2024 , keywords =

  2. [10]

    Proceedings of the 2023

    Dinu, Liviu and Uban, Ana and Cristea, Alina and Dinu, Anca and Iordache, Ioan-Bogdan and Georgescu, Simona and Zoicas, Laurentiu , editor =. Proceedings of the 2023. 2023 , keywords =. doi:10.18653/v1/2023.emnlp-main.473 , abstract =

  3. [11]

    and List, Johann-Mattis , editor =

    Miller, John E. and List, Johann-Mattis , editor =. Detecting. Proceedings of the 17th. 2023 , keywords =. doi:10.18653/v1/2023.eacl-main.190 , abstract =

  4. [12]

    and Uban, Ana Sabina and Iordache, Ioan-Bogdan and Cristea, Alina Maria and Georgescu, Simona and Zoicas, Laurentiu , editor =

    Dinu, Liviu P. and Uban, Ana Sabina and Iordache, Ioan-Bogdan and Cristea, Alina Maria and Georgescu, Simona and Zoicas, Laurentiu , editor =. Pater. Proceedings of the 2024. 2024 , keywords =

  5. [13]

    Mikušová, Nina , year =

  6. [14]

    International Journal of Applied Linguistics , author =

    In the melting pot of web-crawled texts:. International Journal of Applied Linguistics , author =. 2024 , note =. doi:10.1111/ijal.12485 , abstract =

  7. [15]

    Proceedings of DARPA broadcast news workshop , author =

    Performance. Proceedings of DARPA broadcast news workshop , author =. 1999 , keywords =

  8. [16]

    Parameter-

    Lukichev, Daniil and Kryanina, Darya and Bystrova, Anastasia and Fenogenova, Alena and Tikhonova, Maria , keywords =. Parameter-. Proceedings “

  9. [17]

    FLUMINENSIA : časopis za filološka istraživanja , author =

    A. FLUMINENSIA : časopis za filološka istraživanja , author =. 2023 , note =. doi:10.31820/f.35.2.1 , abstract =

  10. [18]

    doi:10.48550/arXiv.2202.09625 , abstract =

    Chen, Shuguang and Aguilar, Gustavo and Srinivasan, Anirudh and Diab, Mona and Solorio, Thamar , month = feb, year =. doi:10.48550/arXiv.2202.09625 , abstract =

  11. [19]

    Proceedings of the

    Aguilar, Gustavo and Kar, Sudipta and Solorio, Thamar , editor =. Proceedings of the. 2020 , keywords =

  12. [20]

    Proceedings of the

    Muñoz Ortiz, Alberto and Vilares, David , year =. Proceedings of the

  13. [21]

    Muñoz-Ortiz, Alberto and Vilares, David , keywords =

  14. [22]

    Computational Linguistics , author =

    Unsupervised. Computational Linguistics , author =. 2009 , note =. doi:10.1162/coli.08-010-R1-07-048 , number =

  15. [24]

    Tuiteamos o pongamos un tuit?

    Stewart, Ian and Yang, Diyi and Eisenstein, Jacob , editor =. Tuiteamos o pongamos un tuit?. Proceedings of the. 2021 , keywords =

  16. [25]

    Computational Linguistics , author =

    Automatic. Computational Linguistics , author =. 2020 , keywords =. doi:10.1162/coli_a_00361 , abstract =

  17. [26]

    and Georgescu, Simona and Mihai, Mihnea-Lucian and Uban, Ana Sabina , editor =

    Cristea, Alina Maria and Dinu, Liviu P. and Georgescu, Simona and Mihai, Mihnea-Lucian and Uban, Ana Sabina , editor =. Automatic. Findings of the. 2021 , keywords =. doi:10.18653/v1/2021.findings-emnlp.243 , abstract =

  18. [27]

    Detecting loan words computationally , isbn =

    Zhang, Liqin and Manni, Franz and Fabri, Ray and Nerbonne, John , editor =. Detecting loan words computationally , isbn =. Variation. 2021 , doi =

  19. [28]

    Identification of

    Goldberg, Yoav and Elhadad, Michael , editor =. Identification of. Computational. 2008 , doi =

  20. [29]

    Annotation of

    Nevěřilová, Zuzana , editor =. Annotation of. Text,. 2016 , keywords =. doi:10.1007/978-3-319-45510-5_32 , abstract =

  21. [30]

    Unsupervised

    Mansikkaniemi, André and Kurimo, Mikko , editor =. Unsupervised. Proceedings of the. 2012 , keywords =

  22. [31]

    ACM Trans

    Loanword. ACM Trans. Asian Low-Resour. Lang. Inf. Process. , author =. 2020 , keywords =. doi:10.1145/3374212 , abstract =

  23. [32]

    Mi, Chenggang and Yang, Yating and Wang, Lei and Zhou, Xi and Jiang, Tonghai , editor =. A. Proceedings of the. 2018 , keywords =

  24. [33]

    Computer Speech & Language , author =

    Loanword identification based on web resources:. Computer Speech & Language , author =. 2023 , keywords =. doi:10.1016/j.csl.2023.101517 , abstract =

  25. [34]

    Automatic

    Köllner, Marisa , month = aug, year =. Automatic

  26. [35]

    Miller, John and Pariasca, Emanuel and Beltran Castañon, Cesar , editor =. Neural. Proceedings of the. 2021 , keywords =

  27. [36]

    Corpus Linguistics and Linguistic Theory , author =

    Modelling loanword success – a sociolinguistic quantitative study of. Corpus Linguistics and Linguistic Theory , author =. 2020 , note =. doi:10.1515/cllt-2017-0010 , abstract =

  28. [37]

    Language in Society , author =

    Common and uncommon ground:. Language in Society , author =. 1993 , keywords =. doi:10.1017/S0047404500017449 , abstract =

  29. [38]

    Languages , author =

    English-. Languages , author =. 2016 , note =. doi:10.3390/languages1010007 , abstract =

  30. [39]

    Bilingual

    Muysken, Pieter , year =. Bilingual

  31. [40]

    Languages , author =

    Code-. Languages , author =. 2020 , note =. doi:10.3390/languages5020022 , abstract =

  32. [41]

    Bilingualism: Language and Cognition , author =

    Testing the nonce borrowing hypothesis:. Bilingualism: Language and Cognition , author =. 2012 , keywords =. doi:10.1017/S1366728911000381 , abstract =

  33. [42]

    Language , author =

    Constraints on. Language , author =. 1979 , note =. doi:10.2307/412586 , abstract =

  34. [43]

    Sometimes

    Poplack, Shana , month = jan, year =. Sometimes. Linguistics , volume =. doi:10.1515/ling.1980.18.7-8.581 , abstract =

  35. [44]

    , editor =

    Thomason, Sarah G. , editor =. Social factors and linguistic processes in the emergence of stable mixed languages , volume =. The. 2003 , note =

  36. [45]

    Detecting

    Alvarez-Mellado, Elena and Lignos, Constantine , editor =. Detecting. Proceedings of the 60th. 2022 , keywords =. doi:10.18653/v1/2022.acl-long.268 , abstract =

  37. [46]

    Procesamiento del Lenguaje Natural , author =

    Overview of. Procesamiento del Lenguaje Natural , author =. 2021 , keywords =

  38. [47]

    Alvarez Mellado, Elena , month = may, year =. An. Proceedings of the

  39. [48]

    Pugh, Robert and Tyers, Francis , year =. The. Proceedings of the

  40. [49]

    Revista Signos

    Configuración lingüística de anglicismos procedentes de. Revista Signos. Estudios de Lingüística , author =. 2018 , note =

  41. [50]

    Anglicisms and

    Martí Solano, Ramón and Ruano San Segundo, Pablo , month = mar, year =. Anglicisms and

  42. [51]

    Multitask

    Pritzen, Julia and Gref, Michael and Zühlke, Dietlind and Schmidt, Christoph Andreas , editor =. Multitask. Proceedings of the. 2022 , keywords =

  43. [52]

    , month = jun, year =

    Ashok, Dhananjay and Lipton, Zachary C. , month = jun, year =

  44. [53]

    Heigold, Georg and Varanasi, Stalin and Neumann, Günter and van Genabith, Josef , editor =. How. Proceedings of the 13th. 2018 , keywords =

  45. [54]

    Empirical

    Namysl, Marcin and Behnke, Sven and Köhler, Joachim , editor =. Empirical. Findings of the. 2021 , keywords =. doi:10.18653/v1/2021.findings-acl.27 , urldate =

  46. [55]

    KONVENS 2016, Ruhr-University Bochum , author =

    What to do about non-standard (or non-canonical) language in. KONVENS 2016, Ruhr-University Bochum , author =. 2016 , keywords =

  47. [56]

    Findings of the

    Sainz, Oscar and Campos, Jon and García-Ferrero, Iker and Etxaniz, Julen and de Lacalle, Oier Lopez and Agirre, Eneko , editor =. Findings of the. 2023 , keywords =. doi:10.18653/v1/2023.findings-emnlp.722 , abstract =

  48. [57]

    Stanislawek, Tomasz and Wróblewska, Anna and Wójcicka, Alicja and Ziembicki, Daniel and Biecek, Przemyslaw , editor =. Named. Proceedings of the 23rd. 2019 , keywords =. doi:10.18653/v1/K19-1058 , abstract =

  49. [58]

    and Nenkova, Ani , month = jan, year =

    Agarwal, Oshin and Yang, Yinfei and Wallace, Byron C. and Nenkova, Ani , month = jan, year =. Entity-. doi:10.48550/arXiv.2004.04123 , abstract =

  50. [59]

    Proceedings of the 59th

    Wang, Xiao and Liu, Qin and Gui, Tao and Zhang, Qi and Zou, Yicheng and Zhou, Xin and Ye, Jiacheng and Zhang, Yongxin and Zheng, Rui and Pang, Zexiong and Wu, Qinzhuo and Li, Zhengyan and Zhang, Chong and Ma, Ruotian and Fei, Zichu and Cai, Ruijian and Zhao, Jun and Hu, Xingwu...

  51. [60]

    Lin, Hongyu and Lu, Yaojie and Tang, Jialong and Han, Xianpei and Sun, Le and Wei, Zhicheng and Yuan, Nicholas Jing , editor =. A. Proceedings of the 2020. 2020 , keywords =. doi:10.18653/v1/2020.emnlp-main.592 , abstract =

  52. [61]

    Transactions of the Association for Computational Linguistics , author =

    Context-aware. Transactions of the Association for Computational Linguistics , author =. 2021 , keywords =. doi:10.1162/tacl_a_00386 , abstract =

  53. [63]

    Ma, Ruotian and Wang, Xiaolei and Zhou, Xin and Zhang, Qi and Huang, Xuanjing , editor =. Towards. Proceedings of the 2023. 2023 , keywords =. doi:10.18653/v1/2023.emnlp-main.281 , abstract =

  54. [64]

    Patterns , author =

    Data and its (dis)contents:. Patterns , author =. 2021 , note =. doi:10.1016/j.patter.2021.100336 , language =

  55. [65]

    Proceedings of the 2024

    Rueda, Andrew and. Proceedings of the 2024. 2024 , keywords =

  56. [66]

    Advances in

    Wang, Alex and Pruksachatkun, Yada and Nangia, Nikita and Singh, Amanpreet and Michael, Julian and Hill, Felix and Levy, Omer and Bowman, Samuel , year =. Advances in

  57. [67]

    Naik, Aakanksha and Ravichander, Abhilasha and Sadeh, Norman and Rose, Carolyn and Neubig, Graham , editor =. Stress. Proceedings of the 27th. 2018 , keywords =

  58. [68]

    and Davis, Ernest and Morgenstern, Leora , month = jun, year =

    Levesque, Hector J. and Davis, Ernest and Morgenstern, Leora , month = jun, year =. The. Proceedings of the

  59. [69]

    1996 , keywords =

    Using the framework , author =. 1996 , keywords =

  60. [70]

    Diachronica , author =

    On detecting borrowing:. Diachronica , author =. 2003 , note =. doi:10.1075/dia.20.2.04min , abstract =

  61. [71]

    Automatic

    Zaitsev, Konstantin and Minchenko, Anzhelika , editor =. Automatic. Proceedings of the first workshop on. 2022 , keywords =

  62. [72]

    Nath, Abhijnan and Mahdipour Saravani, Sina and Khebour, Ibrahim and Mannan, Sheikh and Li, Zihui and Krishnaswamy, Nikhil , editor =. A. Proceedings of the 29th. 2022 , keywords =

  63. [73]

    Terminàlia , author =

    Garbell: l’avaluador automàtic de neologismes catalans (. Terminàlia , author =. 2022 , note =

  64. [74]

    Evaluating

    Amigó, Enrique and Delgado, Agustín , editor =. Evaluating. Proceedings of the 60th. 2022 , pages =. doi:10.18653/v1/2022.acl-long.399 , abstract =

  65. [75]

    Information Retrieval , author =

    A comparison of extrinsic clustering evaluation metrics based on formal constraints , volume =. Information Retrieval , author =. 2009 , keywords =. doi:10.1007/s10791-008-9066-8 , abstract =

  66. [76]

    Sebastiani, Fabrizio , month = sep, year =. An. Proceedings of the 2015. doi:10.1145/2808194.2809449 , abstract =

  67. [77]

    Information Retrieval Journal , author =

    Evaluation measures for quantification: an axiomatic approach , volume =. Information Retrieval Journal , author =. 2020 , keywords =. doi:10.1007/s10791-019-09363-y , abstract =

  68. [78]

    Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2019) , author =

    Automatic. Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2019) , author =

  69. [79]

    Automatic

    Marimon, Montserrat and Gonzalez-Agirre, Aitor and Intxaurrondo, Ander and Martin, Jose Antonio Lopez and Villegas, Marta , year =. Automatic

  70. [80]

    Phonotactics as an

    Mæhlum, Petter and Ivanova, Sardana , editor =. Phonotactics as an. Proceedings of the. 2023 , keywords =

  71. [81]

    Parameter-

    Lukichev, Daniil and Kryanina, Darya and Bystrova, Anastasia and Fenogenova, Alena and Tikhonova, Maria , month = jun, year =. Parameter-. doi:10.28995/2075-7182-2023-22-295-306 , abstract =

  72. [82]

    Machine vs

    Shaitarova, Anastassia and Göhring, Anne and Volk, Martin , editor =. Machine vs. Proceedings of the 24th. 2023 , keywords =

  73. [83]

    Chinchor, Nancy , year =. Fourth

  74. [84]

    Bagga, Amit and Baldwin, Breck , month = aug, year =. Entity-. 36th. doi:10.3115/980845.980859 , urldate =

  75. [85]

    ACM Transactions on Asian and Low-Resource Language Information Processing , author =

    Improving the. ACM Transactions on Asian and Low-Resource Language Information Processing , author =. 2023 , keywords =. doi:10.1145/3572773 , abstract =

  76. [86]

    Zheng, Jonathan and Ritter, Alan and Xu, Wei , month = mar, year =

  77. [87]

    Transactions of the Association for Computational Linguistics , author =

    Data. Transactions of the Association for Computational Linguistics , author =. 2018 , note =. doi:10.1162/tacl_a_00041 , abstract =

  78. [88]

    Zhou, Lexin and Moreno-Casares, Pablo A. and Martínez-Plumed, Fernando and Burden, John and Burnell, Ryan and Cheke, Lucy and Ferri, Cèsar and Marcoci, Alexandru and Mehrbakhsh, Behzad and Moros-Daval, Yael and hÉigeartaigh, Seán Ó and Rutar, Danaja and Schellaert, Wout and Vo...

  79. [89]

    Procesamiento del Lenguaje Natural , author =

    Overview of. Procesamiento del Lenguaje Natural , author =. 2023 , keywords =

  80. [90]

    doccano:

    Nakayama, Hiroki and Kubo, Takahiro and Kamura, Junya and Taniguchi, Yasufumi and Liang, Xu , year =. doccano:

  81. [91]

    Spanish , booktitle =

    Rodríguez González, Félix , editor =. Spanish , booktitle =. 2002 , note =

  82. [92]

    Foundations and Trends in Machine Learning , author =

    An introduction to conditional random fields , volume =. Foundations and Trends in Machine Learning , author =. 2012 , note =

  83. [93]

    Ortografía de la lengua española , publisher =

  84. [94]

    Lemario y palabras nuevas en la edición 23.3 del

    Rodríguez Alberich, Gabriel , year =. Lemario y palabras nuevas en la edición 23.3 del

  85. [95]

    Okazaki, Naoaki , year =

  86. [96]

    Honnibal, Matthew and Montani, Ines , year =

  87. [97]

    python-crfsuite , author =

  88. [98]

    Cañete, José , month = may, year =. Spanish

  89. [99]

    Pérez, Jorge , year =

  90. [100]

    Atlantis , author =

    Anglicisms in contemporary. Atlantis , author =. 1999 , note =

  91. [101]

    Sarker, Sagor , year =

  92. [102]

    Cardellino, Cristian , month = aug, year =. Spanish

  93. [103]

    Diálogo de la Lengua, IX , author =

    An up-to-date review of the literature on. Diálogo de la Lengua, IX , author =. 2017 , pages =

  94. [104]

    Text chunking using transformation-based learning , booktitle =

    Ramshaw, Lance A and Marcus, Mitchell P , year =. Text chunking using transformation-based learning , booktitle =

  95. [105]

    Hofland, Knut , month = may, year =. A. Proceedings of the

  96. [106]

    Learning

    Grave, Edouard and Bojanowski, Piotr and Gupta, Prakhar and Joulin, Armand and Mikolov, Tomas , year =. Learning. Proceedings of the

  97. [107]

    Transactions of the Association for Computational Linguistics , author =

    Enriching. Transactions of the Association for Computational Linguistics , author =. 2017 , pages =

  98. [108]

    The retrieval of false anglicisms in newspaper texts , booktitle =

    Furiassi, Cristiano and Hofland, Knut , year =. The retrieval of false anglicisms in newspaper texts , booktitle =

  99. [109]

    Revista signos , author =

    Configuración lingüística de anglicismos procedentes de. Revista signos , author =. 2018 , note =

  100. [110]

    Los anglicismos de frecuencia sintácticos en español: estudio empírico , journal =

    Rodríguez Medina, María Jesús , year =. Los anglicismos de frecuencia sintácticos en español: estudio empírico , journal =

  101. [111]

    A data-driven approach to anglicism identification in

    Losnegaard, Gyri Smordal and Lyse, Gunn Inger , editor =. A data-driven approach to anglicism identification in. Exploring. 2012 , pages =

  102. [112]

    Revista española de lingüística aplicada , author =

    Email or correo electrónico?. Revista española de lingüística aplicada , author =. 2012 , note =

  103. [113]

    International Journal of English Studies , author =

    Towards a corpus-based analysis of anglicisms in. International Journal of English Studies , author =. 2009 , pages =

  104. [114]

    Language design: journal of theoretical and experimental linguistics , author =

    Anglicisms in. Language design: journal of theoretical and experimental linguistics , author =. 2016 , pages =

  105. [115]

    International Journal of English Studies , author =

    A reassessment of traditional lexicographical tools in the light of new corpora: sports. International Journal of English Studies , author =. 2011 , pages =

  106. [116]

    Revista signos , author =

    Neología sintagmática anglicada en español:. Revista signos , author =. 2018 , note =

  107. [117]

    Journal of Artificial Intelligence Research , author =

    Cross-lingual bridges with models of lexical borrowing , volume =. Journal of Artificial Intelligence Research , author =. 2016 , pages =

  108. [118]

    Colombian Applied Linguistics Journal , author =

    Anglicism:. Colombian Applied Linguistics Journal , author =. 2014 , note =

  109. [119]

    Newly-coined

    Oncíns Martínez, José Luis , editor =. Newly-coined. The anglicization of. 2012 , pages =

  110. [120]

    Using foreign inclusion detection to improve parsing performance , booktitle =

    Alex, Beatrice and Dubey, Amit and Keller, Frank , year =. Using foreign inclusion detection to improve parsing performance , booktitle =

  111. [121]

    Analecta Malacitana (AnMal electrónica) , author =

    A corpus-based study of. Analecta Malacitana (AnMal electrónica) , author =. 2018 , note =

  112. [122]

    Revista Canaria de Estudios Ingleses , author =

    Typographical,. Revista Canaria de Estudios Ingleses , author =. 2017 , note =

  113. [123]

    Proposing a pragmatic distinction for lexical

    Winter-Froemel, Esme and Onysko, Alexander , editor =. Proposing a pragmatic distinction for lexical. The anglicization of. 2012 , pages =

  114. [124]

    Al-Badrashiny, Mohamed and Diab, Mona , month = nov, year =. The. doi:10.18653/v1/W16-5813 , booktitle =

  115. [125]

    Codeswitching language identification using

    Xia, Meng Xuan , month = nov, year =. Codeswitching language identification using. doi:10.18653/v1/W16-5818 , booktitle =

  116. [126]

    Codeswitching

    Shrestha, Prajwol , month = nov, year =. Codeswitching. doi:10.18653/v1/W16-5816 , booktitle =

  117. [127]

    Language

    Sikdar, Utpal Kumar and Gambäck, Björn , month = nov, year =. Language. doi:10.18653/v1/W16-5817 , booktitle =

  118. [128]

    Multilingual

    Samih, Younes and Maharjan, Suraj and Attia, Mohammed and Kallmeyer, Laura and Solorio, Thamar , month = nov, year =. Multilingual. doi:10.18653/v1/W16-5806 , booktitle =

  119. [129]

    Shirvani, Rouzbeh and Piergallini, Mario and Gautam, Gauri Shankar and Chouikha, Mohamed , month = nov, year =. The. doi:10.18653/v1/W16-5815 , booktitle =

  120. [130]

    , month = nov, year =

    Jaech, Aaron and Mulcaire, George and Ostendorf, Mari and Smith, Noah A. , month = nov, year =. A. doi:10.18653/v1/W16-5807 , booktitle =

  121. [131]

    Semi-automatic approaches to

    Andersen, Gisle , editor =. Semi-automatic approaches to. The anglicization of. 2012 , pages =

  122. [132]

    Aguilar, Gustavo and AlGhamdi, Fahad and Soto, Victor and Diab, Mona and Hirschberg, Julia and Solorio, Thamar , month = jul, year =. Named. Proceedings of the. doi:10.18653/v1/W18-3219 , abstract =

  123. [133]

    Automatic detection of

    Alex, Beatrice , year =. Automatic detection of

  124. [134]

    El anglicismo en el español peninsular contemporáneo , volume =

    Pratt, Chris , year =. El anglicismo en el español peninsular contemporáneo , volume =

  125. [135]

    Anglicismos hispánicos , publisher =

    Lorenzo, Emilio , year =. Anglicismos hispánicos , publisher =

  126. [136]

    Onomázein , author =

    El anglicismo léxico en el discurso económico de divulgación científica del español de. Onomázein , author =. 2004 , note =

  127. [137]

    and Gimeno Menéndez, M.V

    Gimeno Menéndez, F. and Gimeno Menéndez, M.V. , year =. El desplazamiento lingüístico del español por el inglés , isbn =

  128. [138]

    , year =

    Gómez Capuz, J. , year =. Los préstamos del español: lengua y sociedad , publisher =

  129. [139]

    Revista alicantina de estudios ingleses , author =

    Towards a typological classification of linguistic borrowing (illustrated with anglicisms in. Revista alicantina de estudios ingleses , author =. 1997 , note =

  130. [140]

    El anglicismo en el español actual , publisher =

    Medina López, Javier , year =. El anglicismo en el español actual , publisher =

  131. [141]

    International Journal of Bilingualism , author =

    Using distributional semantics in loanword research:. International Journal of Bilingualism , author =. 2017 , note =

  132. [142]

    Epos: Revista de filología , author =

    A. Epos: Revista de filología , author =. 2018 , note =

  133. [143]

    Dynamics of language contact:

    Clyne, Michael and Clyne, Michael G and Michael, Clyne , year =. Dynamics of language contact:

  134. [144]

    Constraint-

    Tsvetkov, Yulia and Ammar, Waleed and Dyer, Chris , month = may, year =. Constraint-. doi:10.3115/v1/N15-1062 , booktitle =

  135. [145]

    Linguistics , author =

    The. Linguistics , author =. 1988 , note =

  136. [146]

    Code-switching or borrowing?

    Lipski, John M , year =. Code-switching or borrowing?. Selected proceedings of the second workshop on

  137. [147]

    Anglicisms in

    Onysko, Alexander , year =. Anglicisms in

  138. [148]

    Anglicismos en la prensa económica española , school =

    Vélez Barreiro, Marco , year =. Anglicismos en la prensa económica española , school =

  139. [149]

    Bilingualism: Language and Cognition , author =

    What does the nonce borrowing hypothesis hypothesize? , volume =. Bilingualism: Language and Cognition , author =. 2012 , note =

  140. [150]

    The impact of

    Patzelt, Carolin , year =. The impact of. Multilingual

  141. [151]

    Applying corpus and computational methods to loanword research: new approaches to

    Serigos, Jacqueline Rae Larsen , year =. Applying corpus and computational methods to loanword research: new approaches to

  142. [152]

    Language Variation and Change , author =

    Myths and facts about loanword development , volume =. Language Variation and Change , author =. 2012 , note =

  143. [153]

    Linguistics , author =

    Predicting new words from newer words:. Linguistics , author =. 2010 , pages =. doi:10.1515/ling.2010.043 , number =

  144. [154]

    The anglicization of

    Furiassi, Cristiano and Pulcini, Virginia and Rodríguez González, Félix , year =. The anglicization of

  145. [155]

    English in

    Görlach, Manfred , year =. English in

  146. [156]

    Empirical Approaches to Language Typology , author =

    Loanword typology:. Empirical Approaches to Language Typology , author =. 2008 , note =

  147. [157]

    Language contact, creolization, and genetic linguistics , publisher =

    Thomason, Sarah Grey and Kaufman, Terrence , year =. Language contact, creolization, and genetic linguistics , publisher =

  148. [158]

    International Journal of Lexicography , author =

    Prescriptivism and descriptivism in the treatment of anglicisms in a series of bilingual. International Journal of Lexicography , author =. 2011 , note =

  149. [159]

    Transactions of the Association for Computational Linguistics , author =

    Named entity recognition with bidirectional. Transactions of the Association for Computational Linguistics , author =. 2016 , note =

  150. [160]

    Loanwords in the world's languages: a comparative handbook , publisher =

    Haspelmath, Martin and Tadmor, Uri , year =. Loanwords in the world's languages: a comparative handbook , publisher =

  151. [161]

    What to do about non-standard (or non-canonical) language in

    Plank, Barbara , year =. What to do about non-standard (or non-canonical) language in. Proceedings of the 13th

  152. [162]

    Language Resources and Evaluation , author =

    An unsupervised method for identifying loanwords in. Language Resources and Evaluation , author =. 2015 , note =

  153. [163]

    Yang, Jie and Liang, Shuailong and Zhang, Yue , year =. Design. Proceedings of the 27th

  154. [164]

    Journal of French Language Studies , author =

    Lexical borrowings in. Journal of French Language Studies , author =. 2010 , note =

  155. [165]

    Cognitive Linguistics , author =

    Cognitive. Cognitive Linguistics , author =. 2012 , note =

  156. [166]

    Automatic detection of anglicisms for the pronunciation dictionary generation: a case study on our

    Leidig, Sebastian and Schlippe, Tim and Schultz, Tanja , year =. Automatic detection of anglicisms for the pronunciation dictionary generation: a case study on our. Spoken

  157. [167]

    Grammatical borrowing in cross-linguistic perspective , volume =

    Matras, Yaron and Sakel, Jeanette , year =. Grammatical borrowing in cross-linguistic perspective , volume =

  158. [168]

    The Hague: Mouton , author =

    Languages in. The Hague: Mouton , author =

  159. [169]

    Language , author =

    The analysis of linguistic borrowing , volume =. Language , author =. 1950 , note =

  160. [170]

    23.4 , author =

    Diccionario de la lengua española, ed. 23.4 , author =

  161. [171]

    Colorado Research in Linguistics , author =

    A. Colorado Research in Linguistics , author =

  162. [172]

    Andamios , author =

    La lexicografía del español y el español hispanoamericano , volume =. Andamios , author =. 2014 , note =

  163. [173]

    Nueva revista de filología hispánica , author =

    Americanismo frente a españolismo lingüísticos , volume =. Nueva revista de filología hispánica , author =. 1995 , note =

  164. [174]

    Language variation and change , author =

    Myths and facts about loanword development , volume =. Language variation and change , author =. 2012 , note =

  165. [175]

    Computational Linguistics , author =

    Inter-coder agreement for computational linguistics , volume =. Computational Linguistics , author =. 2008 , note =

  166. [177]

    Wang, Shuhe and Sun, Xiaofei and Li, Xiaoya and Ouyang, Rongbin and Wu, Fei and Zhang, Tianwei and Li, Jiwei and Wang, Guoyin , month = oct, year =

  167. [178]

    Memorization vs

    Elangovan, Aparna and He, Jiayuan and Verspoor, Karin , editor =. Memorization vs. Proceedings of the 16th. 2021 , keywords =. doi:10.18653/v1/2021.eacl-main.113 , abstract =

  168. [179]

    Robustness

    Goel, Karan and Rajani, Nazneen Fatema and Vig, Jesse and Taschdjian, Zachary and Bansal, Mohit and Ré, Christopher , editor =. Robustness. Proceedings of the 2021. 2021 , keywords =. doi:10.18653/v1/2021.naacl-demos.6 , abstract =

  169. [180]

    Computer Speech & Language , author =

    Generalisation in named entity recognition:. Computer Speech & Language , author =. 2017 , keywords =. doi:10.1016/j.csl.2017.01.012 , abstract =

  170. [181]

    Decomposed

    Ma, Tingting and Jiang, Huiqiang and Wu, Qianhui and Zhao, Tiejun and Lin, Chin-Yew , month = apr, year =. Decomposed

  171. [182]

    ner and pos when nothing is capitalized , url =

    Mayhew, Stephen and Tsygankova, Tatiana and Roth, Dan , editor =. ner and pos when nothing is capitalized , url =. Proceedings of the 2019. 2019 , keywords =. doi:10.18653/v1/D19-1650 , abstract =

  172. [183]

    Proceedings of the 29th

    Malmasi, Shervin and Fang, Anjie and Fetahu, Besnik and Kar, Sudipta and Rokhlenko, Oleg , editor =. Proceedings of the 29th. 2022 , pages =

  173. [184]

    Discontinuous

    Vilares, David and Gómez-Rodríguez, Carlos , editor =. Discontinuous. Proceedings of the 2020. 2020 , pages =. doi:10.18653/v1/2020.emnlp-main.221 , abstract =

  174. [185]

    Identifying expressions of opinion in context , abstract =

    Breck, Eric and Choi, Yejin and Cardie, Claire , year =. Identifying expressions of opinion in context , abstract =. Proceedings of the 20th international joint conference on

  175. [186]

    Identifying expressions of opinion in context , abstract =

  176. [187]

    Language Resources and Evaluation , author =

    Annotating. Language Resources and Evaluation , author =. 2005 , keywords =. doi:10.1007/s10579-005-7880-9 , abstract =

  177. [188]

    Part-of-

    Schmid, Helmut , month = aug, year =. Part-of-

  178. [189]

    Brill, Eric , year =. A. Speech and

  179. [190]

    Ratnaparkhi, Adwait , year =. A. Conference on

  180. [191]

    Overview for the

    Molina, Giovanni and AlGhamdi, Fahad and Ghoneim, Mahmoud and Hawwari, Abdelati and Rey-Villamizar, Nicolas and Diab, Mona and Solorio, Thamar , editor =. Overview for the. Proceedings of the. 2016 , pages =. doi:10.18653/v1/W16-5805 , urldate =

  181. [192]

    Overview for the

    Solorio, Thamar and Blair, Elizabeth and Maharjan, Suraj and Bethard, Steven and Diab, Mona and Ghoneim, Mahmoud and Hawwari, Abdelati and AlGhamdi, Fahad and Hirschberg, Julia and Chang, Alison and Fung, Pascale , editor =. Overview for the. Proceedings of the. 2014 , pages =...

  182. [193]

    and Koprinska, Irena and Honnibal, Matthew , editor =

    O'Keefe, Timothy and Pareti, Silvia and Curran, James R. and Koprinska, Irena and Honnibal, Matthew , editor =. A. Proceedings of the 2012. 2012 , pages =

  183. [194]

    Proceedings of the 15th

    Pavlopoulos, John and Sorensen, Jeffrey and Laugier, Léo and Androutsopoulos, Ion , editor =. Proceedings of the 15th. 2021 , pages =. doi:10.18653/v1/2021.semeval-1.6 , abstract =

  184. [195]

    Da San Martino, Giovanni and Yu, Seunghak and Barrón-Cedeño, Alberto and Petrov, Rostislav and Nakov, Preslav , editor =. Fine-. Proceedings of the 2019. 2019 , pages =. doi:10.18653/v1/D19-1565 , abstract =

  185. [196]

    Proceedings of the

    Ben Jannet, Mohamed and Adda-Decker, Martine and Galibert, Olivier and Kahn, Juliette and Rosset, Sophie , editor =. Proceedings of the

  186. [197]

    Zhong, Xiaoshi and Cambria, Erik , year =. Time. Proceedings of the 2018. doi:10.1145/3178876.3185997 , language =

  187. [198]

    Ratinov, Lev and Roth, Dan , editor =. Design. Proceedings of the. 2009 , keywords =

  188. [199]

    Lingvisticæ Investigationes , author =

    A survey of named entity recognition and classification , volume =. Lingvisticæ Investigationes , author =. 2007 , keywords =. doi:https://doi.org/10.1075/li.30.1.03nad , language =

  189. [200]

    Ding, Ning and Xu, Guangwei and Chen, Yulin and Wang, Xiaobin and Han, Xu and Xie, Pengjun and Zheng, Haitao and Liu, Zhiyuan , editor =. Few-. Proceedings of the 59th. 2021 , pages =. doi:10.18653/v1/2021.acl-long.248 , abstract =

  190. [201]

    Extended

    Sekine, Satoshi and Sudo, Kiyoshi and Nobata, Chikashi , editor =. Extended. Proceedings of the

  191. [202]

    Goldberg, Yoav , month = mar, year =. Two

  192. [203]

    Proceedings of the 59th

    Fu, Jinlan and Huang, Xuanjing and Liu, Pengfei , editor =. Proceedings of the 59th. 2021 , pages =. doi:10.18653/v1/2021.acl-long.558 , abstract =

  193. [204]

    Findings of the

    Katz, Uri and Vetzler, Matan and Cohen, Amir and Goldberg, Yoav , editor =. Findings of the. 2023 , keywords =. doi:10.18653/v1/2023.findings-emnlp.218 , abstract =

  194. [205]

    IEEE TRANSACTIONS ON KNOWLEDGE AND DATA ENGINEERING , author =

    A. IEEE TRANSACTIONS ON KNOWLEDGE AND DATA ENGINEERING , author =

  195. [206]

    Katiyar, Arzoo and Cardie, Claire , editor =. Nested. Proceedings of the 2018. 2018 , pages =. doi:10.18653/v1/N18-1079 , abstract =

  196. [207]

    Ohta, Tomoko and Tateisi, Yuka and Kim, Jin-Dong , month = mar, year =. The. Proceedings of the second international conference on

  197. [208]

    Ohta, Tomoko and Tateisi, Yuka and Kim, Jin-Dong , year =. The. Proceedings of the second international conference on. doi:10.3115/1289189.1289260 , abstract =

  198. [209]

    2014 , keywords =

    Journal of Biomedical Informatics , author =. 2014 , keywords =. doi:10.1016/j.jbi.2013.12.006 , abstract =

  199. [210]

    Pradhan, Sameer and Moschitti, Alessandro and Xue, Nianwen and Ng, Hwee Tou and Björkelund, Anders and Uryupina, Olga and Zhang, Yuchen and Zhong, Zhi , editor =. Towards. Proceedings of the. 2013 , pages =

  200. [211]

    2021 , note =

    npj Systems Biology and Applications , author =. 2021 , note =. doi:10.1038/s41540-021-00200-x , abstract =

  201. [212]

    Keyphrase

    Mu, Funan and Yu, Zhenting and Wang, LiFeng and Wang, Yequan and Yin, Qingyu and Sun, Yibo and Liu, Liqun and Ma, Teng and Tang, Jing and Zhou, Xing , month = feb, year =. Keyphrase. doi:10.48550/arXiv.2002.05407 , abstract =

  202. [213]

    ACM Transactions on Knowledge Discovery from Data , author =

    Nested. ACM Transactions on Knowledge Discovery from Data , author =. 2022 , pages =. doi:10.1145/3522593 , abstract =

  203. [214]

    , editor =

    Finkel, Jenny Rose and Manning, Christopher D. , editor =. Nested. Proceedings of the 2009. 2009 , pages =

  204. [216]

    Bridging the

    Rohanian, Omid and Taslimipoor, Shiva and Kouchaki, Samaneh and Ha, Le An and Mitkov, Ruslan , editor =. Bridging the. Proceedings of the 2019. 2019 , pages =. doi:10.18653/v1/N19-1275 , abstract =

  205. [217]

    Natural Language Engineering , author =

    Focus of negation:. Natural Language Engineering , author =. 2021 , note =. doi:10.1017/S1351324920000388 , abstract =

  206. [218]

    Negation

    Jiménez Zafra, Salud María , month = jun, year =. Negation

  207. [219]

    Egyptian Informatics Journal , author =

    The impact of using different annotation schemes on named entity recognition , volume =. Egyptian Informatics Journal , author =. 2021 , keywords =. doi:10.1016/j.eij.2020.10.004 , abstract =

  208. [220]

    Representing and

    Lapponi, Emanuele and Read, Jonathon and Ovrelid, Lilja , month = dec, year =. Representing and. 2012. doi:10.1109/ICDMW.2012.23 , abstract =

  209. [221]

    Learning the

    Morante, Roser and Liekens, Anthony and Daelemans, Walter , editor =. Learning the. Proceedings of the 2008. 2008 , pages =

  210. [223]

    Unsupervised

    Giannakopoulos, Athanasios and Musat, Claudiu and Hossmann, Andreea and Baeriswyl, Michael , editor =. Unsupervised. Proceedings of the 8th. 2017 , pages =. doi:10.18653/v1/W17-5224 , abstract =

  211. [224]

    Unsupervised

    Fusco, Francesco and Staar, Peter and Antognini, Diego , editor =. Unsupervised. Proceedings of the 2022. 2022 , keywords =. doi:10.18653/v1/2022.emnlp-industry.1 , abstract =

  212. [225]

    Unsupervised

    Dowlagar, Suman and Mamidi, Radhika , editor =. Unsupervised. Proceedings of the 17th. 2020 , keywords =

  213. [226]

    Programming and Computer Software , author =

    Methods for automatic term recognition in domain-specific text collections:. Programming and Computer Software , author =. 2015 , keywords =. doi:10.1134/S036176881506002X , abstract =

  214. [227]

    IEEE journal of biomedical and health informatics , author =

    A. IEEE journal of biomedical and health informatics , author =. 2022 , pmid =. doi:10.1109/JBHI.2021.3123192 , abstract =

  215. [228]

    Procesamiento del Lenguaje Natural , author =

    Negation. Procesamiento del Lenguaje Natural , author =. 2021 , pages =

  216. [229]

    Computational Linguistics , author =

    Modality and. Computational Linguistics , author =. 2012 , note =. doi:10.1162/COLI_a_00095 , number =

  217. [230]

    IEEE Transactions on Neural Networks and Learning Systems , author =

    A. IEEE Transactions on Neural Networks and Learning Systems , author =. 2022 , note =. doi:10.1109/TNNLS.2022.3213168 , abstract =

  218. [231]

    Proceedings of the AAAI Conference on Artificial Intelligence , author =

    Neural. Proceedings of the AAAI Conference on Artificial Intelligence , author =. 2017 , note =. doi:10.1609/aaai.v31i1.10995 , abstract =

  219. [232]

    Dataset and

    Li, Peng and Li, Wei and He, Zhengyan and Wang, Xuguang and Cao, Ying and Zhou, Jie and Xu, Wei , month = sep, year =. Dataset and

  220. [233]

    Zhou, Xiaoqiang and Hu, Baotian and Chen, Qingcai and Tang, Buzhou and Wang, Xiaolong , editor =. Answer. Proceedings of the 53rd. 2015 , pages =. doi:10.3115/v1/P15-2117 , urldate =

  221. [234]

    Computational Linguistics , author =

    Survey:. Computational Linguistics , author =. 2017 , note =. doi:10.1162/COLI_a_00302 , abstract =

  222. [235]

    End-to-end learning of semantic role labeling using recurrent neural networks , url =

    Zhou, Jie and Xu, Wei , editor =. End-to-end learning of semantic role labeling using recurrent neural networks , url =. Proceedings of the 53rd. 2015 , pages =. doi:10.3115/v1/P15-1109 , urldate =

  223. [236]

    He, Luheng and Lee, Kenton and Lewis, Mike and Zettlemoyer, Luke , editor =. Deep. Proceedings of the 55th. 2017 , pages =. doi:10.18653/v1/P17-1044 , abstract =

  224. [237]

    Peng, Fuchun and Feng, Fangfang and McCallum, Andrew , month = aug, year =. Chinese

  225. [238]

    Chen, Xinchi and Qiu, Xipeng and Zhu, Chenxi and Liu, Pengfei and Huang, Xuanjing , editor =. Long. Proceedings of the 2015. 2015 , pages =. doi:10.18653/v1/D15-1141 , urldate =

  226. [239]

    Elephant:

    Evang, Kilian and Basile, Valerio and Chrupa. Elephant:. Proceedings of the 2013. 2013 , pages =

  227. [240]

    Strzyz, Michalina and Vilares, David and Gómez-Rodríguez, Carlos , editor =. Viable. Proceedings of the 2019. 2019 , pages =. doi:10.18653/v1/N19-1077 , abstract =

  228. [241]

    Constituent

    Gómez-Rodríguez, Carlos and Vilares, David , editor =. Constituent. Proceedings of the 2018. 2018 , pages =. doi:10.18653/v1/D18-1162 , abstract =

  229. [242]

    He, Zhiyong and Wang, Zanbo and Wei, Wei and Feng, Shanshan and Mao, Xianling and Jiang, Sheng , month = nov, year =. A. doi:10.48550/arXiv.2011.06727 , abstract =

  230. [243]

    Gehrmann, Sebastian and Adewumi, Tosin and Aggarwal, Karmanya and Ammanamanchi, Pawan Sasanka and Aremu, Anuoluwapo and Bosselut, Antoine and Chandu, Khyathi Raghavi and Clinciu, Miruna-Adriana and Das, Dipanjan and Dhole, Kaustubh and Du, Wanyu and Durmus, Esin and Dušek, Ond...

  231. [244]

    Gehrmann, Sebastian and Adewumi, Tosin and Aggarwal, Karmanya and Ammanamanchi, Pawan Sasanka and Aremu, Anuoluwapo and Bosselut, Antoine and Chandu, Khyathi Raghavi and Clinciu, Miruna-Adriana and Das, Dipanjan and Dhole, Kaustubh and Du, Wanyu and Durmus, Esin and Dušek, Ond...

  232. [245]

    ruder.io, A blog about natural language processing and machine learning

    Challenges and. ruder.io, A blog about natural language processing and machine learning. , author =. 2021 , keywords =

  233. [246]

    Yuan, Jun and Vig, Jesse and Rajani, Nazneen , month = mar, year =. 27th. doi:10.1145/3490099.3511146 , abstract =

  234. [247]

    Errudite:

    Wu, Tongshuang and Ribeiro, Marco Tulio and Heer, Jeffrey and Weld, Daniel , editor =. Errudite:. Proceedings of the 57th. 2019 , keywords =. doi:10.18653/v1/P19-1073 , abstract =

  235. [248]

    Comparisons of sequence labeling algorithms and extensions , isbn =

    Nguyen, Nam and Guo, Yunsong , month = jun, year =. Comparisons of sequence labeling algorithms and extensions , isbn =. Proceedings of the 24th international conference on. doi:10.1145/1273496.1273582 , abstract =

  236. [249]

    Dissecting

    Papay, Sean and Klinger, Roman and Padó, Sebastian , editor =. Dissecting. Proceedings of the 2020. 2020 , keywords =. doi:10.18653/v1/2020.emnlp-main.396 , abstract =

  237. [250]

    Phonotactics as an

    Mæhlum, Petter and Ivanova, Sardana , editor =. Phonotactics as an. Proceedings of the. 2023 , pages =

  238. [251]

    Robustness to

    Bodapati, Sravan and Yun, Hyokun and Al-Onaizan, Yaser , month = nov, year =. Robustness to. Proceedings of the 5th. doi:10.18653/v1/D19-5531 , abstract =

  239. [252]

    IEEE Transactions on Knowledge and Data Engineering , author =

    Entity. IEEE Transactions on Knowledge and Data Engineering , author =. 2015 , keywords =. doi:10.1109/TKDE.2014.2327028 , abstract =

  240. [253]

    Romanica Olomucensia , author =

    Phraseological neoforms from the. Romanica Olomucensia , author =. 2023 , keywords =. doi:10.5507/ro.2023.002 , abstract =

  241. [254]

    Yu, Juntao and Bohnet, Bernd and Poesio, Massimo , month = may, year =. Neural. Proceedings of the

  242. [256]

    Computational Linguistics , author =

    A. Computational Linguistics , author =. 2002 , keywords =. doi:10.1162/089120102317341756 , abstract =

  243. [257]

    Vivek, Rajan and Ethayarajh, Kawin and Yang, Diyi and Kiela, Douwe , month = sep, year =. Anchor. doi:10.48550/arXiv.2309.08638 , abstract =

  244. [258]

    Context-aware

    Chen, Shuguang and Neves, Leonardo and Solorio, Thamar , month = sep, year =. Context-aware. doi:10.48550/arXiv.2309.08999 , abstract =

  245. [259]

    Pretrained

    Yates, Andrew and Nogueira, Rodrigo and Lin, Jimmy , month = jun, year =. Pretrained. Proceedings of the 2021. doi:10.18653/v1/2021.naacl-tutorials.1 , abstract =

  246. [260]

    Prompting

    Blevins, Terra and Gonen, Hila and Zettlemoyer, Luke , month = jul, year =. Prompting. Proceedings of the 61st. doi:10.18653/v1/2023.acl-long.367 , abstract =

  247. [261]

    2021 , keywords =

    Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track , author =. 2021 , keywords =

  248. [262]

    Proceedings of the eLex 2015 conference , author =

    Combining a rule-based approach and machine learning in a good-example extraction task for the purpose of lexicographic work on contemporary standard. Proceedings of the eLex 2015 conference , author =. 2015 , keywords =

  249. [263]

    Predicting corpus example quality via supervised machine learning , abstract =. Proc. Electronic Lexicography in the 21st Century Conference (eLex) , author =. 2015 , keywords =

  250. [264]

    Proceedings of the Electronic Lexicography in the 21st Century Conference , author =

    Using a. Proceedings of the Electronic Lexicography in the 21st Century Conference , author =. 2015 , keywords =

  251. [265]

    Rule-based and machine learning approaches for second language sentence-level readability , url =

    Pilán, Ildikó and Volodina, Elena and Johansson, Richard , month = jun, year =. Rule-based and machine learning approaches for second language sentence-level readability , url =. Proceedings of the. doi:10.3115/v1/W14-1821 , urldate =

  252. [266]

    International Journal of Lexicography , author =

    Identification and automatic extraction of good dictionary examples: the case(s) of. International Journal of Lexicography , author =. 2019 , keywords =. doi:10.1093/ijl/ecy014 , number =

  253. [267]

    Complexity , author =

    Automated. Complexity , author =. 2021 , note =. doi:10.1155/2021/2553199 , abstract =

  254. [268]

    International Journal of Bilingualism , author =

    The. International Journal of Bilingualism , author =. 2023 , note =. doi:10.1177/13670069231168535 , abstract =

  255. [269]

    and Matwin, Stan , editor =

    Nadeau, David and Turney, Peter D. and Matwin, Stan , editor =. Unsupervised. Advances in. 2006 , keywords =. doi:10.1007/11766247_23 , abstract =

  256. [270]

    Segura-Bedmar, Isabel and Martínez, Paloma and Herrero-Zazo, María , month = jun, year =. Second

  257. [271]

    Proceedings of the

    Daza, Daniel and Cochez, Michael and Groth, Paul , month = may, year =. Proceedings of the. doi:10.18653/v1/2022.spnlp-1.4 , abstract =

  258. [272]

    Intercultural communication studies , author =

    Language. Intercultural communication studies , author =. 2002 , keywords =

  259. [273]

    What do we really know about

    Vajjala, Sowmya and Balasubramaniam, Ramya , month = jun, year =. What do we really know about. Proceedings of the

  260. [274]

    International Journal of Bilingualism , author =

    Language contact phenomena in multiword units:. International Journal of Bilingualism , author =. 2023 , note =. doi:10.1177/13670069231190209 , abstract =

  261. [275]

    Humanities and Social Sciences Communications , author =

    Tracking the acceptance of neologisms in. Humanities and Social Sciences Communications , author =. 2023 , note =. doi:10.1057/s41599-023-01977-4 , abstract =

  262. [276]

    Wintner, Shuly and Shehadi, Safaa and Zeira, Yuli and Osmelak, Doreen and Nov, Yuval , month = aug, year =. Shared. doi:10.48550/arXiv.2308.15209 , abstract =

  263. [277]

    Enhancing

    Esuli, Andrea and Sebastiani, Fabrizio , editor =. Enhancing. Human. 2011 , keywords =. doi:10.1007/978-3-642-20095-3_46 , abstract =

  264. [278]

    Sentence-

    Esuli, Andrea and Marcheggiani, Diego and Sebastiani, Fabrizio , year =. Sentence-. Proceedings of the

  265. [279]

    Lester, Brian , month = nov, year =. iobes:. Proceedings of. doi:10.18653/v1/2020.nlposs-1.16 , abstract =

  266. [280]

    Training

    Suzuki, Jun and McDermott, Erik and Isozaki, Hideki , month = jul, year =. Training. Proceedings of the 21st. doi:10.3115/1220175.1220203 , urldate =

  267. [281]

    Evaluating

    Esuli, Andrea and Sebastiani, Fabrizio , editor =. Evaluating. Multilingual and. 2010 , doi =

  268. [282]

    Evaluating

    Esuli, Andrea and Sebastiani, Fabrizio , editor =. Evaluating. Multilingual and. 2010 , keywords =. doi:10.1007/978-3-642-15998-5_12 , abstract =

  269. [283]

    The truth of the

    Sasaki, Yutaka , year =. The truth of the

  270. [284]

    ACM Computing Surveys , author =

    A review of the. ACM Computing Surveys , author =. 2023 , pages =. doi:10.1145/3606367 , abstract =

  271. [285]

    , month = sep, year =

    Cleverdon, Cyril W. , month = sep, year =. The significance of the. Proceedings of the 14th annual international. doi:10.1145/122860.122861 , urldate =

  272. [286]

    Acoustic vowel analysis using multilevel regression models:

    Bäumler, Linda and Hartmann, Frederik , editor =. Acoustic vowel analysis using multilevel regression models:. Corpus. 2023 , doi =

  273. [287]

    Proceedings of the

    Ben Jannet, Mohamed and Adda-Decker, Martine and Galibert, Olivier and Kahn, Juliette and Rosset, Sophie , month = may, year =. Proceedings of the

  274. [288]

    Jannet, Mohamed Ameur Ben and Adda-Decker, Martine and Galibert, Olivier and Kahn, Juliette and Rosset, Sophie , keywords =

  275. [289]

    Cognition , author =

    Cognitive influences in language evolution:. Cognition , author =. 2019 , keywords =. doi:10.1016/j.cognition.2019.02.007 , abstract =

  276. [290]

    Evaluating

    Nouvel, Damien and Ehrmann, Maud and Rosset, Sophie , year =. Evaluating. Named. doi:10.1002/9781119268567.ch6 , note =

  277. [291]

    and Auzanne, Cedric G

    Garofolo, John S. and Auzanne, Cedric G. P. and Voorhees, Ellen M. , year =. The. Content-

  278. [292]

    Generating

    Galibert, Olivier and Jannet, Mohamed Ameur Ben and Kahn, Juliette and Rosset, Sophie , month = may, year =. Generating. Proceedings of the

  279. [293]

    Makhoul, John and Kubala, Francis and Schwartz, Richard and Weischedel, Ralph , keywords =

  280. [294]

    How to evaluate

    Jannet, Mohamed Ameur Ben and Galibert, Olivier and Adda-Decker, Martine and Rosset, Sophie , month = sep, year =. How to evaluate. Interspeech 2015 , publisher =. doi:10.21437/Interspeech.2015-322 , abstract =

  281. [295]

    Chinchor, Nancy and Sundheim, Beth , year =. Fifth

  282. [296]

    Proceedings of the 2nd

    Palen-Michel, Chester and Holley, Nolan and Lignos, Constantine , month = nov, year =. Proceedings of the 2nd. doi:10.18653/v1/2021.eval4nlp-1.5 , abstract =

  283. [297]

    Tsvetkov, Yulia and Dyer, Chris , month = jul, year =. Lexicon. Proceedings of the 53rd. doi:10.3115/v1/P15-2021 , urldate =

  284. [298]

    Sequence

    Wu, Winston and Duh, Kevin and Yarowsky, David , month = aug, year =. Sequence. Findings of the. doi:10.18653/v1/2021.findings-acl.353 , abstract =

  285. [299]

    International Journal of Lexicography , author =

    Multiword. International Journal of Lexicography , author =. 2019 , keywords =. doi:10.1093/ijl/ecy012 , abstract =

  286. [300]

    Proceedings of the 2021

    Lin, Bill Yuchen and Gao, Wenyang and Yan, Jun and Moreno, Ryan and Ren, Xiang , month = nov, year =. Proceedings of the 2021. doi:10.18653/v1/2021.emnlp-main.302 , abstract =

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.