{"id":"a85b3753-0814-4e34-bdac-8ed09530ce4b","arxiv_id":"2608.00713","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A six-year, open, daily-updated database of over two million unassimilated English borrowings in the Spanish press, with an honest evaluation of its detector's real-world precision.","lead":"This paper describes Observatorio Lazaro, a system that has automatically tracked English borrowings in the Spanish digital press since 2020, producing a database of over two million borrowing occurrences. It matters because it turns a slow, manual, dictionary-based field into a continuously updated, open resource for studying how Spanish press language changes in real time.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Corrected hapax and productivity figures apply Table 9's per-tier precision to the whole 2020–2026 database, yet the audit covers only the post-Aug-2022 detector; CRF-era precision for 20.5% of occurrences is unmeasured and may differ.","rationale":"The paper is an honest and unusually transparent resource description: it reports raw precision, breaks errors down by type, gives per-tier estimates, and explicitly lists limitations. I do not question the resource's existence or the integrity of the audit. My concern is narrower. The headline corrected type-level statistics in Section 6.2 (53.6% hapax, P=0.012, beta=0.611) are obtained by applying Table 9's per-tier precision values to all 68,424 lemmas, including the 20.5% of the database produced by the pre-August-2022 CRF model. Section 5.3 explicitly says the CRF model's deployed precision has not been measured. Because the CRF was trained on a smaller headline corpus, its error profile in the rare tail could differ substantially; the nonce tier already has 0.54 precision under the current model, so the corrected hapax share is sensitive to exactly the quantity that is unknown. This is not a hidden flaw, but it is unhedged: unlike recall, the paper does not label these type-level figures as upper or lower bounds. The proposed audit of CRF-era spans is feasible from the stored database and would settle whether the transfer holds. If it fails, the type-level conclusions need qualification; if it passes, the current conditional acceptance should stand. This is why I would keep the reader's CONDITIONAL verdict rather than move it.","tokens_in":22384,"tokens_out":13012,"duration_ms":121450,"concrete_test":"Audit a new stratified sample of 1,000 spans drawn only from CRF-era detections (April 2020–August 2022), following the exact frequency-tier protocol of Table 9, and compare per-tier precision to Table 9 using a chi-square or Fisher's exact test per tier. If CRF-era precision in the nonce or low tiers is significantly below 0.54 or 0.63, recompute the corrected hapax share and the Heaps/Herdan exponents with period-specific precision corrections, and report whether the corrected 53.6% hapax share and the open-class conclusion survive. If the CRF-era tier precisions are statistically indistinguishable from Table 9, the concern does not land.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 5.3 states plainly that the deployed precision audit characterizes only the current BiLSTM-CRF detector (August 2022 onward), covering 79.5% of occurrences, and that the precision of the superseded CRF model behind the other 20.5% has not been measured. Section 6.2 then takes Table 9's per-tier precisions and applies them to the full 2020–2026 database to derive the corrected hapax share (53.6%), P=0.012, C=0.738, and beta=0.611. The CRF model was trained on a smaller, headline-only corpus; its error profile, particularly in the nonce tier that dominates the type inventory and where the current model already has precision 0.54, is very likely different and is not necessarily bounded by the current detector's numbers. Because the type-level statistics in the abstract and Section 6 are a central advertised result, this unvalidated transfer of precision estimates is the weakest load-bearing step. The recall issue is also real but is explicitly hedged as a lower bound in Section 6.2; the CRF-era precision transfer is not hedged at all and is directly testable from the stored data.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents Observatorio Lázaro, a continuously updated automatic monitor of unassimilated lexical borrowings (predominantly anglicisms) in Spanish digital news. It documents the acquisition pipeline, the BiLSTM-CRF detector, the database of 2,007,647 borrowing tokens in 1,880,377 articles over 993.4 million running tokens (2020–2026), and the public web/API access layer. The evaluation includes a held-out span-level F1 of 0.86, an inter-annotator agreement of κ=0.91, and a manual audit of 1,000 stratified spans that yields a token-weighted precision of 0.87, a type-weighted precision of 0.59, and a precision- and recall-corrected density of about 2.14 anglicisms per thousand tokens. Section 6 reports hapax shares (58.7% raw, 53.6% corrected), productivity constants, section-level density contrasts, and temporal stability over the single-model window 2023–2025. The paper explicitly acknowledges that deployed recall is unmeasured and that a mid-2022 detector/outlet change creates a discontinuity.","tokens_in":22592,"tokens_out":12088,"duration_ms":106664,"significance":"If the corrected statistics are supported, this is a valuable resource paper: it offers the first open, daily-updated, diachronic database of anglicism usage in the Spanish press, with a documented pipeline, a public API, a downloadable snapshot, and a rare attempt to quantify deployed precision rather than only lab performance. The paper's frank treatment of unmeasured recall, the 2022 discontinuity, and lemmatization uncertainty is a genuine strength, as is the release of the detector through a pip-installable library and the Zenodo snapshot. However, the headline statistical claims—corrected hapax share, corrected productivity exponents, and the true density estimate—currently rest on extrapolating the precision audit across a detector change and on assuming that test-set recall transfers to deployment. These are fixable but load-bearing issues, which is why I recommend major revision rather than acceptance at this stage.","major_comments":[{"comment":"Section 5.3 states that the deployed precision audit characterizes only the current BiLSTM-CRF detector (August 2022 onward), covering 79.5% of occurrences, and that the precision of the superseded CRF model for the remaining 20.5% has not been measured. Section 6.2 then applies Table 9's per-tier precisions to the full 2020–2026 database to produce the corrected hapax share of 53.6%, P=0.012, C=0.738, and β=0.611. The CRF was trained on a smaller, headline-only corpus and its error profile may differ, especially in the nonce tier where the current model already has strict precision 0.54. Because these corrected statistics appear in the abstract and in Section 6, the transfer is load-bearing. The authors should either audit a sample of CRF-era spans and recompute the corrections, or restrict all corrected type-level and productivity statistics to the period covered by the audit and report raw counts for the pre-August-2022 portion.","section":"§5.3 / §6.2"},{"comment":"The token-weighted precision of 0.87 is computed by weighting per-tier precision values by the tier's share of the token stream, but the audit selected 1,000 spans belonging to 1,000 distinct lemmas, i.e., one occurrence per type. The per-tier precision is therefore a type-level estimate; treating it as an occurrence-level estimate assumes correctness is constant within a lemma across all its occurrences. That assumption is questionable for ambiguous surface forms such as 'horror' or 'look', which are native Spanish words in some contexts and borrowings in others. Since the token-weighted figure is the basis for the corrected density of 2.14 per thousand tokens, the paper should either re-estimate token-weighted precision from an occurrence-level sample or provide evidence that within-lemma variation is negligible.","section":"§5.3, Tables 7–9"},{"comment":"Section 5.3 derives the headline estimate of 2.14 anglicisms per thousand tokens by scaling the precision-corrected detection rate by the test-set recall of 0.82, while acknowledging that deployed recall is unmeasured; Section 8 repeats that recall cannot be measured directly on the deployed data. The abstract and Section 4.2 nevertheless present the density as a stable point value without this condition. Because the detector is static and miss rates on borrowings that entered Spanish after training are a plausible source of downward bias, the density should be reported in the abstract and elsewhere as an order-of-magnitude estimate conditional on recall transfer, or as an explicitly labeled lower or upper bound, consistently with the hedging used in Section 6.4.","section":"§5.3, §4.2, Abstract"}],"minor_comments":[{"comment":"Section 1 cites an earlier estimate of around 2% of the vocabulary in El País in 1991, while the paper's own result is about two per thousand tokens (0.2%); please clarify whether the historical figure is 2% of tokens, 2% of types, or 2 per thousand, since the current juxtaposition appears to imply a tenfold decline where the text elsewhere suggests the modern figure is higher.","section":"§1 vs. §4.2/§5.3"},{"comment":"The abstract says monitoring started in April 2020, but Appendix A lists first-seen dates in January and February 2020 for several core outlets; please make the dates consistent.","section":"Abstract vs. Appendix A"},{"comment":"The sentence 'All four constants fall' immediately follows the reporting of only P, C, and β; if a fourth constant, such as the fitted intercept of the Heaps curve, is intended, it should be named, otherwise the sentence should read 'all three.'","section":"§6.2"},{"comment":"Section 4.2 reports an overall density of approximately 2,020 per million tokens, while Section 6.4 reports 1,817 per million on the composition-stable core outlets; because the denominators differ, the paper should state explicitly that the latter is a raw, core-outlet rate so readers do not read the two numbers as inconsistent.","section":"§4.2 vs. §6.4"},{"comment":"The per-tier precisions in Table 9 are rounded to two decimals, but at the sample sizes in Table 7 they do not always correspond to integer true-positive counts (e.g., 0.63 × 250 = 157.5); please report the raw numerators or add a rounding note for reproducibility.","section":"Table 9"},{"comment":"The corrected values P=0.012, C=0.738, and β=0.611 are reported without the exact procedure by which per-tier precisions were applied to lemma counts and to the token stream; since the correction is described as changing the shape of the accumulation curve rather than merely rescaling it, a worked formula or the analysis code should be provided.","section":"§6.2"}],"recommendation":"major_revision","confidential_remarks":"This is a resource paper whose central contribution—the open, continuously updated database—is sound in design and honestly evaluated. The main risk is the transferability of the precision audit across the 2022 detector change and from type-level to token-level inference; both are addressable within the manuscript's scope. I see no issues with citation practice or overlap with prior work beyond what the paper itself discloses."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The short version: this is a real resource and the paper should go to peer review. It is the first continuously updated, open database of anglicism usage in the Spanish press, and the evaluation is unusually candid. The main thing to fix in revision is that the Section 6 statistics apply the current detector's per-tier precision to the whole 2020–2026 database, even though the audit only covers the post-August-2022 model, which accounts for 79.5% of occurrences. The paper states this plainly in Section 5.3, then proceeds in Section 6.2 without a hedge, quietly assuming the CRF-era error profile matches the BiLSTM-CRF's.\n\nWhat is actually new: the operational pipeline, the deployed precision audit, the per-tier precision numbers, and the error-corrected diachronic characterization. The resource itself—over two million borrowing tokens, 68k types, daily updates, open API and CSV dumps, with a Zenodo snapshot—is the contribution; the detector was already published. I credit the honesty here. They manually audit 1,000 stratified spans, report raw precision of 0.68, derive token-weighted 0.87 and type-weighted 0.59, and explicitly say recall is unmeasured in the wild because only sentences with detections are stored. That is the right way to document a resource, not the way to bury it.\n\nThe load-bearing soft spot is the precision transfer. The type-level numbers in the abstract and Section 6—corrected hapax share 53.6%, P=0.012, beta=0.611—are built on Table 9 applied to the full database, including the 20.5% of occurrences produced by the superseded CRF model. That model was trained on a smaller headline-only corpus, and its error profile, especially in the nonce tier where the current model already has precision 0.54, is likely different. This is directly testable from the stored data, so it is not an unfixable problem; it just should be measured or explicitly bounded before the type-level results are advertised. The recall issue is real but the paper flags it as a lower-bound assumption, which is fair. No confidence intervals on the corrected density is a minor omission, not a fatal one.\n\nWho this is for: contact linguists, Spanish lexicographers, and anyone building monitor corpora. I would cite it. The paper is a serious, self-aware piece of resource documentation, and the author's treatment of the limitations is among the more honest I have seen in this genre.\n\nRecommendation: send it out. It deserves referee time, and the revision should push for either CRF-era precision numbers or a clear statement that type-level statistics are provisional pending that measurement. I disagree with any desk-reject reflex; the resource itself is the point, and it is valuable.","headline":"A genuinely useful, honestly documented resource that deserves a serious referee; the main caveat is that the type-level statistics apply the current detector's precision profile to the earlier CRF-era data without hedging.","tokens_in":23158,"tokens_out":1521,"would_cite":true,"duration_ms":14883,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Observatorio Lázaro is a self-populating, openly queryable monitor of anglicisms in the Spanish press, holding over two million borrowings with quantified precision.","keywords":["anglicisms","lexical borrowing","Spanish press","diachronic corpus","neology monitoring","sequence labeling","language resource","borrowing detection"],"falsifier":"Take a random sample of complete articles from the stable core period of 2023-2025, annotate every unassimilated borrowing by hand, and compare with what the pipeline stored: if the wild recall is materially below the lab figure of 0.82, the corrected density of 2.14 per thousand tokens understates the true anglicism rate and the flat trend could hide a real increase; if precision on the pre-2022 detector period differs from the audit, the type-level statistics need recomputation.","tokens_in":22120,"feed_emoji":"📰","tokens_out":12465,"duration_ms":98231,"temperature":0.7,"pith_summary":"The paper claims that Observatorio Lázaro is a working, continuously updated monitor of unassimilated English borrowings in the Spanish press, and that six years of its output form a reliable diachronic database. It reports more than two million borrowing tokens in 1.88 million articles, with a rate near two anglicisms per thousand tokens that remains stable across the comparable period. The significance is that borrowing dictionaries are static and one-off corpora cannot follow a volatile phenomenon, so a self-populating, openly queryable record would let linguists watch anglicisms enter, spread, and fade in real time. The claim is backed by a held-out F1 of 0.86, inter-annotator agreement of 0.91, and a 1,000-span precision audit whose token-weighted precision is 0.87. The paper also gives a first statistical profile: a steeply skewed, open vocabulary dominated by rare types, concentrated in fashion, technology, and lifestyle sections.","feed_headline":"Spanish press anglicisms: 2 million uses, stable rate of 2 per 1,000","feed_subtitle":"A daily-run detector with published error rates makes six years of anglicism births and trends openly queryable.","key_machinery":"The load-bearing mechanism is an end-to-end daily pipeline: article retrieval, text cleaning and tokenization, span-level sequence labeling for unassimilated borrowings, lemmatization, and storage in a database served through a public website and API. The central object is the unassimilated lexical borrowing span, a foreign-origin word or multiword expression used in otherwise monolingual Spanish text and not yet adapted to Spanish spelling or morphology. The evaluation rests on a frequency-tiered precision audit: because borrowings follow a Zipfian distribution, the paper samples 1,000 distinct lemmas across five frequency tiers and weights each tier's measured precision by its actual share of types and tokens, which is what yields the contrast between token-weighted precision of 0.87 and type-weighted precision of 0.59. That correction, together with the detector's test-set recall, converts raw detections into population estimates such as the density of 2.14 per thousand tokens.","core_discovery":"On its own terms, the paper's discovery is that a neural borrowing detector, run daily on a fixed set of news outlets, can produce a valid longitudinal record rather than only a benchmark. The record contains 2,007,647 borrowing tokens across 1,880,377 articles and 993.4 million running tokens from 2020 to 2026. Evaluated on held-out text the detector reaches a span-level F1 of 0.86, and a manual audit of 1,000 frequency-stratified spans gives a token-weighted precision of 0.87 and a type-weighted precision of 0.59. After discounting false positives and scaling by the detector's test-set recall, the paper estimates a true anglicism density of 2.14 per thousand tokens, about one in every 500, and treats it as stable over the single-model window starting in 2023. It further claims that the borrowing vocabulary is an open and growing class: after precision correction, 53.6% of types are attested once, the Heaps exponent is 0.611, and borrowing density varies by more than an order of magnitude across newspaper sections.","pith_inferences":["A direct consequence of the paper's own recall limitation, worth making explicit: the stable 2023-2025 density is an upper bound on any real decline and a lower bound on any real rise, because a static detector will tend to miss the newest borrowings, so true contemporary usage could be higher than two per thousand.","A testable extension of the paper's reasoning is that the sharp section gradient, from about 10.5 borrowings per thousand tokens in fashion to under 0.7 in politics, points to register and domain rather than global language contact as the main driver, which the database's outlet and section fields make directly testable.","The same pipeline could be transplanted to other recipient languages or donor languages, but the paper's own finding that non-English borrowings are heavily under-detected suggests such a transplant would need rebalanced training data before its non-English counts could be trusted.","Because type-weighted precision is only 0.59, any study of lexical innovation built on this resource will be very sensitive to the precision correction; re-auditing the nonce tier on a larger sample could move the corrected hapax share of 53.6% substantially."],"forward_implications":["Diachronic studies of anglicism birth and spread in Spanish can now run on open, continually updated data instead of static dictionaries or hand-annotated snapshots.","Frequency and trend analyses can treat the data as nearly benchmark-grade, while rare-type and neology studies must apply the paper's per-tier precision corrections.","Longitudinal comparisons should be restricted to the stable core outlets and the single-model window from 2023 onward, because the 2022 detector and outlet changes create a comparability boundary.","Lexicographers and language planners gain a candidate-detection feed of new anglicisms with first-attestation dates, contexts, and outlet and section distributions.","The measured density of about two unassimilated anglicisms per thousand tokens provides a current baseline for the Spanish press, replacing dated estimates from earlier decades."],"supporting_citations":[{"why":"Supplies the annotated training corpus and the BiLSTM-CRF detector whose benchmark F1 of 0.86 underlies the resource's evaluation.","marker":"Álvarez-Mellado and Lignos (2022)"},{"why":"Provides the earlier annotated corpus of emerging anglicisms and the CRF model that produced the pre-2022 portion of the database.","marker":"Álvarez-Mellado (2020a)"},{"why":"Frames automatic detection of unassimilated borrowings in the Spanish press as a shared benchmark task that the pipeline builds on.","marker":"Álvarez-Mellado et al. (2021)"},{"why":"Supplies the inter-annotator agreement threshold used to judge the training annotations reliable.","marker":"Artstein and Poesio (2008)"},{"why":"Provides the lemmatizer used to group surface forms into lemma types, which affects all type-level statistics.","marker":"De Smedt and Daelemans (2012)"},{"why":"Provides the tokenizer and sentence splitter that define the running-token unit and the stored context sentences.","marker":"Montani et al. (2023)"},{"why":"Is the citable, versioned snapshot of the exact 2020-2026 database described in the paper.","marker":"Álvarez Mellado (2026)"}],"fun_headline_variants":["Spanish press anglicisms: 2M uses, stable 2 per 1,000 tokens","Observatorio Lazaro: 2M anglicisms in Spanish press, open data","Daily scan logs 2M anglicisms in Spanish press over 6 years","Open database: 2M anglicisms in Spanish press, stable rate","Six-year anglicism record: 2M uses, open API, stable rate"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Everything rests on assuming the detector makes the same kinds of mistakes out in the wild as it did on the 1,000 hand-checked spans and the lab test set, even though the system only stores sentences where it found something and never measures what it missed.","fun_headline_variants_meta":{"raw":{"variants":["Spanish press anglicisms: 2M uses, stable 2 per 1,000 tokens","Observatorio Lazaro: 2M anglicisms in Spanish press, open data","Daily scan logs 2M anglicisms in Spanish press over 6 years","Open database: 2M anglicisms in Spanish press, stable rate","Six-year anglicism record: 2M uses, open API, stable rate"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000684,"raw_usage":{"total_tokens":3206,"prompt_tokens":1150,"completion_tokens":2056,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":766,"completion_tokens_details":{"reasoning_tokens":1945}},"tokens_in":766,"tokens_out":2056,"duration_ms":12131,"temperature":1.0,"reasoning_tokens":1945,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T15:16:48.394205+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of complete articles from the stable core period of 2023-2025, annotate every unassimilated borrowing by hand, and compare with what the pipeline stored: if the wild recall is materially below the lab figure of 0.82, the corrected density of 2.14 per thousand tokens understates the true anglicism rate and the flat trend could hide a real increase; if precision on the pre-2022 detector period differs from the audit, the type-level statistics need recomputation.","supporting_citations":[],"review_version":2}