{"id":"79ea1603-1c85-4b8d-859c-c725f64b6032","arxiv_id":"2412.03557","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"FICE, a title-entity measure combining term freshness and rarity, correlates strongly with 5-year cumulative citation counts for groups of ACL Anthology papers.","lead":"The authors propose a new measure, FICE, which scores a paper by how fresh and how rare the scientific terms in its title are. They report that groups of ACL papers with higher FICE scores tend to receive more citations over the following five years.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"FICE's freshness weights are fit on the full corpus, so post-publication entity success leaks into the purported citation correlation.","rationale":"The reader's stated weakest assumption was the reliability of Gaussian extrapolation for term lifetimes, and that is a real concern. However, I see a more fundamental issue: the fitted lifetime curve for each entity uses the entire corpus, including documents published after the paper being scored. This means future information about entity adoption is baked into the freshness weight at t0, independently of whether the Gaussian model is a good fit. Even a perfectly fit Gaussian would leak post-publication success into FICE, so the central correlation may be inflated by construction. The recommended test is a causal refit using only data available at t0. This test also subsumes the reader's concern, because it removes both the future-data leakage and the need to extrapolate beyond the observed period from post-t0 data. I do not see a need to move the verdict: the paper is already CONDITIONAL, and this concern strengthens the conditions rather than changing the verdict class. The code is publicly available, which makes the proposed test straightforward.","tokens_in":8023,"tokens_out":9988,"duration_ms":110641,"concrete_test":"Recompute FICE for the same ACL quotas using a strictly causal estimation: for each paper at publication year t0, refit df(e,t) for every entity in its title using only documents published up to t0 (for rare entities, use a regularized single-Gaussian fit on the truncated data), obtain te from that truncated fit, and recompute Eq. (1)-(3). Rerun the Spearman correlations in Table 2. If the coefficients drop below roughly 0.3 or lose significance, the headline correlation depends on post-publication entity success; if they remain near 0.7, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing assumption is not just that Gaussian extrapolation approximates term lifetimes; it is that the lifetime model is fit on the full corpus and then evaluated at each paper's publication time t0. Section 4.4 fits df(e,t) for each entity using the complete 1952-2020 yearly curve; Section 4.5 then computes r(e,t0) in Eq. (1) using te obtained from that full-corpus fit. For any paper published before the corpus end, the denominator sum over [ts, te] includes documents published after t0 (plus the predicted tail beyond 2020). Thus an entity's later adoption rate can raise the denominator, lower r(e,t0), and inflate freshness 1 - r(e,t0). FICE at t0 is therefore not a contemporaneous property of the title; it is a retrospective quantity that already encodes how successful the entity became after t0. Since future entity adoption is plausibly correlated with future citations, the headline Spearman coefficients (0.766 and 0.748 in Table 2) may be partly an artifact of the fitting protocol rather than a genuine association between title-entity freshness and citation impact. The paper reports no causal/truncated refit (fit only on data up to t0), so the reader cannot separate the claimed association from this leakage.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes a new bibliometric indicator, Freshness and Informativity Weighted Cognitive Extent (FICE), which weights the scientific entities appearing in paper titles by a freshness factor (1 minus a lifetime ratio) and by a frequency-based informativity factor. The lifetime of each entity is modeled by fitting its yearly document frequency with a composite of Gaussian profiles, and the lifetime ratio compares the cumulative document frequency up to a paper's publication year with the cumulative frequency over the entire modeled lifetime. Using the ACL Anthology, the authors report that entity-based cognitive extent grows more slowly than the number of papers, and that average FICE per quota correlates strongly with the log of a 5-year cumulative citation count (Spearman 0.766 and 0.748 for quota sizes 125 and 250 in Table 2). The paper also compares FICE with three simplified baselines and claims that FICE shows the strongest correlation.","tokens_in":8326,"tokens_out":5539,"duration_ms":56613,"significance":"If the correlation is robust, the result would be interesting: aggregate title-vocabulary freshness and informativity would predict collective 5-year citation impact, and FICE would extend the original cognitive extent in a principled way. The paper has several concrete strengths: the code is publicly available, three entity recognizers are compared on a manually annotated benchmark, several baseline formulations are tested, and the authors explicitly note that the correlation is collective rather than applicable to individual papers. The main significance risk is that the headline correlation may be an artifact of fitting the lifetime model on the full corpus: the freshness weights for papers published before the end of the corpus use document frequencies observed after the paper's publication year, so post-publication entity success leaks into the metric. The statistical evidence is also less robust than the text suggests, because the correlation at quota size 500 is not significant and one baseline outperforms FICE at that quota.","major_comments":[{"comment":"The lifetime ratio r(e,t0) is computed with ts and te obtained from a Gaussian fit to the full 1952–2020 document-frequency curve. For any paper published at t0 before the corpus end, the denominator in Eq. (1) includes observed documents published after t0 as well as the extrapolated tail beyond 2020. Consequently, freshness 1 − r(e,t0) encodes how successful the entity became after the paper was published, and because future entity adoption is plausibly correlated with future citations, the Spearman coefficients in Table 2 may be inflated. This is a real leakage problem, not merely a theoretical one. I request a truncated refit in which df(e,t) is fit using only years up to t0, or an explicit reinterpretation of FICE as a retrospective hybrid measure, with correspondingly tempered language about prediction.","section":"§4.4–4.5, Eq. (1)"},{"comment":"The claim that \"FICE exhibits the strongest correlation against all baseline models\" is not supported at |Q|=500: in that row the lifetime-ratio-only baseline has ρ=0.744, which is larger than FICE's ρ=0.717, and FICE's p-value is 0.109, i.e., not significant at the 0.05 level. The headline strong correlation therefore rests almost entirely on the two smaller quota sizes. Please report the number of bins used for each correlation, provide confidence intervals, and also compute a per-paper rank correlation without binning into quota averages, so that the reader can distinguish a genuine aggregate association from an artifact of binning.","section":"Table 2, |Q|=500 row"},{"comment":"The Gaussian composite fitting procedure is not validated. The paper does not report goodness-of-fit statistics, residual diagnostics, uncertainty in the fitted parameters, or a comparison with alternative lifetime models, and te is often located beyond the observable period by construction. Because r(e,t0) and hence FICE depend directly on te, a poor extrapolation would make the freshness weights artifacts of the fitting model rather than properties of the corpus. Please include a hold-out validation, a null or alternative model comparison, and a report of the distribution of te values across entities.","section":"§4.4"},{"comment":"The definition of C5(2015) as the sum of citations received in years 2015–2019 appears to use a fixed calendar window, regardless of each paper's publication year. A paper published in 2019 thus contributes only one year of citations, while a paper published in 2010 contributes citations received well after its first five years. This mixes publication years and citation windows, confounding the relation between FICE at publication time and citation impact. Please clarify the window; if it is indeed fixed, compute per-paper first-five-year citation counts or include a publication-year control.","section":"§4.6, §5.2"},{"comment":"The statement that entity-based cognitive extent \"increases at a slower rate\" is contradicted by the |Q|=125 row, where the slope rises from 1.19 in 1980–2000 to 6.80 in 2000–2020. The decreasing slope holds only for |Q|=250 and |Q|=500, so the conclusion is quota-dependent as reported. Please either present a trend test that accounts for quota size and time period, or explicitly report that the slowdown is not uniform across quota sizes.","section":"§5.1, Table 1"}],"minor_comments":[{"comment":"There is a typo, \"Sciientific\" for \"Scientific,\" in the last paragraph of Section 4.2.","section":"§4.2"},{"comment":"The header contains \"Speareman\" instead of \"Spearman.\"","section":"Table 2"},{"comment":"Figure 4 is used for both the entity-based cognitive extent growth plot and the FICE–C5 correlation plot, which is confusing; please renumber the figures.","section":"Figures"},{"comment":"The caption says the FICE values are calculated using \"undisambiguated entities,\" while Section 5.2 says disambiguation does not affect the correlation; please clarify which entity set is plotted and whether the reported coefficients are for disambiguated or undisambiguated entities.","section":"Figure 4 caption"},{"comment":"The reported F1 scores for SciBERT and SpaCy (0.05 and 0.07) are far below typical published NER results on scientific text; please include the annotation guideline details, the benchmark construction, and the exact prompts so readers can assess whether the comparison is fair.","section":"§4.2"},{"comment":"The error bars are described as Gaussian standard deviations, but citation distributions are highly skewed; please use quartiles or bootstrapped confidence intervals for the binned averages.","section":"§5.2"}],"recommendation":"major_revision","confidential_remarks":"The leakage concern from the reviewer stress test does land: the full-corpus fit used in Section 4.4 makes freshness a retrospective quantity. That said, the issue is fixable within the manuscript's scope by adding a truncated refit, and the paper otherwise makes a reasonable empirical contribution with released code. The internal inconsistencies in Table 1 and Table 2 should also be corrected before the paper can be accepted."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Chen,\n\nThe headline: the paper's central correlation is not trustworthy as stated. FICE's freshness weight at publication time t0 is computed from a document-frequency model fit on the full 1952–2020 corpus. For a paper published before 2020, the denominator in Eq. (1) includes documents published after t0, so an entity that later becomes popular gets a lower lifetime ratio and thus a higher freshness. Since future popularity is plausibly correlated with future citations, the Spearman coefficients in Table 2 (0.766 and 0.748) are at least partly an artifact of the fitting protocol. The stress-test note holds up on close reading.\n\nWhat the paper does well: it proposes a clean composite metric (freshness times informativity), compares it against three ablations, and ships code. The replication of Milojević's slower-growth trend on ACL titles is a useful data point. The entity recognition comparison (GPT-4 vs SciBERT vs spaCy) is honest and gives a concrete reason for the choice.\n\nSoft spots beyond the leakage: (1) the Gaussian extrapolation of entity lifetimes is not validated anywhere; te can lie outside the observed period and small fitting errors compound. (2) The correlation is computed on binned averages with |Q|=125, 250, 500; p-values rise sharply with bin size and the bins are not independent. (3) C5 is a fixed 2015–2019 calendar window that pools papers from different publication years, mixing age effects into the citation count. None of these are fatal alone, but they add uncertainty.\n\nThe fix is straightforward and the authors should have run it: fit df(e,t) only on data up to t0, recompute FICE at t0, and re-test the correlation. If the correlation survives truncated refits, the result would be interesting. As is, the metric is retrospective and the predictive claim in the conclusion is not supported.\n\nI would send this to peer review because the idea is new and the flaw is correctable, but I would not cite the correlation in its current form.\n\nBest.","headline":"FICE's correlation with citations is likely inflated because the lifetime model is fit on the full corpus, making the metric retrospective rather than predictive.","tokens_in":8779,"tokens_out":3066,"would_cite":false,"duration_ms":30090,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper proposes a quota-level title-vocabulary measure, FICE, that weights scientific entities by freshness and informativity, and reports a strong positive Spearman correlation with the logarithm of 5-year average cumulative citation…","keywords":["cognitive extent","citation impact","entity recognition","document frequency","freshness","informativity","citation prediction","title analysis"],"falsifier":"Compare the fitted predictions against reality: take entities whose fitted extinction year falls inside the corpus window, so their later document frequencies are observable, and check whether the Gaussian-composite fit predicts the actual yearly counts. Then recompute FICE using observed lifetimes only and re-run the Spearman test; if the strong correlation disappears or drops sharply, the reported relation depends on the extrapolation rather than on title vocabulary.","tokens_in":7836,"feed_emoji":"📈","tokens_out":10710,"duration_ms":98526,"temperature":0.7,"pith_summary":"The paper tries to establish that the collective title vocabulary of a body of papers carries information about how many citations those papers will receive over the following five years. It defines FICE, a quota-level statistic that weights each recognized scientific entity in a title by two quantities: freshness, the share of the entity's predicted lifetime still ahead of it, and informativity, how uncommon the entity is within its contemporary vocabulary. Using entities extracted from the titles of computational-linguistics papers, the authors report that average FICE per quota correlates strongly with the logarithm of the average 5-year cumulative citation count, with Spearman coefficients of 0.766 at quota size 125 and 0.748 at quota size 250. If the result holds, title-vocabulary dynamics can act as a collective early indicator of research impact, even though the correlation is not claimed to apply to individual papers.","feed_headline":"Freshness of title terms tracks 5-year citation counts","feed_subtitle":"A weighted measure of how new and how rare title entities are correlates at 0.77 with average citations per quota.","key_machinery":"FICE (Freshness and Informativity Weighted Cognitive Extent) is the central object. It is a weighted sum over a quota of papers: for each recognized scientific entity in a title, multiply informativity $w(e,t_0)$ by freshness $1-r(e,t_0)$, and sum across titles. The load-bearing part is the lifetime ratio $r$, computed from a model of each entity's yearly document frequency as a composite of Gaussian profiles; the model supplies the extinction year beyond the observed period, so freshness is not just presence or absence but where the entity sits in its predicted career. Informativity $w$ is a normalized inverse document frequency taken across the entities in the same title, ensuring rare entities count more. FICE carries the correlation result because it converts a raw vocabulary count into a measure of how new and how rare a quota's vocabulary is.","core_discovery":"FICE extends the earlier notion of cognitive extent, which counted unique phrases per quota, by replacing the dichotomous count with a weighted sum. For each paper title in a quota, every scientific entity contributes $w(e,t_0)(1-r(e,t_0))$, where the lifetime ratio $r$ is cumulative document frequency up to publication divided by the total over the entity's entire modeled lifetime, and the informativity weight $w$ normalizes cumulative document frequency across all entities in the same title. The authors model each entity's yearly document frequency as a composite of Gaussian profiles, extrapolate the fit to find the extinction year, and use per-year citation records to compute the average 5-year cumulative citation count per quota. In the corpus studied, they find a strong positive Spearman correlation between average FICE and the logarithm of that citation count, and they show FICE beats three simpler variants: the dichotomous entity count, the informativity weight alone, and the freshness weight alone. They also reproduce the earlier observation that the unique entity count per quota grows more slowly than the paper count.","pith_inferences":["A natural next test is to control for publication year, because both FICE and citation counts move upward over time; a year-controlled version would show how much of the correlation is vocabulary signal rather than shared time trend.","The same lifetime-fitting machinery could be reused to map how individual scientific terms age and fall out of use, independent of citation prediction.","A cross-corpus replication on titles from other disciplines would show whether the strong correlation is a property of scientific titles in general or is specific to this corpus and its entity extractor."],"forward_implications":["FICE can act as a collective, citation-free early indicator of which research topics are accumulating 5-year citation impact.","The dichotomous entity count has weak or negative correlation with the citation measure, so the signal comes from the freshness and informativity weights rather than from vocabulary size alone.","The slower-than-exponential growth of unique scientific entities per quota, previously observed in other fields, also holds in the computational-linguistics corpus studied here.","Because FICE is defined for any scholarly text, the same weighted cognitive extent can be applied to abstracts, full texts, or other document units without changing the formula.","The paper is explicit that the collective correlation does not imply that individual authors can raise citation counts by coining new entity names, since transient entities contribute little to FICE."],"supporting_citations":[{"why":"Defines cognitive extent as unique phrases per quota and reports the slower-growth trend that FICE extends and reproduces.","marker":"[1]"},{"why":"Supplies the corpus metadata, titles and publication years, used for entity counts and quota construction.","marker":"[18]"},{"why":"Supplies the language model adopted for zero-shot scientific entity recognition after benchmarking.","marker":"[19]"},{"why":"Provides the annotation guidelines for the benchmark used to select the entity recognizer.","marker":"[22]"},{"why":"Provides the similarity model used for entity disambiguation through a thresholded classification step.","marker":"[15]"},{"why":"Supplies the optimizer used to fit the Gaussian-composite document frequency curves that determine entity lifetimes.","marker":"[23]"},{"why":"Supplies the per-year citation counts that define the 5-year cumulative citation outcome.","marker":"[24]"}],"fun_headline_variants":["Freshness-weighted title terms predict citation counts","Title novelty and rarity metric ties to citations","New measure links title freshness to citation impact","Rare and new title terms signal future citations","FICE: A metric for title freshness and informativity"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the Gaussian-curve extrapolation used to predict when an entity stops appearing approximates the entity's real lifetime; if that predicted extinction year is wrong, the freshness weights are artifacts and the reported citation correlation would not be a property of the corpus.","fun_headline_variants_meta":{"raw":{"variants":["Freshness-weighted title terms predict citation counts","Title novelty and rarity metric ties to citations","New measure links title freshness to citation impact","Rare and new title terms signal future citations","FICE: A metric for title freshness and informativity"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000268,"raw_usage":{"total_tokens":1637,"prompt_tokens":982,"completion_tokens":655,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":598,"completion_tokens_details":{"reasoning_tokens":585}},"tokens_in":598,"tokens_out":655,"duration_ms":7041,"temperature":1.0,"reasoning_tokens":585,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T22:15:48.776178+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare the fitted predictions against reality: take entities whose fitted extinction year falls inside the corpus window, so their later document frequencies are observable, and check whether the Gaussian-composite fit predicts the actual yearly counts. Then recompute FICE using observed lifetimes only and re-run the Spearman test; if the strong correlation disappears or drops sharply, the reported relation depends on the extrapolation rather than on title vocabulary.","supporting_citations":[{"cited_title":"Milojević, Quantifying the cognitive extent of sci- ence, Journal of Informetrics 9 (2015) 962–973","cited_arxiv_id":null,"evidence_quote":"Defines cognitive extent as unique phrases per quota and reports the slower-growth trend that FICE extends and reproduces."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the corpus metadata, titles and publication years, used for entity counts and quota construction."},{"cited_title":"semanticscholar.org/product/api, 2024","cited_arxiv_id":null,"evidence_quote":"Supplies the per-year citation counts that define the 5-year cumulative citation outcome."}],"review_version":1}