{"id":"15ca637b-9c6d-4499-8a2c-7f377e6296ea","arxiv_id":"2412.02342","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"A rule-based text-mining method applied to French press archives identified 88 new downstream space companies, adding roughly a third to a known company database.","lead":"This paper proposes a step-by-step, rule-based computer method for finding French companies that sell products and services built on satellite data, by scanning newspaper articles and a national company registry. Applied to 22 years of French press, it flagged 88 previously unknown companies in this 'downstream space' business.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline 88 includes 30 'by-product' companies that fail Rule 4 and are found only by an undocumented expert co-citation step; strictly, the rule-based pipeline yields 58.","rationale":"I read the paper in good faith: it is a transparent methodological application, the rules are described in unusual detail, and the limitations of press coverage are openly acknowledged. The reader's CONDITIONAL verdict is reasonable. However, the single most load-bearing defect I find is not primarily the representativeness of the seed database, although that is a real structural limit. It is the composition of the headline '88.' The paper itself distinguishes 'strict implementation' (58) from 'by-products' (30), and the by-product mechanism is not a rule: it is an expert co-citation review applied to a 22,862-company intermediate list, with no specified inclusion criterion. Counting those 30 as products of the rule-based approach conflates the algorithmic pipeline with the manual sorting that surrounds it. This directly affects the paper's central quantitative claim and its reproducibility. My proposed check is a single replication exercise: count strict-rule outputs separately from codified by-products. If the strict count is 58, the abstract and conclusion should be revised, and the enrichment percentage should be recomputed. This does not change the appropriate verdict from the reader's CONDITIONAL—it reinforces it—so I set verdict_should_be to UNCHANGED. I marked agreement as partial because the reader identified adjacent concerns (numerical inconsistencies, missing audit trail, in-sample evaluation) but did not spotlight the by-product conflation as the core weakness of the 88 figure.","tokens_in":18765,"tokens_out":9537,"duration_ms":109801,"concrete_test":"Re-run the method exactly as specified through Rule 4 on the same press corpus and Sirene dictionary, and count only companies that pass all four rules. Then codify a by-product protocol—for example, take every company in the 22,862-company intermediate list that appears within 30 words of a Rule-4-passing company in the same article, classify it by the paper's website/context criteria, and count separately. If the strict count is 58 and the by-product count is unstable or not reproducible, the abstract's 88 should be revised to '58 via rules plus 30 via co-citation review,' and the enrichment percentage recomputed from the revised total.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative result—88 new downstream companies 'using our rule-based named entity recognition approach'—mixes two distinct detection mechanisms. Section 4 is explicit: strict application of Rules 1–4 yields 58 companies; the other 30 are 'by-products,' companies that fail Rule 4 (they do not contain any regular expression from Table 4) and were noticed only because they are co-cited with Rule-4 companies during the manual sort of the 1,475-company final list. The by-product stage is not defined as a rule: neither the set of co-cited intermediate-list companies examined nor the co-citation criterion is specified, so the method cannot be replicated to produce 30. Section 3.5 says the authors 'do not entirely exclude' non-regular-expression companies, but the inclusion rule is not operationalized. The abstract and conclusion nevertheless attribute all 88 to the rule-based NER approach. This matters because the headline number is the paper's main evidence of value; if the strict method is 58, the enrichment claim changes (88/246 is about 36%, 88/334 is about 26%, neither matching the stated 33%, and 58/246 is about 24%). The reader's conditional verdict already requires data/code release; the by-product ambiguity is an internal reproducibility defect that should be fixed before the 88 figure is used.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a rule-based Named Entity Recognition pipeline to identify French downstream space companies from press articles. The method filters the French Sirene company registry by APE codes and legal statuses (Rule 1), keeps only capitalized words in a downstream-space query corpus (Rule 2), retains matches within 30 words of context or query keywords (Rule 3), and finally keeps names containing hand-picked regular expressions (Rule 4). The pipeline reduced 6.9 million legal units to 1,475 candidates, which were then manually sorted by an expert. The paper reports 58 new companies from the strict application of the rules plus 30 'by-products' discovered through co-citation during sorting, for a total of 88 new downstream companies, and claims this enriches the known-company database by 33%.","tokens_in":19008,"tokens_out":4810,"duration_ms":54161,"significance":"If the central claim were properly supported, the paper would make a useful methodological contribution: it addresses a real gap in identifying downstream space firms that standard industrial classifications miss, and it provides a transparent, rule-based alternative to statistical NER with detailed intermediate counts, honest reporting of losses (31 of 220 known companies never appear in the press), a full query and rule list, a comparison with spaCy, and practical time estimates. These are genuine strengths. However, the headline result and its evaluation currently rest on an undocumented manual by-product step, in-sample performance metrics, and no released code or company list, which limits the paper's evidential value until those issues are repaired.","major_comments":[{"comment":"The headline figure of 88 is not the output of the rule-based method as defined. The text states that strict implementation of Rules 1-4 yielded 58 companies, while the 30 additional 'by-products' are companies that failed Rule 4 and were included only because an expert noticed their co-citation with Rule-4 companies during manual sorting. No operational criterion is given for co-citation: neither the set of intermediate-list companies examined nor the relationship, distance, or evidentiary threshold for inclusion is specified, so the 30 by-products cannot be reproduced. The abstract and conclusion nevertheless attribute all 88 companies to 'our rule-based named entity recognition approach.' The enrichment figures are also internally inconsistent: Table 5 sums to 334 but the text says 344; 88/220 is 40%, 88/246 is about 36%, and 88/334 is about 26%, none matching the stated 33%. The authors should report the strict-rule count and the by-product count separately, define a reproducible co-citation protocol or drop the by-products from the main claim, and reconcile the percentages.","section":"Section 4; Table 5; Abstract"},{"comment":"The Conservation ratio of 120/128 is measured against the same set of known companies that was used to construct the rules. The APE codes (Rule 1), legal statuses, the Word Context list (Rule 3), and the regular expressions (Rule 4) were all derived from the 220-company reference database and a 500-article sample, and the 128 companies are a subset of those 220. The performance metrics therefore describe in-sample calibration, not predictive or discovery performance. The authors are transparent that they did not split the data, but the conclusion that the rules are effective should be framed as calibration; a stronger validation would use the 26 'Other' companies or newly collected firms as a holdout set. This is load-bearing because the paper's methodological value claim is that the rules identify companies beyond those used to build them.","section":"Section 5.2.1; Section 3.2; Section 3.4; Section 3.5"},{"comment":"The manual sorting of the 1,475 candidates is a single-expert classification with no inter-rater reliability measure, no codebook beyond a two-step procedure, and no audit trail. Given that only 58 of the 1,475 (about 4%) were classified as downstream companies and that the 30 by-products were added during the same subjective review, the classification step carries substantial weight in the main result. A reproducibility-focused methodology paper should provide a classification codebook, record for each accepted company the evidence used (press description, website keywords, or call-for-projects mention), and ideally report agreement with a second coder on a subsample. Without this, the numerical results cannot be independently verified or applied by other researchers.","section":"Section 4; Section 6"},{"comment":"The paper promises guidelines for replication but does not release the article corpus, the code for the matching algorithm, the construction of the Sirene dictionary, or the final list of 334 companies with their sources. Since the method's contribution is methodological and the evaluation depends on exact intermediate lists and matching decisions, the absence of these materials prevents any independent audit of the 58-company strict result, the 30 by-products, or the comparison with spaCy. The authors should provide a supplementary data and code package, or at minimum the final company list and the code for Rules 1-4, before the reported numbers can be accepted as more than an illustrative application.","section":"Section 6; data availability"}],"minor_comments":[{"comment":"The total number of companies in the final database is given as 344 in one place, 334 in Table 5, and 334 in the Discussion; these figures should be reconciled.","section":"Section 4; Table 5; Section 7"},{"comment":"The first sentence of Section 3.3 calls Rule 2 'The third rule,' although Rule 2 is the second rule in the numbering; this appears to be a typo.","section":"Section 3.3"},{"comment":"The number of potential downstream companies after Rule 3 is reported as 22,862 in Section 4 and as 22,863 in Section 5.2.1; the discrepancy should be corrected.","section":"Section 5.2.1; Section 4"},{"comment":"The comparison with spaCy is not apples-to-apples: spaCy labels organizations and miscellaneous entities without access to the Sirene dictionary, while the rule-based method starts from a dictionary matching step. Clarify that the 83 known companies are those whose names appeared among spaCy's labeled entities, rather than a dictionary-based identification, and discuss how this affects the interpretation of the comparison.","section":"Section 5.2.2"},{"comment":"The text refers to 'Figure 5.1' where it should refer to the figure on known-company losses, which is labeled Figure 5; please fix the cross-reference.","section":"Section 2.1; Figure 5"}],"recommendation":"major_revision","confidential_remarks":"The reader's and skeptic's concern about the 88 versus 58 discrepancy is well grounded and should be the central focus of the revision. The paper has a genuinely useful methodological core, but the headline result is currently overstated and the evaluation is in-sample. If the authors separate the strict-rule output from the by-product step, provide a reproducible by-product criterion, reconcile the enrichment percentages, and release the code and company list, the paper could be acceptable for publication in an applied economics or domain-application venue. I would not recommend rejection at this stage."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Start with the punchline. The headline discovery—88 new downstream space companies identified by a rule-based NER pipeline—is not exactly what the abstract claims. Read Section 4 carefully: strict application of the four rules yields 58. The other 30 are 'by-products' that failed Rule 4 and were spotted manually because they were co-cited with Rule-4 companies in the intermediate list. That distinction is in the paper, so the authors are not hiding it, but the abstract and conclusion attribute all 88 to the rule-based approach. That's a framing problem, not a fraud.\n\nWhat's actually new and useful is the detail. The query construction, the Sirene dictionary filtering by APE codes and legal forms, the capital-letter restriction, the 30-word context window, the regular expression list—every step has intermediate counts. The evaluation against the 220 known companies shows where known companies are lost (press absence, misquotes, English-only articles) and is honest about the losses. The spaCy comparison, while simple, gives a sanity check. For anyone trying to build a company list in a niche industry without clean activity codes, this is a concrete, transferable recipe.\n\nThe soft spots are real but proportionate. No list, corpus, or code is released, so the 58 and the 88 cannot be independently checked. The by-product stage is not operationalized: no rule describes what co-citation qualifies, so that 30 is not reproducible. The evaluation is in-sample—rules were tuned on the 220 known companies and then measured against them, so the conservation ratio of 120/128 is optimistic. There are also minor internal inconsistencies: the abstract's 33% enrichment doesn't match the numbers in Table 5 (88/246 ≈ 36%, 58/246 ≈ 24%), and the text says 22,863 in one place and 22,862 in another. None of these are fatal, but together they keep the central claim conditional.\n\nThe reader's conditional verdict is fair. I'd send this to a serious referee, and if I were the editor, I'd make acceptance contingent on releasing the company list and code, specifying the by-product criterion, and fixing the numerical inconsistencies. The paper is for applied economists and others who compile firm registers from unstructured text; for that audience, it earns a careful read.","headline":"A transparent methodology paper whose headline 88 is really 58 plus 30 manually spotted co-cited firms—worth reviewing, but the abstract oversells the rule-based pipeline.","tokens_in":19588,"tokens_out":3063,"would_cite":false,"duration_ms":33523,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A rule-based named-entity-recognition pipeline applied to French press articles and the national business registry identified 88 new downstream space companies, expanding the known sector database by more than a third.","keywords":["downstream space sector","named entity recognition","rule-based NLP","text mining","company identification","French press","space economy","industry classification"],"falsifier":"Hold out a random 30 of the 128 known companies that appear in the query articles, rebuild Rules 1, 3, and 4 using only the other 98, and count how many held-out companies the rebuilt pipeline recovers; if recovery is much below the reported 120 of 128, the rules are overfit to the seed database rather than capturing the sector.","tokens_in":18505,"feed_emoji":"🛰️","tokens_out":10301,"duration_ms":104276,"temperature":0.7,"pith_summary":"The paper proposes a reusable text-mining method to find French companies in the downstream space sector, firms earning revenue from products and services built on satellite data such as communications, Earth observation, and navigation. The method is rule-based named-entity recognition: starting from 220 known downstream companies, it builds a dictionary from the national business registry using their industry codes and legal forms, then scans 48,900 French press articles and applies three further rules (capitalized words, a 30-word semantic context, and recurring name patterns) to shrink millions of possible firms to 1,475 candidates for expert review. In the first application, expert sorting confirmed 58 new downstream companies from that list, and co-citation during sorting uncovered 30 more, for 88 new companies and a more-than-a-third enrichment of the reference database. The paper's evaluation finds the rule cascade keeps 120 of the 128 known companies that appear in the query articles, while a generic statistical NER model recovers only 83 of those 128. If this holds, the method gives economists and space agencies a practical way to track a sector that official activity classifications cannot delimit, and a template that can be ported to other countries or industries.","feed_headline":"Rule-based text mining finds 88 new French space firms","feed_subtitle":"Four hand-crafted rules on 48,900 press articles expand the known downstream-space company database by one-third.","key_machinery":"The load-bearing object is the Sirene dictionary: a subset of the French national business registry cut down to the 33 APE (official activity) codes and the 8 legal statuses that appear among the 220 known downstream companies. Around that dictionary the method wraps four rules: (1) keep only registry entries matching those codes and legal forms; (2) keep only words in the press text that start with a capital letter, with accents removed and punctuation preserved; (3) keep only matched company names that appear within 30 words of a word on the business-context list or the query's own keywords, so the citation is semantically anchored to downstream space; and (4) keep only names containing one of 18 recurring character strings observed in downstream company names. The cascade converts 6.9 million registered firms into 1,475 hand-checkable candidates, and that reduction, from national registry to expert-sortable list, is the mechanism that makes the identification tractable.","core_discovery":"On the paper's own terms, the central result is that a transparent, expert-crafted NER cascade can do what standard activity classifications and generic statistical NER do not: find downstream space companies in France that were previously unknown. Starting from 220 known downstream firms, the authors build a dictionary by filtering the French national business registry (6.9 million active legal units) to the 33 official activity codes and 8 legal forms present among the known firms, yielding 650,000 candidates. Matching these against the capitalized words of 48,900 French press articles on downstream space activity leaves 30,084 companies; requiring a mention within thirty words of a business or space-context word leaves 22,862; requiring one of eighteen recurring name patterns (such as data, geo, ima, sat, tele) leaves 1,475. Expert sorting of that list confirms 58 new downstream companies, and co-citation during sorting adds 30 more, for 88 new companies total and a final 2022 database of 334 firms, more than a quarter of them detected by the method. In the evaluation, the rule cascade keeps 120 of the 128 known companies that appear in the query articles, while a generic statistical NER model labels 268,015 entities yet recovers only 83 of those 128.","pith_inferences":["Beyond the paper: because Rule 4's regular-expression list and Rule 1's code list are learned from the 2022 known-company snapshot, the pipeline's precision will decay as New Space applications diversify; an annual recalibration, as the paper recommends, is not just a convenience but a correctness condition.","Beyond the paper: the 30 co-cited 'by-product' companies suggest the press corpus contains an implicit network, firms mentioned together in the same article are likely to share programs, markets, or supplier links, and the paper's final paragraph points to this unexploited structure; converting co-citation into a graph would give an independent test of downstream membership.","Beyond the paper: the reported 4% precision among the 1,475 Rule-4 candidates implies the method is best read as a candidate generator, not a classifier; feeding the 88 new names plus their article contexts into a small trained NER model would likely close much of the gap toward the fully automated system the paper sketches."],"forward_implications":["The downstream-space company database can be updated annually by rerunning the four rules on a single year of press articles, with a substantial drop from the initial 12-day application time because the rules and query stay unchanged.","The same methodology can be transplanted to other countries: replace the registry, query language, and known-company seed list, and the rule cascade will produce a candidate list for the local downstream space sector.","For any industry whose firms straddle many official activity codes, the rule cascade offers a way to detect new entrants that sector-level input-output analysis misses.","The method's conservation ratio (120 of 128 known companies survive the text rules) indicates that most downstream companies that do appear in the query articles can be detected, while the 31 known companies never found in the press mark a hard ceiling on press-based identification."],"supporting_citations":[{"why":"Supplies the New Space framing and the definition of downstream space activities that the identification method targets.","marker":"[Bousedra, 2023]"},{"why":"Defines the downstream segment and the value-chain view that the paper argues official activity classifications cannot capture.","marker":"[OECD, 2022]"},{"why":"Supplies the application-area categories (communications, Earth observation, navigation) used to design the press query.","marker":"[Booz & Company, 2014]"},{"why":"The closest prior use of NER for the satellite domain and the statistical approach the rule-based method is compared against.","marker":"[Maurya et al., 2022]"},{"why":"Shows a trained NER model can outperform generic NER on satellite-domain entities, motivating the paper's proposed merge of rule-based and statistical approaches.","marker":"[Jafari et al., 2020]"},{"why":"Provides a French financial-news NER corpus that justifies choosing the pretrained French news model used as the statistical baseline.","marker":"[Jabbari et al., 2020]"},{"why":"Compares rule-based and machine-learning NER and supports the paper's choice of a rule-based design for a task where context is essential.","marker":"[Gorinski et al., 2019]"},{"why":"Shows that text can serve as an alternative to official industry classification, which motivates the whole identification strategy.","marker":"[Hoberg and Phillips, 2016]"},{"why":"Grounds the paper's use of press text as economic data and situates the method within text-as-data methodology.","marker":"[Gentzkow et al., 2019]"}],"fun_headline_variants":["Rule-based NER discovers 88 new French space firms","Text mining unearths 88 new downstream space firms","How rule-based NER spotted 88 new French space companies","88 new downstream space firms found by rule-based NER","One-third more space firms: rule-based NER finds 88"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole pipeline assumes that the 220 known downstream companies are a fair mold for the entire population, same activity codes, legal forms, capitalization habits, and name patterns, and that press coverage is a reliable net; since 31 of those 220 companies never appear in either press database, any downstream company that looks different or stays out of the news will be missed no matter how well the rules are tuned.","fun_headline_variants_meta":{"raw":{"variants":["Rule-based NER discovers 88 new French space firms","Text mining unearths 88 new downstream space firms","How rule-based NER spotted 88 new French space companies","88 new downstream space firms found by rule-based NER","One-third more space firms: rule-based NER finds 88"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000939,"raw_usage":{"total_tokens":4000,"prompt_tokens":918,"completion_tokens":3082,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":534,"completion_tokens_details":{"reasoning_tokens":2998}},"tokens_in":534,"tokens_out":3082,"duration_ms":24896,"temperature":1.0,"reasoning_tokens":2998,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T23:33:41.981595+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Hold out a random 30 of the 128 known companies that appear in the query articles, rebuild Rules 1, 3, and 4 using only the other 98, and count how many held-out companies the rebuilt pipeline recovers; if recovery is much below the reported 120 of 128, the rules are overfit to the seed database rather than capturing the sector.","supporting_citations":[],"review_version":1}