{"id":"80243484-a3f7-4c7a-af55-5964a1113536","arxiv_id":"2501.18129","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Seventy bibliometric gender-bias studies show heterogeneous author name disambiguation and gender identification methods, prompting a new SoDA Cards reporting template to improve comparability.","lead":"After reviewing 70 published studies on gender bias in scholarly metrics, this paper documents inconsistent choices in author name disambiguation and gender identification and proposes a standardized reporting template called SoDA Cards. A smart generalist might read it to understand whether conflicting bibliometric gender-gap findings come from real differences or from different methods.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Method variability is documented, but the claim that it limits reliability and comparability is not tested; no sensitivity analysis links AND/gender-identification choices to different gender-bias conclusions.","rationale":"I read the paper as a methodological review plus a proposal. The documentation of variability is credible: the kappa scores are reported, the codebook is explicit, and the distribution of methods is concrete evidence. However, the central claim has two components: variability exists, and variability is consequential. The first is supported by the data; the second is not. The most load-bearing unsecured link is the consequence claim. It is testable by reanalysis, and the paper does not perform or cite such a test on its own corpus. The sample-representativeness issue identified by the reader is real, but even a perfectly representative sample would not answer whether different methods change the conclusions. My concern is therefore complementary rather than identical, which is why I mark agreement as partial. Because the reader's verdict (CONDITIONAL) already reflects incomplete validation of the proposal, my concern does not move the verdict; it specifies an additional condition that should be met before the central claim is fully accepted.","tokens_in":23165,"tokens_out":4952,"duration_ms":50565,"concrete_test":"Re-analyze one of the surveyed studies' datasets (e.g., Caplar et al. 2017, or a common MAG/OpenAlex corpus) under multiple pipelines: (a) no disambiguation (authorship-level), (b) name-based AND, (c) algorithmic AND; and gender via (i) Genderize.io/heuristics, (ii) an algorithmic classifier, (iii) a manual/gold-standard subset. Compare the estimated gender effect sizes (e.g., citation or publication gaps) and their statistical significance across pipelines. If the estimates are stable, the observed methodological variability is not consequential for the central claim; if they differ materially, the paper's premise is supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central inference—that the observed variability in author name disambiguation (AND) and gender identification is large enough to limit reliability and comparability—is not directly demonstrated. Tables 2 and 3 show that methods differ (51.4% of papers did no AND; 64.3% used a single gender-identification method), and Sections 2.1 and 2.2 cite prior work (e.g., Kim and Diesner 2016; Santamaria and Mihaljevic 2018) showing that individual methods can be error-prone. However, the paper does not show that using different AND methods (e.g., name-based vs. algorithmic vs. gold-standard) or different gender-identification methods (e.g., Genderize.io vs. manual/gold-standard assignment) materially changes the estimated gender gaps in the surveyed studies' own analyses. Without such a sensitivity analysis, the 'limits reliability and comparability' premise underlying the SoDA Cards proposal is an assertion rather than a demonstrated result. The observed variation establishes heterogeneity of practice, but not its consequences; the urgency argument depends on the latter.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reviews 70 peer-reviewed publications from 2009-2023 on gender bias in bibliometrics, annotating author name disambiguation (AND) and gender identification methods with a codebook and Cohen's kappa. The main empirical finding is that 51.4% of sampled papers performed no AND and 64.3% used a single gender identification method, with name-based heuristics being the most frequent approach. Based on this variability, the authors argue that methodological inconsistencies limit reliability and comparability, and they propose Scholarly Data Analysis (SoDA) Cards, a reporting template adapted from Model Cards and Datasheets, to standardize documentation of data sources, AND, gender identification, causal controls, and results.","tokens_in":23384,"tokens_out":5794,"duration_ms":52975,"significance":"If the empirical distribution is reliable, the paper provides a useful, systematically annotated map of current methodological practice in a policy-relevant area. The SoDA Cards proposal is a concrete, low-cost transparency mechanism that the bibliometrics community could adopt, and the paper is commendable for publishing its codebook, inter-annotator agreement scores, and a filled example card. However, the contribution's central premise—that observed variability undermines reliability and comparability—is not directly tested, and internal inconsistencies in the reported percentages weaken the quantitative evidence. The paper would be significantly strengthened by a sensitivity analysis or a more measured framing of the claims.","major_comments":[{"comment":"The percentages reported in the text and in Table 3 are internally inconsistent: heuristics are reported as 42.27% in the text but as 41.84% (41/98) in the table; manual search appears as 23.71% in the text but as 22.45% (22/98) in the table; algorithm appears as 18.56% in the text but as 18.37% (18/98) in the table; and gold standards is reported as 15.46% although the table's own count of 17/98 equals 17.35%. These discrepancies directly affect the paper's central descriptive claims and must be reconciled.","section":"4.2.1, Table 3"},{"comment":"The paper asserts that methodological variability 'limits reliability and comparability,' but no analysis in the manuscript demonstrates this. The distribution of methods establishes heterogeneity, and the cited prior work (e.g., [38], [76]) shows that individual methods can be error-prone, but there is no evidence that the studies' conclusions about gender bias would change under alternative AND or gender-identification methods. Since the urgency of the SoDA Cards proposal rests on this premise, the paper should either provide a sensitivity analysis on a subset of the reviewed papers or explicitly reframe the claim as a hypothesis requiring further study.","section":"4.1.1, 4.2.1, abstract"},{"comment":"The sampling procedure is not reported in sufficient detail for the review to be reproducible: the exact Google Scholar query strings, retrieval dates, screening workflow, and counts of excluded papers are absent, and the assertion that 'most citations of these influential papers utilize similar methodological practices' is not backed by citation data. This is important because the representativeness of the 70-paper sample underpins the headline distributions in Tables 2 and 3.","section":"3.1, Appendix A.1"}],"minor_comments":[{"comment":"The heading 'Gender idenitfication' contains a typo; it should read 'Gender identification.'","section":"Section 5.4 heading"},{"comment":"The field labels contain spacing artifacts such as 'T emporal scope' and 'Geographical scope' with a space before 'scope'; additionally, the filled example in Figure 6 is labeled 'Model Card example' although the paper introduces it as a SoDA Card, which should be consistent.","section":"Figures 6 and 7"},{"comment":"The text states the sample covers papers 'published between 2009 to 2023' and the title says 'past 12 years,' but Figure 3 shows a distribution starting in 2010; this is a minor inconsistency that should be corrected.","section":"Section 3.1"},{"comment":"The gold standard row reports 4.28% for 3/70, which should be 4.29% for consistency with the other rounded percentages.","section":"Table 2"},{"comment":"The sentence 'This, combined with the difficulty in inferring gender from Asian names, led them to exclude researchers from China ... and Singapore' would benefit from a citation to the specific passage in [35], as the current text paraphrases without a page or section reference.","section":"Section 4.2.3"}],"recommendation":"major_revision","confidential_remarks":"The internal percentage inconsistencies are the kind of data-quality issue that suggests the tables need a careful re-audit. The paper's core proposal is reasonable, but the framing as a demonstrated problem rather than a hypothesis overreaches. The absence of detailed search strings is fixable with an online appendix."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nRead the Lee et al. paper. The honest summary: it's a solid, useful review of 70 studies of gender bias in scholarly metrics, with new empirical distributions of author name disambiguation (AND) and gender identification methods (Tables 2 and 3). The finding that 51% of papers do no disambiguation and 64% rely on a single gender-identification method, with heuristics most common, is real and worth having on record. The SoDA Cards proposal is a sensible adaptation of Model Cards/Datasheets to the bibliometrics context, and the worked example (Caplar et al.) makes the template concrete. Kappa scores (0.71-0.85) are acceptable for this kind of annotation. Credit where due: the authors ship a codebook, a GitHub repo, and are transparent about their own limitations.\n\nThe soft spots are real but not fatal. First, the sample is deliberately citation-biased (Google Scholar, high-cited papers, no preprints); the authors acknowledge this but then lean on it to claim their distributions reflect 'dominant practices.' That inference is shaky. Second, the appendix lists keywords but not the actual search strings, query dates, or inclusion/exclusion screening, so the 'systematic review' label is weaker than it looks. Third, there are small numerical inconsistencies: heuristics reported as 41.84% in Table 3 but 42.27% in the text; manual search 22.45% vs 23.71%; algorithm 18.37% vs 18.56%. And Table 2's 'No Disambiguation' list includes reference 87 twice. None of this undermines the qualitative picture, but it needs cleaning.\n\nThe bigger conceptual issue, which the stress-test note flags, is real: the paper documents heterogeneity of practice but never shows that different AND or gender-identification choices materially change the gender-gap estimates. The abstract and Section 2.5 assert that methodological inconsistency 'limits reliability and comparability.' That is plausible and supported by prior work on individual methods being error-prone, but this paper doesn't run or cite a sensitivity analysis that connects method choice to conclusions. So the urgency argument for SoDA Cards rests partly on an assertion. I'd like to see either a small empirical demonstration (e.g., re-analyze one dataset with two AND methods) or a softer, more careful claim about comparability.\n\nOverall: this is a paper for bibliometricians and meta-science researchers, and for anyone designing gender-bias studies. It deserves a serious referee; with revisions (numbers fixed, search protocol documented, claim moderated) it would be a solid contribution. I'd send it out.","headline":"A genuinely useful review of methodological variability in gender-bias bibliometrics with a smart reporting template, though the claim that variability breaks comparability is asserted more than proven.","tokens_in":23891,"tokens_out":3128,"would_cite":false,"duration_ms":29109,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that unreliable author name disambiguation and gender identification methods undermine gender-bias bibliometrics, and proposes a SoDA Cards framework for standardized reporting.","keywords":["gender bias","bibliometrics","author name disambiguation","gender identification","SoDA Cards","scholarly data analysis","methodological standardization","reproducibility"],"falsifier":"Apply the authors' own codebook to a broader unselected corpus of gender-and-citation studies, including preprints and low-citation work: if the method distribution clusters on one or two dominant approaches, the no-consensus claim is falsified. A second test: re-run one landmark gender-bias dataset through every combination of disambiguation and gender-identification method; if the conclusion about gender bias is unchanged across all combinations, the practical urgency of standardization weakens substantially.","tokens_in":23000,"feed_emoji":"📊","tokens_out":6921,"duration_ms":67382,"temperature":0.7,"pith_summary":"Gender-bias bibliometrics rests on two error-prone data-processing steps, deciding which author is which and deciding each author's gender, and this review of 70 studies argues that current practice is too varied to support reliable, comparable conclusions. More than half of the sampled studies (51.4%) performed no author name disambiguation, and 64.3% relied on a single gender-identification method, with name-based heuristics the most common choice. The paper's remedy is the Scholarly Data Analysis (SoDA) Card, a structured reporting template that documents how names were disambiguated, how gender was assigned, how unknown labels were handled, and which causal factors were controlled. If the field adopts the cards, individual findings become easier to compare and aggregate, which the paper argues is a prerequisite for evidence-informed policy on gender inequality in academia.","feed_headline":"51% of gender-bias bibliometric studies skip name disambiguation","feed_subtitle":"A 70-paper review shows no consensus on methods, then proposes SoDA Cards to standardize reporting.","key_machinery":"The carrying object is the Scholarly Data Analysis (SoDA) Card, a structured reporting template with sections for study specification, corpus profile, author name disambiguation, gender identification, analysis, and results. It works by requiring explicit answers to questions that most reviewed studies leave implicit: which disambiguation method was used and whether it was evaluated, which name part fed gender inference, which gender categories were allowed, what share of authors stayed unidentified, how unknown labels were handled, and which causal factors such as career length were controlled. The card rests on the paper's annotation taxonomy, which sorts disambiguation into five categories and gender identification into four, and on the annotation process that produced inter-annotator agreement of 0.71-0.85 (Cohen's kappa) before discrepancies were resolved to full consensus.","core_discovery":"The paper's central claim is that the methodological pipeline, not just the finding, determines what gender-bias bibliometrics can say. In a curated sample of 70 peer-reviewed works from 2009 to 2023, the paper finds no dominant standard: 51.4% of papers did no author name disambiguation and analyzed authorship-level records, 21.4% used algorithmic disambiguation, 12.9% name-based heuristics, 10.0% manual searches, and 4.3% gold-standard identity data. For gender identification, 64.3% of papers used a single method, name-based heuristics were the most frequent approach (41 of 98 counted method-instances, 41.8%), and gold-standard self-reported gender was used least often. The paper argues that this variability, combined with the documented difficulty of Asian names and the common practice of dropping authors with unassignable gender, makes existing results hard to compare, and it proposes the SoDA Card as a documentation standard to fix that.","pith_inferences":["Beyond the paper: the disclosure requirement itself may improve data quality, because researchers who must report 'no disambiguation performed' face pressure to justify it.","Beyond the paper: a natural test is to apply SoDA Cards retrospectively to the 70 reviewed papers; the result would be a reusable benchmark of pipeline choices that future gender-bias studies could control for.","Beyond the paper: the four gender-identification types do not capture every tool's internal behavior, so a practical extension would require recording tool name, version, and parameter settings, since two tools both labeled 'algorithmic' can disagree on the same name."],"forward_implications":["Studies that fill a SoDA Card become directly comparable, enabling meta-analyses that aggregate gender-gap effect sizes instead of stacking incompatible pipelines.","Journal and funder adoption of the card would create longitudinal records of which methods dominate, letting the field track whether practice is improving.","The card requires reporting the percentage of unidentified gender and how unknown labels were handled, making the disproportionate exclusion of Asian or unisex names visible rather than silent.","Reliable estimates of where gender bias does and does not exist are the basis for policy interventions, so the card indirectly strengthens evidence-based decision-making."],"supporting_citations":[{"why":"Shows how initial-based name disambiguation distorts coauthorship network measurements, grounding the claim that method choice changes gender-bias results.","marker":"[38]"},{"why":"Reports that the self-citation gender gap disappears after controlling for career length, motivating the paper's call for causal-factor controls.","marker":"[58]"},{"why":"A canonical large-scale study in the corpus whose name-based gender identification and authorship-level analysis the paper categorizes.","marker":"[46]"},{"why":"A survey of name-to-gender inference services used by the paper to establish that gender-identification tools vary in reliability.","marker":"[77]"},{"why":"The model-reporting standard that SoDA Cards adapt to bibliometrics.","marker":"[59]"},{"why":"The dataset-documentation convention that motivates the paper's transparency-by-template approach.","marker":"[28]"},{"why":"A corpus study that explicitly excluded researchers from several Asian countries, used as evidence of how name-handling choices shape samples.","marker":"[35]"}],"fun_headline_variants":["SoDA Cards to standardize gender-bias research","Half of gender-bias studies ignore name disambiguation","Gender-bias metrics need SoDA Card transparency","Uncomparable gender-bias studies get SoDA fix"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the 70 highly cited, keyword-matched, peer-reviewed papers represent the field as a whole; if the many papers with fewer citations or non-matching keywords actually share a common methodology, the 'no consensus' finding loses its force.","fun_headline_variants_meta":{"raw":{"variants":["SoDA Cards to standardize gender-bias research","Half of gender-bias studies ignore name disambiguation","Gender-bias metrics need SoDA Card transparency","Uncomparable gender-bias studies get SoDA fix"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000738,"raw_usage":{"total_tokens":3314,"prompt_tokens":978,"completion_tokens":2336,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":594,"completion_tokens_details":{"reasoning_tokens":2285}},"tokens_in":594,"tokens_out":2336,"duration_ms":17012,"temperature":1.0,"reasoning_tokens":2285,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T00:33:16.017255+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Apply the authors' own codebook to a broader unselected corpus of gender-and-citation studies, including preprints and low-citation work: if the method distribution clusters on one or two dominant approaches, the no-consensus claim is falsified. A second test: re-run one landmark gender-bias dataset through every combination of disambiguation and gender-identification method; if the conclusion about gender bias is unchanged across all combinations, the practical urgency of standardization weakens substantially.","supporting_citations":[{"cited_title":"Distortive effects of initial-based name disambiguation on measurements of large-scale coauthorship networks","cited_arxiv_id":null,"evidence_quote":"Shows how initial-based name disambiguation distorts coauthorship network measurements, grounding the claim that method choice changes gender-bias results."},{"cited_title":"Fegley, Jana Diesner, and Vetle I","cited_arxiv_id":null,"evidence_quote":"Reports that the self-citation gender gap disappears after controlling for career length, motivating the paper's call for causal-factor controls."},{"cited_title":"Sugimoto","cited_arxiv_id":null,"evidence_quote":"A canonical large-scale study in the corpus whose name-based gender identification and authorship-level analysis the paper categorizes."}],"review_version":1}