{"id":"ef0e46b9-72c9-4a98-ae32-2876e64af985","arxiv_id":"2608.13444","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Gender prediction is illegitimate but gender imputation can still yield valid disparity measurements for traditional sexism, though not for oppositional sexism.","lead":"This paper argues that predicting someone's gender is always illegitimate, yet imputing gender can still produce valid measurements of certain real-world gender disparities. It distinguishes two forms of sexism to explain when imputation helps and when it harms.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central validity claim depends on an untested alignment between imputation outputs and the perceptions of actual discriminators; without direct validation, the paper's 'can yield valid measurements' remains an assertion, not a demonstrated capability.","rationale":"The reader's weakest_assumption correctly targets the empirical premise that discrimination tracks perceived gender and that imputation captures those cues. I agree with that diagnosis, and I sharpen it: even granting the construct, the imputation model is a distinct measurement instrument whose outputs must match the perceptions of the specific decision-makers whose behavior generates the disparity. This alignment is testable, and the paper's own Section 5.2 points to the needed test ('surveying recruiters') without providing results. The paper is a careful conceptual contribution, and its normative argument about illegitimacy is well supported by the cited critical literature; the weak point is the positive validity claim, which remains a possibility argument. Because the paper already frames the claim as conditional and offers recommendations for achieving validity, the reader's CONDITIONAL verdict is appropriate. I would not move it to ACCEPT without direct empirical validation, nor to REJECT, because the argument is coherent and the required validation is feasible. Verdict unchanged from the reader's conditional.","tokens_in":22422,"tokens_out":4819,"duration_ms":52236,"concrete_test":"Run a validation study in a resume-screening setting: collect a sample of names with actual hiring outcomes; obtain perceived-gender judgments from a panel of recruiters; run a standard name-based imputation model (e.g., genderize.io or a recent neural classifier) on the same names; and compute the gender disparity in outcomes using (a) recruiter-perceived gender and (b) imputed gender, using the continuous-output aggregation recommended in Sec. 5.4. Pre-register a tolerance (e.g., relative difference >20% or sign flip). If the two estimates disagree beyond tolerance, the validity claim fails in this concrete setting; if they agree across several settings, the claim gains the empirical support it currently lacks.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's positive thesis—that gender imputation can yield valid disparity measurements for traditional sexism (Sec. 4.3)—requires that the construct measured by an imputation model coincide with the construct that actually triggers discrimination: perceived gender. Section 4.2 asserts this via the catcalling example ('a nonbinary person misperceived as a woman might still experience catcalling'), and Section 4.2 footnote 9 acknowledges the open question 'perceived by whom?', deferring to 'a possibly discriminatory decision-maker.' Section 5.2 then recommends training and validating imputation models on directly collected perceived-gender data, e.g., surveying recruiters. That recommendation marks the condition for validity, but the paper provides no empirical evidence that any existing or proposed imputation model satisfies it. Models are commonly trained on self-reported labels, sex-assigned-at-birth name databases, or crowd annotations, none of which necessarily match the perceptions of the decision-makers whose behavior constitutes the disparity. If the model's predicted perceived gender diverges from those perceptions, the resulting aggregate disparity estimate is invalid for exactly the construct the paper says imputation can measure. The claim 'can yield valid measurements' is therefore conditional on an alignment that is plausible but unverified.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This position paper argues that algorithmic gender prediction is always illegitimate because it restricts individuals' agency to self-determine their gender, but that gender imputation—prediction used for aggregate disparity measurement—can nevertheless yield valid measurements of traditional sexism, since traditional sexism operates on perceived gender and social position rather than on self-identified gender identity. The authors draw on transfeminist literature, especially Serano's distinction between traditional sexism and oppositional sexism, and on measurement theory (Adcock and Collier, Messick) to separate questions of legitimacy from questions of validity. They offer recommendations for when imputation might be justified, including narrowly scoped gender constructs, using appropriate inputs, and aggregating estimates, and they illustrate the framework with three case studies: auditing gender bias in generative image models, measuring gender disparities in film, and imputing gender from names. The paper is conceptual; it provides no empirical validation of the central validity claim.","tokens_in":22646,"tokens_out":5942,"duration_ms":62826,"significance":"If the argument succeeds, the paper makes a valuable contribution by giving researchers a vocabulary to reconcile two positions that are often treated as contradictory: the ethical rejection of gender prediction and the practical need for demographic labels in disparity auditing. The distinction between legitimacy and validity, and the mapping of these concepts onto traditional versus oppositional sexism, productively clarifies a real conflation in the literature. The recommendations—use narrowly scoped perceived-gender constructs, train on directly collected perceived-gender data, aggregate rather than individualize, and use continuous estimators—are concrete and actionable. The paper is careful, well-referenced, and includes thoughtful positionality, ethical, and adverse-impact statements. Its main weakness is that the positive claim \"can yield valid measurements\" is asserted through illustrative examples and case-study commentary rather than demonstrated through any empirical validity check or fully specified measurement conditions. The contribution is therefore best understood as a taxonomy and research agenda; the title's promise of valid measurements remains conditional.","major_comments":[{"comment":"The positive validity claim is conditional on an unverified alignment between imputation outputs and the perceptions of the discriminators whose behavior constitutes the disparity. The paper asserts that traditional sexism is triggered by perceived gender and social position, and that imputation can capture these constructs (e.g., the catcalling example in §4.2), but it offers no empirical evidence that any imputation model's predicted labels coincide with the perceptions of actual decision-makers in a given setting. The authors' own footnote 9 (\"perceived by whom?\") and their recommendation in §5.2 to collect perceived-gender data by surveying recruiters concede that this alignment is an open empirical question. Because the title claims imputation \"can yield valid measurements,\" this evidence—or an explicit reformulation of the claim as a conditional feasibility argument—is needed to support the central thesis.","section":"§4.2, footnote 9; §5.2"},{"comment":"The relation between trans-misogyny and the two-sexism distinction is under-specified. The authors state that imputation can help address \"traditional sexism, including instances of trans-misogynistic discrimination,\" but trans-misogyny, under the Serano account they adopt, is a compound of traditional sexism and oppositional sexism. If imputation is illegitimate and unsuitable for measuring oppositional sexism, the reader needs to know whether the disparity estimate captures only the traditional-sexism component of trans-misogyny and whether such a partial measurement is still valid or is instead liable to mislead. This ambiguity affects the scope of the central claim and should be resolved explicitly.","section":"§4.2"},{"comment":"The notion of \"validity\" used for disparity measurements is not operationalized sufficiently to assess the title's claim. The paper invokes content, convergent, and consequential validity, but it does not state what evidence would establish that an imputed-gender disparity measure is valid, nor under what measurement-error conditions it fails. For example, the paper could specify a formal condition such as the imputation being nondifferential with respect to the disparity construct, or provide a concrete validation procedure such as comparing imputed-perceived-gender disparities to audit-study benchmarks or to self-reported perceived-gender data. Without this, \"valid\" remains ambiguous between a philosophical claim and a statistical one, and the second half of the paper depends on that distinction.","section":"§4.1, §4.3"}],"minor_comments":[{"comment":"The abstract contains a typo: \"predicting gender iswrong\" should read \"predicting gender is wrong.\"","section":"Abstract"},{"comment":"The phrase \"she coinstraditional sexismto refer\" is missing spaces and should read \"she coins 'traditional sexism' to refer.\"","section":"§4.2"},{"comment":"The author name \"V ogel\" appears with an extra space in the text and in the reference list; it should be \"Vogel.\"","section":"References and §6.3"},{"comment":"The sentence about Spanish versus English-language Chinese name datasets is difficult to parse; consider rewriting to clarify the intended contrast between name-gender associations across linguistic populations.","section":"§5.2"},{"comment":"The table classifies \"using self-reported gender identity to measure online recruiter discrimination\" as invalid, which is correct but likely surprising to readers; a brief explanation in the main text or table caption would help clarify that the invalidity stems from a construct mismatch, not from the data source.","section":"Appendix C, Table 3"}],"recommendation":"major_revision","confidential_remarks":"The paper is a conceptual contribution, and I want to be fair about the appropriate standard of evidence: a philosophical argument that imputation can be valid in principle may not require a full empirical study. However, the title and abstract make a categorical claim (\"can yield valid measurements\") that goes beyond the evidence provided. A revision that either adds one worked validation example or explicitly reframes the claim as a conditional feasibility argument would resolve my main concern. I see no issues with citation integrity or novelty; the paper builds appropriately on prior work in trans studies, measurement theory, and algorithmic fairness."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a read. The contribution is a clean separation of two questions that usually get tangled: whether gender prediction is illegitimate (normatively wrong because it restricts gender self-determination) and whether it is invalid (descriptively wrong as measurement). On top of that, the paper brings in Serano's traditional versus oppositional sexism and argues that imputation, though always illegitimate, can in principle yield valid aggregate measurements of traditional sexism—discrimination based on perceived gender and social position—while being fundamentally ill-suited to oppositional sexism. That synthesis is new and genuinely useful. It gives critics and practitioners a shared vocabulary instead of talking past each other. The three case studies (generative image models, film, name-based inference) make the abstractions concrete, and the recommendations are practical: scope the gender concept to the setting, collect perception-aligned data, aggregate, weigh benefits against harms.\n\nThe soft spot is exactly where the stress-test lands. The positive thesis depends on an empirical premise: that the cues an imputation model uses are the cues that actually trigger discrimination—that a face, a name, or an appearance-based prediction matches the perception of the decision-maker whose behavior constitutes the disparity. The paper states this premise (footnote 9 asks 'perceived by whom?') and then recommends, in Sec 5.2, training and validating models on directly collected perceived-gender data. That is the right move, but it also means the title's 'can yield valid measurements' is a conditional promise, not a demonstration. No existing pipeline is shown to satisfy the alignment condition. I don't think that kills the paper—it is a conceptual argument, and the conditions are spelled out—but a referee should push the authors to be more explicit that the empirical adequacy is open and to sketch a concrete validation study, e.g., comparing imputed labels to recruiter perceptions in a resume audit.\n\nThere is self-citation (Dong et al. 2025, Wang 2025, Wallach et al. 2025), but it is not padding; those are the directly relevant prior results. The paper is careful, honest about its limitations, and engages seriously with trans scholarship.\n\nBottom line: the framework holds together, and the paper deserves a serious referee. If I were editing, I would send it out; the main revision request would be to temper the headline claim or clearly mark the empirical alignment as an open condition.","headline":"A genuinely useful conceptual separation of legitimacy from validity for gender imputation, with an empirical premise that remains untested and should be flagged in review.","tokens_in":23169,"tokens_out":4054,"would_cite":true,"duration_ms":42467,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Algorithmic gender prediction is always illegitimate, but gender imputation used in the aggregate can still yield valid measurements of traditional sexism—discrimination against women and femininity—even as it remains harmful toward…","keywords":["gender imputation","gender prediction","measurement validity","legitimacy","traditional sexism","oppositional sexism","algorithmic fairness","transgender and nonbinary people"],"falsifier":"A correspondence audit could settle it: submit matched resumes or applications in which targets share the same gender identity but differ in perceived-gender cues, such as name and photo, and record real decision-makers' responses. If callback or hiring rates track perceived gender, the validity premise holds; if they track identity instead, the premise fails. A second check compares imputed-perceived-gender disparity estimates against self-reported-identity estimates in the same population; systematic divergence in settings where identity-based discrimination is known to operate would falsify the claim that imputation validly measures sexism there.","tokens_in":22190,"feed_emoji":"⚖️","tokens_out":8400,"duration_ms":76134,"temperature":0.7,"pith_summary":"Gender imputation—predicting gender from names, faces, or other cues to fill in missing demographic data—is a common tool for measuring gender disparities in hiring, film, and image generation. Yet transgender and nonbinary scholars have argued that algorithmically assigning gender is morally wrong. This paper tries to reconcile those positions by separating two questions: whether gender prediction is illegitimate, meaning it denies people the authority to self-determine their gender, and whether it is invalid, meaning it produces unusable measurements. Its central claim is that imputation is always illegitimate but can sometimes be valid—specifically for measuring traditional sexism, discrimination that targets women and femininity, because that form of discrimination operates on perceived gender rather than on gender identity. The paper argues that scoping imputation narrowly, aggregating results, and reserving it for cases with no reasonable alternative can turn an unethical practice into a usable measurement tool without legitimizing gender prediction itself.","feed_headline":"Gender imputation can be illegitimate yet still valid","feed_subtitle":"Separating legitimacy from validity shows when aggregate imputed gender can audit real disparities.","key_machinery":"The argument is carried by two conceptual distinctions and one definition. The first distinction separates legitimacy—whether a prediction normatively deserves social authority, judged here by whether it restricts a person's capacity to self-determine gender—from validity, whether a measurement accurately captures the concept it claims to measure. The second, from transfeminist literature, distinguishes traditional sexism (discrimination against women and femininity) from oppositional sexism (discrimination against gender deviance and non-normativity, including transphobia and cissexism). The definition is that of gender imputation as a subset of gender prediction: prediction serves a limited auditing or descriptive end and is interpreted only in the aggregate, at the structural level rather than as claims about individuals. Together these let the paper argue that traditional sexism is carried by perceived gender and social position, so aggregate estimates of perceived gender can validly index it, while the same estimates remain illegitimate for the same reason they are harmful: they assume the very authority over gender that self-determination denies them.","core_discovery":"The paper's central claim is that illegitimacy and invalidity are distinct failings, and judging one does not settle the other. Drawing on transfeminist theory, it distinguishes traditional sexism—the privileging of masculinity and targeting of women and femininity—from oppositional sexism, which targets gender deviance and includes transphobia, homophobia, and cissexism. Because traditional sexism is triggered by how a person is perceived and socially positioned rather than by how they identify, an imputation model that estimates perceived gender from names, faces, or appearance can in principle measure the disparities this sexism produces, and can continue to do so even though every such prediction is illegitimate under the paper's self-determination criterion. The paper does not claim all imputation is valid: it is often invalid for the same reasons it is illegitimate, and it can never validly measure oppositional sexism, whose harms it necessarily misses. Valid use requires shifting the target from gender identity to a narrowly scoped concept such as \"perceived gender based on a resume,\" aggregating continuous predictions rather than discretizing them into fixed labels, and reserving imputation for settings where anti-discrimination benefits cannot be obtained through other reasonable means.","pith_inferences":["This analysis suggests a testable empirical program: correspondence audits that pair the same gender identity with different perceived-gender cues, such as names or photos, could verify that decision-makers' discriminatory behavior tracks perceived gender, which would ground the paper's validity premise in data rather than assertion.","The legitimacy–validity split plausibly extends to race and age imputation, where similar binds arise; the paper leaves this extension open, and the same scoped-measurement analysis could clarify when surname-and-location imputation is defensible.","Read as a measurement agenda, the paper's recommendations set an evaluation standard: any imputation-based disparity claim should report uncertainty under misclassification error and be validated against a self-reported subsample, making \"valid\" verifiable rather than asserted.","If discrimination in a specific domain tracks gender identity rather than perception, such as employment decisions based on disclosed identity, the paper's validity claim would not transfer to that domain, so practitioners need a diagnostic per setting."],"forward_implications":["Researchers may use imputed gender to measure traditional-sexism disparities, such as the share of perceived women in a workplace or film, without contradicting the ethical ban on gender prediction, because the ban concerns legitimacy, not validity.","Imputation scores should be kept continuous and aggregated, never discretized into fixed gender labels or used for individual-level decisions, because that preserves convergent validity with self-reported measures.","The target concept must be narrowly scoped to the specific gender perception at stake, such as \"perceived gender from a written name,\" and imputation is unjustified whenever self-reported or other legitimate data could be obtained instead.","Oppositional sexism—harms targeting transgender and nonbinary people—cannot be measured by gender imputation, so new methods are needed to evaluate it, and its neglect is itself a form of harm.","Even valid imputation remains illegitimate, so deployment requires weighing benefits against harms and minimizing the restrictions on gender self-determination."],"supporting_citations":[{"why":"Supplies the distinction between traditional sexism and oppositional sexism, which organizes the paper's validity argument.","marker":"Serano 2007"},{"why":"Provides the background-concept and systematized-concept framework for measurement validity that defines what \"valid\" means in the paper.","marker":"Adcock and Collier 2001"},{"why":"States the categorical rejection of trans-inclusive gender prediction that the paper must accommodate on the legitimacy side.","marker":"Keyes 2018"},{"why":"Shows commercial facial analysis reduces gender to presentation, grounding both the illegitimacy critique and the possibility of measuring perceived gender.","marker":"Scheuerman, Paul, and Brubaker 2019"},{"why":"Supplies the four validity barriers for name-based gender prediction that the paper proposes to mitigate by scoping and aggregation.","marker":"Gautam et al. 2024"},{"why":"Provides the precedent that the relevant meaning of a demographic category depends on the use case, which the paper applies to gender.","marker":"Hanna et al. 2020"},{"why":"Shows aggregate disparity measures can be computed from unobserved protected attributes, forming the statistical basis for valid imputation-based measurement.","marker":"Chen et al. 2019"},{"why":"Demonstrates that continuous model outputs mitigate discretization-induced bias in demographic prediction, grounding the convergent-validity recommendation.","marker":"Dong et al. 2025"},{"why":"Supplies the measurement-validity criteria for evaluating generative AI systems, which the paper adapts to gender disparity measurement.","marker":"Wallach et al. 2025"}],"fun_headline_variants":["Imputed gender can measure sexism but never transphobia","Gender imputation: illegitimate yet valid for sexism audits","The legitimacy-validity split in gender prediction","Why gender imputation can be valid despite being wrong","Imputed labels: valid for disparity measures, not transphobia"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that traditional sexism is triggered by perceived gender and social position, and that the cues an imputation model uses—names, faces, appearance—are the same cues that trigger sexist discrimination, so aggregate imputed perception tracks the real disparity.","fun_headline_variants_meta":{"raw":{"variants":["Imputed gender can measure sexism but never transphobia","Gender imputation: illegitimate yet valid for sexism audits","The legitimacy-validity split in gender prediction","Why gender imputation can be valid despite being wrong","Imputed labels: valid for disparity measures, not transphobia"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000343,"raw_usage":{"total_tokens":1947,"prompt_tokens":1066,"completion_tokens":881,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":682,"completion_tokens_details":{"reasoning_tokens":800}},"tokens_in":682,"tokens_out":881,"duration_ms":8617,"temperature":1.0,"reasoning_tokens":800,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T10:50:17.842265+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A correspondence audit could settle it: submit matched resumes or applications in which targets share the same gender identity but differ in perceived-gender cues, such as name and photo, and record real decision-makers' responses. If callback or hiring rates track perceived gender, the validity premise holds; if they track identity instead, the premise fails. A second check compares imputed-perceived-gender disparity estimates against self-reported-identity estimates in the same population; systematic divergence in settings where identity-based discrimination is known to operate would falsify the claim that imputation validly measures sexism there.","supporting_citations":[],"review_version":1}