{"id":"21119728-0c80-4190-9515-610aa6b09ecc","arxiv_id":"2608.12669","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Generative AI representations should be judged by whether they undermine equal participation in society, not by whether they are descriptively accurate.","lead":"This paper argues that the real problem with generative AI's depictions of social groups is not factual inaccuracy but misrecognition, meaning representations that deny people equal standing in society. It proposes Nancy Fraser's idea of participatory parity as the standard for judging when AI outputs are unjust, beyond accuracy.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's proposed standard, participatory parity, is deliberately non-algorithmic and assigned to public contestation, but the paper does not show how contestation can adjudicate contested cases or correct for power asymmetries; this gap threatens the claim that recognitional justice offers…","rationale":"The paper is a serious, well-grounded conceptual argument. It correctly identifies three genuine limitations of accuracy-based representational fairness: accurate outputs can be status-reinforcing, authority over accuracy is contested, and social groups are not stable targets. The turn to Fraser's two-dimensional theory is apt and well explained. Independent support includes the use of recognized political theory, not machine-checked proofs or reproducible code, so the argument stands or falls on the strength of its normative inference. The load-bearing step is the move from 'accuracy is insufficient' to 'participatory parity should be the standard.' This step is insecure because the paper's own statements show the proposed standard is underdetermined. The authors write that 'there is no objective marker' and that judgments 'must be worked out discursively and dialogically'; when discussing the doctors/nurses example they note that 'there are reasonable arguments to be made against such an approach too. The issue would need to be decided collectively.' These admissions are not fatal in themselves—a normative standard can be non-algorithmic and still orient deliberation—but the paper does not specify the conditions under which collective decision-making would be legitimate and non-dominating. That is precisely the same authority problem the paper levels against accuracy-based approaches, merely relocated from 'who decides what is accurate?' to 'who decides what undermines parity?' Because the central claim is prescriptive for AI governance, the absence of an account of the public reason process is a real weakness. My test would be to force the framework to adjudicate its own central example; if it cannot do so, the claim of better normative tools is premature. This aligns with the reader's weakest assumption, though I would phrase it as an under-specified legitimacy condition rather than pure indeterminacy. The verdict should remain CONDITIONAL: the paper makes a valuable contribution as a conceptual reframing, but its practical governance claim is not yet supported.","tokens_in":16885,"tokens_out":7369,"duration_ms":84700,"concrete_test":"Analytic test: reconstruct the argument from 'accuracy is not the right criterion' to 'parity of participation is the right standard' and pin down the needed premise—that parity can be applied to generative AI outputs without importing exactly the contested group-boundary and authority judgments the paper uses to reject accuracy. Apply the two-level test to the paper's own doctors/nurses example and specify what evidence would show the accurate output 'contributes to a hierarchy of social status.' If the framework cannot adjudicate this canonical case without adding premises about which patterns are unjust or which public is authoritative, the central claim fails to deliver better normative tools.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Load-bearing concern: the central claim—that generative AI should be judged by Fraser's participatory parity rather than representational accuracy—requires that the standard be applicable to concrete outputs. In 'A Two-Dimensional Theory of Justice', the authors concede that 'there is no objective marker that signals when parity of participation has been achieved' and that judgments 'must be worked out discursively and dialogically.' They also concede, in the doctors/nurses example, that there are 'reasonable arguments' on both sides and 'the issue would need to be decided collectively.' This is an admission that the standard underdetermines the very cases it was introduced to resolve. The paper does not specify who participates in the required public contestation, how disagreements are settled, how to prevent already-dominant groups from controlling the process, or when a verdict is legitimate. Without such an account, the shift to recognitional justice risks relocating the authority problem it identifies in accuracy-based approaches: instead of an unclear authority to decide what is accurate, we have an unspecified authority to decide what undermines parity. For an AIES audience concerned with governance, this is the weakest point because the paper convincingly argues that accuracy is not the right criterion but does not show that participatory parity is a usable criterion rather than a placeholder for further political theory.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper argues that generative AI's representational harms are best understood not as problems of descriptive accuracy (\"representational fairness\") but as problems of status subordination (\"recognitional justice\"). Drawing on Nancy Fraser's two-dimensional theory of justice, the authors propose participatory parity as the normative standard against which generative AI outputs should be evaluated: the question is whether outputs strengthen or undermine the ability of all members of society to participate as equals. The paper develops three critiques of accuracy-oriented remedies—accurate representations can perpetuate unjust hierarchies, authority over accurate representation is unresolved, and social groups are too contested and fluid to serve as stable referents—and concludes that judgments about recognitional justice must be worked out discursively and dialogically in public.","tokens_in":17106,"tokens_out":3599,"duration_ms":41716,"significance":"If the paper's central claim holds, it provides a valuable conceptual bridge between political theories of recognition and the fair AI/value alignment literatures, and it sharpens the diagnosis of why accuracy-based fixes are insufficient. The paper is clearly written, carefully structured, and engages responsibly with current work on cultural and pluralistic alignment. Its strengths include a concrete set of examples (doctors/nurses, cultural erasure, persona prompting), a fair characterization of alternative views, and an explicit acknowledgement of the limits of its own proposal. Its main weakness is that the proposed positive standard—participatory parity—is left at a high level of abstraction: the paper does not specify the procedures, participants, or legitimacy conditions for the public contestation it says is required. This makes the contribution primarily diagnostic and programmatic rather than a fully specified governance framework, and it leaves open whether recognitional justice is genuinely more actionable than the accuracy standard it replaces.","major_comments":[{"comment":"The central claim—that generative AI should be held to the standard of participatory parity rather than representational accuracy—requires that the standard be applicable to concrete outputs. The paper explicitly concedes that \"there is no objective marker that signals when parity of participation has been achieved\" and that judgments \"must be worked out discursively and dialogically,\" but it does not specify who participates in this contestation, how disagreements are settled, how to prevent dominant groups from controlling the process, or when a verdict is legitimate. The doctors/nurses example illustrates the gap: after noting \"reasonable arguments\" on both sides, the paper says only that \"the issue would need to be decided collectively.\" Given that the paper's stated goal is to provide \"better conceptual and normative tools for governing\" generative AI, this underdetermination is load-bearing. The authors should either sketch the minimal procedural and legitimacy conditions for public contestation or explicitly revise the claim to say that participatory parity is a diagnostic ideal rather than a governance standard.","section":"A Two-Dimensional Theory of Justice"},{"comment":"The paper says that \"the overarching argument we are advancing in this paper doesn't rest on the details of Fraser's account,\" yet the Conclusion identifies recognitional justice with participatory parity, a distinctly Fraserian notion. This creates a tension: if the argument does not depend on Fraser's specific framework, what is the minimal content of recognitional justice that remains? If it does depend on Fraser's framework, the paper should engage more directly with well-known objections to participatory parity, such as the difficulty of resolving incommensurable value claims in pluralistic societies. The authors should clarify the relation between Fraser's account and their own normative proposal.","section":"A Two-Dimensional Theory of Justice; Conclusion"},{"comment":"The critique of participatory approaches in the discussion of authority notes that institutions can retain control over participant selection, scope of deliberation, and final outcomes, and that participation can be tokenistic. The paper responds that \"this limitation does not make participation irrelevant\" because it exposes political choices. This is a reasonable point, but it does not answer the analogous power-asymmetry problem for the paper's own recommended remedy: public contestation over what undermines parity of participation. In the absence of any account of how marginalized groups can secure a meaningful voice in that contestation, the proposal risks relocating rather than solving the authority problem it identifies in accuracy-based approaches. A brief discussion of conditions under which public contestation can be expected to be non-dominated would materially strengthen the argument.","section":"From Fair Representation to Just Recognition"}],"minor_comments":[{"comment":"There is a typo in the Introduction: \"with respesct\" should read \"with respect.\"","section":"Introduction"},{"comment":"The reference to Rauh et al. contains \"????\" as the year and lacks complete publication details; it should be completed or removed.","section":"References"},{"comment":"The phrase \"On this 'consumerist view' view\" contains a duplicated word; it should read \"On this 'consumerist view',\".","section":"Distribution versus Recognition"},{"comment":"The paper uses \"recognitional justice\" and \"just recognition\" interchangeably in places; fixing on a single term would improve readability.","section":"From Fair Representation to Just Recognition"}],"recommendation":"major_revision","confidential_remarks":"The paper is a well-crafted position piece with a clear contribution to the AI ethics literature, and I see no technical circularity. My main reservation is that the central constructive claim is under-specified in a way that affects its actionability. The revision path is clear: either specify the procedural conditions for public contestation or moderate the claim that participatory parity provides better 'tools for governing' generative AI. I would be willing to review a revised version."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nWhat you should know: this is a conceptual, not empirical, paper, and it is a good one. There is no data or code to check; the work stands or falls on argument. It makes a genuine synthesis: apply Nancy Fraser's participatory parity to generative AI representational harms, rather than accuracy-based fairness. The three critiques of the accuracy frame — contested group boundaries, unclear authority, and harmful accurate representations — are clearly argued and well illustrated with existing work. The connection between fair ML's distributive/representational split and political theory's redistribution/recognition split is done carefully and without caricature. The citation pattern is fair: it cites critical, participatory, and pluralistic alignment work, and it does not oversell novelty.\n\nThe main soft spot is exactly what the stress-test note says. The paper's positive standard, participatory parity, is deliberately non-algorithmic and is referred to public contestation, but the paper does not specify who contests, how disagreements are settled, or how to keep dominant groups from shaping the outcome. The doctors/nurses example is honest — the authors say reasonable arguments exist on both sides and the issue must be decided collectively — but that concession shows the standard underdetermines the contested cases it was introduced to resolve. The paper also assumes Fraser's framework rather than defending it against alternatives; that is fine for an application paper, but it means the central claim is conditional on accepting that framework.\n\nI do not think this is fatal. As a reframing, the paper succeeds: it convincingly shows that accuracy is neither necessary nor sufficient for representational justice in generative AI, and it gives the field a better vocabulary for asking the political question. But the gap is real, and the authors should be pressed to say more about governance and contestation before this becomes an actionable standard.\n\nFor whom: AI ethics researchers, value alignment people, and anyone working on representational harms. It is a useful reading-group piece and worth citing for the critique of accuracy. I would send it to peer review rather than desk-reject; it is exactly the kind of conceptual clarification that a venue like AIES should publish, with a request that the authors add a short discussion of the contestation gap.","headline":"A clear conceptual case for replacing accuracy-based representational fairness with Fraser's participatory parity in generative AI, but the standard's practical operation is left to an unspecified public contestation — worth publishing, with that limit named.","tokens_in":17607,"tokens_out":2846,"would_cite":true,"duration_ms":31786,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that generative AI outputs should be held to the standard of recognitional justice—whether they support parity of participation—rather than representational accuracy.","keywords":["representational fairness","recognitional justice","generative AI","parity of participation","value alignment","cultural pluralism","participatory governance","AI ethics"],"falsifier":"Show a contested AI depiction of a social group where a well-structured public deliberative process either cannot converge on whether the depiction undermines equal standing, or converges only because dominant groups out-vote or out-shout the affected group. For example, run a structured citizens' jury on an image generator's depiction of a religious minority's symbols and observe whether reasoned argument changes verdicts; if the process stays polarized, the account's remedy lacks content.","tokens_in":16684,"feed_emoji":"⚖️","tokens_out":7728,"duration_ms":75761,"temperature":0.7,"pith_summary":"Generative AI's expressive power makes questions of how it depicts social groups central to AI fairness, but existing approaches evaluate that depiction for accuracy. This paper argues that accuracy is the wrong target: the real harm is misrecognition, and the standard should be recognitional justice, defined by whether outputs allow all members of society to participate as equals. It shows that even accurate representations can entrench unjust hierarchies, that social groups lack stable boundaries against which accuracy can be measured, and that no one has uncontested authority to decide what counts as misrepresentation. The upshot is that deciding whether an AI output is just cannot be automated or settled by experts; it must be worked out through public contestation, so fairness research should treat accuracy as diagnostic only.","feed_headline":"Generative AI should be judged by equal participation, not accuracy","feed_subtitle":"The norm for generative AI is whether portrayals let all people participate as equals, not whether they are accurate.","key_machinery":"The load-bearing object is the norm of participatory parity, drawn from the two-dimensional theory of justice, which treats just distribution and just recognition as separate but co-constitutive conditions for equal social standing. In the paper's use, the norm supplies a two-level test: an AI output, or a proposed remedy, is unjust if it lowers one group's standing relative to others, or if the remedy subordinates some members of the group or others. The key move is that the norm is non-monological: it cannot be computed by an algorithmic metric, and its application must be worked out discursively and dialogically through public contestation. This is what carries the argument from accuracy-based evaluation to recognitional justice.","core_discovery":"The paper's central claim is that the normative standard for generative AI's representations of social groups should be recognitional justice rather than representational fairness. The deep question is not whether an output is accurate but whether it strengthens or undermines parity of participation—the ability of all members of society to interact with one another as peers. The paper defends this by showing that accuracy-oriented remedies fail in three ways: accurate representations can reproduce unjust social patterns; who decides representational accuracy is politically contested; and fixing a group as a target of optimization essentializes an internally plural, changing culture. Accordingly, inaccuracy is at most a diagnostic indicator of possible misrecognition, not the harm itself. The remedy is not more faithful data or better measurement but collective, contestatory public reasoning about status and equal standing.","pith_inferences":["A natural next step the paper does not develop is to turn the parity standard into a procedural audit: test whether affected communities can contest model outputs and get them changed.","If the argument is right, a model that reproduces majority self-understandings with high accuracy may fail recognitional justice even when it scores well on pluralism benchmarks weighted by population frequencies.","The same logic could be applied to quality-of-service gaps: when a model serves marginalized language speakers worse, the harm is not only differential quality but a denial of equal standing, entangling distribution and recognition in ways the paper only footnotes.","A testable empirical extension would compare the status effects of accuracy-optimized outputs versus outputs shaped through deliberative, participatory processes."],"forward_implications":["AI fairness evaluations of generative models should treat accuracy as a diagnostic clue, not as the objective to optimize.","Interventions should be judged by whether they reduce status subordination, so deliberately non-accurate representations can be just when they repair hierarchies.","Participatory governance moves from a nice-to-have to a necessary condition: who is included, who sets the scope, and whether input changes the model become core fairness questions.","Cultural and pluralistic alignment benchmarks that rate fidelity to survey data are insufficient; they must also ask whose standing the output supports.","The same output can be just or unjust depending on context, so fairness claims cannot be reduced to a static distribution of attributes."],"supporting_citations":[{"why":"Supplies the norm of participatory parity and the two-dimensional account of justice that the whole argument adopts as its standard.","marker":"Fraser et al. 2003"},{"why":"Extends the recognition account to globalized, pluralistic societies, framing the social conditions generative AI operates in.","marker":"Fraser 2008"},{"why":"Establishes the distribution/representation distinction in fair ML that the paper rebalances for generative AI.","marker":"Barocas, Hardt, and Narayanan 2023"},{"why":"Grounds the critique that 'bias' in language technology is under-conceptualized and not captured by accuracy.","marker":"Blodgett et al. 2020"},{"why":"Provides the taxonomized examples and the point that accurate outputs can reproduce harmful hierarchies.","marker":"Katzman et al. 2023"},{"why":"Supplies the taxonomy of representational harms that the paper takes as useful but incomplete.","marker":"Corvi et al. 2025"},{"why":"Documents cultural erasure in language models, the clearest case of omission as recognitional harm.","marker":"Qadri et al. 2025"},{"why":"Supports the claim that cultures are polyvocal and contested, so no stable accuracy target exists.","marker":"Benhabib 1999"},{"why":"Provides the methodological parallel that optimizing a formal target cannot by itself establish justice.","marker":"Green 2022"},{"why":"Argues statistical probability smoothing privileges dominant cultural expressions, used to show accuracy optimization risks essentialism.","marker":"Mwesigwa 2025"}],"fun_headline_variants":["AI's real fairness test: equal status, not accurate portrayals","For generative AI, recognition beats representation","Judging AI by parity: when accuracy isn't the harm","From accuracy to recognition: fixing AI's status problem","Generative AI should measure equal standing, not likeness"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's argument stands or falls on the premise that public debate can actually settle when a representation denies people equal standing, despite there being no objective test for equal standing.","fun_headline_variants_meta":{"raw":{"variants":["AI's real fairness test: equal status, not accurate portrayals","For generative AI, recognition beats representation","Judging AI by parity: when accuracy isn't the harm","From accuracy to recognition: fixing AI's status problem","Generative AI should measure equal standing, not likeness"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000742,"raw_usage":{"total_tokens":3293,"prompt_tokens":909,"completion_tokens":2384,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":525,"completion_tokens_details":{"reasoning_tokens":2306}},"tokens_in":525,"tokens_out":2384,"duration_ms":17625,"temperature":1.0,"reasoning_tokens":2306,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:54:28.072909+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Show a contested AI depiction of a social group where a well-structured public deliberative process either cannot converge on whether the depiction undermines equal standing, or converges only because dominant groups out-vote or out-shout the affected group. For example, run a structured citizens' jury on an image generator's depiction of a religious minority's symbols and observe whether reasoned argument changes verdicts; if the process stays polarized, the account's remedy lacks content.","supporting_citations":[],"review_version":1}