{"id":"fb812b81-0ab2-4934-8fa4-d1e4e132a646","arxiv_id":"2504.18807","paper_version":1,"verdict":"UNVERDICTED","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A feminist position paper argues that researcher-driven digital cloning needs continuous, explicit user consent and decentralized data governance.","lead":"This paper critiques digital cloning in academic research, arguing that simulating users from scraped data raises consent and representation problems. It proposes decentralized data repositories and dynamic consent dashboards as more ethical alternatives.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central policy claim rests on the unsupported premise that digital clones 'simulate behaviors and agency'; if clones are merely statistical predictors of public data, the argument collapses into a generic privacy concern and the proposed human-participant review is unmotivated or overbroad.","rationale":"The reader and I identify the same weakest link: the agency-simulation premise is load-bearing and unsubstantiated. I agree with the UNVERDICTED verdict because the paper is a self-described position piece, not an empirical claim, and our concern does not change that classification. The concern is not that the argument is internally inconsistent; it is that the policy recommendation is conditional on a factual characterization of the technology. The substitution test would settle whether the premise actually carries the argument or whether the conclusion can be reached via weaker, well-established data-ethics premises. This is a good-faith critique: the paper is clearly written, engages relevant feminist literature, and its proposals (dynamic consent, decentralized repositories) are reasonable. The issue is the gap between 'predicting behavior from data' and 'simulating agency'—the very gap the paper needs, and needs to justify, to motivate a distinct ethical framework.","tokens_in":3935,"tokens_out":6130,"duration_ms":63362,"concrete_test":"Perform a substitution test on Sections 1–4: replace every occurrence of 'digital clone' with 'statistical model trained on public user data' and check whether the argument and the Section 6 policy conclusion survive unchanged. If they survive, the agency-simulation premise is doing no work and the paper's specific contribution collapses into general data ethics. If they do not survive, identify which harms disappear under the substitution, then examine Puri et al. [6] (or run the simulation) to confirm that those harms are actually instantiated by cloned agents rather than merely asserted. In particular, check whether cloned agents have goals/beliefs and produce outputs attributable to specific users; if they only fit aggregate distributions, the premise fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's conclusion in Section 6—that digital cloning research should be exposed to human-participant ethical approval—depends on the assertion in Section 1 that 'digital clones do more than aggregate data—they simulate behaviors and agency' and on Section 2's claim that clones 'actively simulate user agency, predicting behaviors and interactions.' This is the premise that makes cloning ethically different from ordinary analysis of public data. The paper offers no operational definition of agency and no direct evidence that the cited systems satisfy it. Puri et al. [6] is an agent-based model with LLM components; Schmidt et al. [7] generated LLM personas. Whether these constitute 'simulation of agency' as opposed to statistical extrapolation is not established. Toupin [10] is invoked for the claim that digital footprints are extensions of the self, but that is a normative position, not an empirical demonstration that clones are agents. The paper self-identifies as a position paper (Section 2), so empirical proof is not required for every step; however, the policy intervention it proposes has a factual predicate that should be at least precisely specified and supported by examples that genuinely instantiate it. If clones are just predictive models, the unique harms described—identity misrepresentation, manipulation, persistent persona—either reduce to familiar privacy/bias concerns or need new evidence.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper is a four-page position paper arguing that researcher-driven digital cloning—the creation of computational agents that replicate users' digital histories to simulate and predict behavior—raises ethical concerns that are not captured by existing data-protection and human-subjects frameworks. Drawing on feminist theory (Suchman, Toupin, Nunes), the author contends that clones are not merely aggregates of data but active simulations of user agency, and that using public data to build them without explicit consent misrepresents users, flattens contextual complexity, and risks reinforcing systemic biases. The paper proposes two remedies: decentralized, ethically governed data-donation repositories and dynamic consent models with participatory dashboards. It concludes that digital cloning research should be subject to human-participant ethical approval.","tokens_in":4161,"tokens_out":3923,"duration_ms":43391,"significance":"If accepted, the paper's central recommendation would require HCI and adjacent fields to classify simulation studies that clone users as human-subjects research, with concrete consequences for IRB review, data governance, and consent infrastructure. The paper is clearly written and well-sourced, and it connects a timely methodological practice (LLM-driven personas and agent-based simulations) to a substantive ethical debate. Its main strength is its synthesis of feminist critiques of agency into a concrete call for procedural change. However, the argument is built on a factual premise about what digital clones do that is not specified or empirically supported, and the proposed policy intervention lacks a clear scope. The paper is an honest position piece, but the central claim needs to be sharpened before the recommendation can be assessed.","major_comments":[{"comment":"The claim that digital clones 'do more than aggregate data—they simulate behaviors and agency' is load-bearing for the entire argument, but the paper gives no operational definition of agency and no direct evidence that the cited systems (Puri et al. [6], Schmidt et al. [7]) instantiate it. If clones are only sophisticated statistical predictors, the specific harms described (identity misrepresentation, manipulation, persistent persona) reduce to familiar privacy or bias concerns, and the proposed human-participant review becomes either unmotivated or overbroad. Please specify the minimal factual premise required for the ethical argument—e.g., that clones are interactive, persistent, and identifiable—and either defend that premise in the cited examples or reframe the argument so it does not depend on the contested term 'agency.'","section":"Sections 1 and 2"},{"comment":"The central recommendation—'Digital cloning research should be exposed to human-participant ethical approval'—presupposes a clear boundary between digital cloning and ordinary analysis of public data. The paper does not define what counts as a digital clone for regulatory purposes, beyond saying it 'actively simulate[s] user agency.' Without an operational boundary, the recommendation either covers all machine learning on public user data or is unenforceable. Please provide criteria (e.g., interactivity, persistence, identifiability, use in simulation) that distinguish cloning from standard data mining and make the scope of the policy concrete.","section":"Sections 2 and 6"},{"comment":"The proposal of dynamic consent models for digital cloning does not address the feasibility of retrospective consent for existing large-scale datasets. Puri et al. [6], the paper's central example, cloned histories of over 10,000 users scraped from public platforms; it is not explained how researchers would contact all users to obtain consent or what should happen to already-built clones if consent is withheld. Without a transition plan for existing data, the proposal is incomplete as a policy recommendation. Please address how dynamic consent would apply retrospectively.","section":"Section 5"}],"minor_comments":[{"comment":"The title contains 'Digit al Cloning' and the affiliation 'Univers ity of Amsterdam'; these spacing artifacts should be fixed in the camera-ready version.","section":"Title and author block"},{"comment":"The description of Puri et al. [6] states that the ethics board deemed explicit consent unnecessary due to public availability, but the paper does not say whether the user data was anonymized before cloning; this detail matters because the argument assumes clones are linked to identifiable individuals.","section":"Section 2"},{"comment":"The claim that clones 'may overemphasize frequently repeated behaviors while overlooking passive or evolving user interactions, reinforcing echo chambers and distorting online discourse' is an empirical hypothesis without citation; it should be explicitly marked as a projection or supporting example rather than a demonstrated effect.","section":"Section 4"},{"comment":"The 'decentralized data donation repositories' are described as operating 'exclusively for non-commercial academic research,' but the paper does not discuss governance, funding, or how community oversight would be enforced; a sentence on these practicalities would strengthen the proposal.","section":"Section 5"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a short position paper that would be suitable for a venue like CHI, and the topic is timely. The central policy recommendation, however, rests on a factual premise that needs to be specified and supported; without that, the recommendation is difficult to evaluate. I would support publication after the identified load-bearing issues are addressed in revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a quick read if you care about research ethics for LLM personas or agent-based models of social media. The paper is not a new empirical or formal result; it is a synthesis of feminist HCI/critical AI arguments applied to a specific practice, and on those terms it is clear and mostly well sourced.\n\nWhat it does well: the distinction between scraping public data and building a persistent, interactive replica of a person is genuinely important and not just a privacy restatement. The Meta disclaimer and Reddit/Stack Overflow protest examples make that concrete. The paper is honest that it is a position statement, and it does not oversell its originality.\n\nWhere it is soft: the load-bearing premise is that digital clones 'simulate behaviors and agency' rather than just statistically extrapolate from data. That premise appears in Section 1, is repeated in Section 2, and carries the conclusion in Section 6. The paper never defines agency operationally and gives no evidence that the systems it cites (Puri et al., Schmidt et al.) instantiate it. If clones are just pattern recognizers, then the specific harm—identity misrepresentation, persistent persona—collapses into familiar privacy/bias concerns, and the call for human-participant review becomes overbroad. The stress-test note got this right.\n\nTwo more minor issues: the causal claims about echo chambers and overemphasis on frequent behaviors are stated without support, and the proposed solutions (decentralized repositories, dynamic consent dashboards) are drawn directly from the cited literature, so the paper is a synthesis rather than a new design. None of this is fatal for a position paper, but the author should be asked to either define the sense of agency she needs or reframe the argument to survive the collapse of that premise.\n\nBottom line: the paper is a legitimate contribution to the conversation, arguably worth a slot in a workshop or short-paper venue. It deserves a serious referee, on the condition that the referee pushes on the agency premise. I would not cite it in my own empirical work, but I would cite it if I were writing on ethics of synthetic participants. Bring it to reading group if you want a focused discussion of whether 'digital cloning' is a distinct ethical category.","headline":"A short, well-written position paper that makes a plausible ethical case for consent in digital cloning research, but its policy conclusion rests on an unargued premise about what clones actually simulate.","tokens_in":4622,"tokens_out":2093,"would_cite":false,"duration_ms":19382,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Researcher-driven digital cloning should be governed like research with human participants.","keywords":["Digital Clones","User Agency","Feminist HCI","Consent","Simulation Studies","Ethical AI","AI solutionism","Research ethics"],"falsifier":"Conduct a controlled comparison in which a digital clone and a demographic baseline each predict a user's responses in context-sensitive situations, such as expressing political views in private versus public settings. If the clone's predictions are no more accurate or context-sensitive than the baseline, then clones have not been shown to simulate agency rather than aggregate patterns, and the paper's central premise is empirically weakened.","tokens_in":3738,"feed_emoji":"🤖","tokens_out":8747,"duration_ms":73377,"temperature":0.7,"pith_summary":"This position paper argues that researcher-driven digital cloning—building computational agents that simulate users from scraped public data—is ethically distinct from ordinary data analysis, because clones simulate behavior and agency rather than merely aggregate data. The paper contends that treating public data as freely usable without consent obscures this difference, and risks flattening human complexity, reinforcing systemic biases, and denying users control over persistent digital replicas of themselves. Drawing on feminist theories of relational agency, it proposes decentralized data donation repositories and dynamic consent dashboards as correctives. If the paper is right, research ethics boards should treat cloning studies like human-subjects research and require explicit consent even when the source data is publicly available.","feed_headline":"Treat digital cloning like research with human participants","feed_subtitle":"A feminist critique argues clones simulate user agency, so using public data without consent is ethically inadequate.","key_machinery":"The central object is the 'digital clone'—a computational agent that replicates a user's digital history, such as posts and interactions, to simulate and predict behavior. The load-bearing mechanism is a feminist relational account of agency, which holds that agency does not reside in discrete agents but emerges from sociomaterial arrangements; on this view, a clone built from a user's data functions as an extension of that user's agency. This substitution—data replica as agentic extension—is what converts a privacy concern into a consent-and-representation concern that warrants human-subjects oversight.","core_discovery":"The central claim is that digital clones, as used in simulation studies, simulate user agency and therefore cloning is not a form of passive data collection. Because agency is relational and context-dependent—shaped by power structures, gender, and social position—the paper argues that clones which strip away context misrepresent users, especially marginalized users, and can perpetuate systemic biases. Researcher-driven cloning that scrapes public data without explicit consent therefore fails the ethical standards expected of research with human participants. The paper concludes that digital cloning research should be subject to human-participant ethical approval, and recommends governance mechanisms such as dynamic consent and decentralized, non-commercial data repositories.","pith_inferences":["The relational-agency premise, if accepted, extends beyond academic simulation studies to commercial LLM-persona research, implying that companies building user simulations from public text would also face a consent obligation.","A testable extension of the proposed governance is whether participants in voluntary donation repositories behave differently from users whose data is scraped, which would indicate whether consent changes the fidelity of simulated behavior.","The argument implies a legal corollary: a clone is a new derived artifact, not the original data, so existing data-protection regimes may already require consent for cloning even where scraping the raw data is lawful.","If the paper is right, the ethical burden also shifts to the designers of simulation frameworks, who would need to build context-awareness and consent mechanisms into the tools themselves, not just the studies that use them."],"forward_implications":["Research ethics boards would classify researcher-driven digital cloning as human-subjects research, requiring informed consent even when the source data is publicly available.","Simulation studies that clone users would have to justify any waiver of explicit consent, or switch to user-driven donation repositories with transparent terms.","Persistent clones would need deletion and withdrawal mechanisms, honoring users' right to be forgotten after they revoke consent.","Dynamic consent dashboards—notifying users of data use, outcomes, and withdrawal options—would become standard infrastructure for behavior-simulation research."],"supporting_citations":[{"why":"Supplies the central case study: cloning a misinformation-sharing network of over 10,000 users without explicit consent.","marker":"[6]"},{"why":"Exemplifies the AI-solutionist practice of using LLM-generated personas to replace human participants.","marker":"[7]"},{"why":"Provides the relational-agency theory that grounds the claim that clones extend user agency rather than merely represent data.","marker":"[9]"},{"why":"Supports the claim that digital footprints are extensions of the self, especially for marginalized communities.","marker":"[10]"},{"why":"Supplies the posthuman feminist framework for how AI systems shape and manipulate user behavior.","marker":"[5]"},{"why":"Establishes deepfakes as precedent for identity-replication harm, used to analogize digital-clone harms.","marker":"[4]"},{"why":"Informs the proposed dynamic consent dashboard mechanism for transparency and user control.","marker":"[1]"}],"fun_headline_variants":["Your data clone needs your consent","Digital clones aren't passive data","Cloning users? Treat it like human research","Agency simulated, consent ignored"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire argument depends on the premise that digital clones do more than aggregate data—that they actively simulate behavior and agency; if clones are merely statistical pattern-matchers with no meaningful simulation of agency, the distinct harm the paper identifies collapses into a generic privacy concern.","fun_headline_variants_meta":{"raw":{"variants":["Your data clone needs your consent","Digital clones aren't passive data","Cloning users? Treat it like human research","Agency simulated, consent ignored"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000131,"raw_usage":{"total_tokens":1029,"prompt_tokens":749,"completion_tokens":280,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":365,"completion_tokens_details":{"reasoning_tokens":230}},"tokens_in":365,"tokens_out":280,"duration_ms":3684,"temperature":1.0,"reasoning_tokens":230,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:08:02.376806+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Conduct a controlled comparison in which a digital clone and a demographic baseline each predict a user's responses in context-sensitive situations, such as expressing political views in private versus public settings. If the clone's predictions are no more accurate or context-sensitive than the baseline, then clones have not been shown to simulate agency rather than aggregate patterns, and the paper's central premise is empirically weakened.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the central case study: cloning a misinformation-sharing network of over 10,000 users without explicit consent."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the relational-agency theory that grounds the claim that clones extend user agency rather than merely represent data."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supports the claim that digital footprints are extensions of the self, especially for marginalized communities."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the posthuman feminist framework for how AI systems shape and manipulate user behavior."}],"review_version":1}