{"id":"bb7b0b5a-9900-42cd-b5aa-ae9ef8916f11","arxiv_id":"2505.05298","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper advocates designing LLMs as 'reasonable parrots' that challenge users through argumentative dialogue to enhance critical thinking.","lead":"This position paper proposes that large language models should be designed to argue with users, not just answer them, in order to strengthen critical thinking. It introduces 'reasonable parrots', multi-persona systems grounded in argumentation theory, and illustrates the idea with prototype dialogues.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The proposal's load-bearing premise—that argumentative dialogical moves improve users' critical thinking—is asserted without empirical support; the prototype shows only that LLMs can produce such moves, not that users benefit.","rationale":"The reader identified the weakest assumption correctly, and I agree with the CONDITIONAL verdict. The paper is an honest position paper that explicitly calls its prototype a sketch, so there is no internal contradiction or formal defect to flag. The load-bearing step is where the normative design claim meets social science: 'argue with us by design' is only valuable if users actually benefit from argumentative dialogue. The cited evidence does not establish that benefit—Costello et al. is domain-specific belief updating, and Ma et al. concerns decision accuracy in a constrained task, not durable critical-thinking gains from multi-parrot moves. The prototype demonstrates feasibility of generating dialogical moves, not efficacy. This is an addressable empirical gap rather than a reason for rejection, so the paper should remain CONDITIONAL, pending the proposed test. No change to the reader's verdict is needed.","tokens_in":9084,"tokens_out":4935,"duration_ms":54269,"concrete_test":"Run a preregistered between-subjects experiment (N≥120 per arm) in which participants discuss a contested issue under three conditions: (A) the four-parrot reasonable-parrots system described in §4, (B) a standard LLM Q&A interface, and (C) a self-reflection writing prompt without AI. Measure pre/post changes in (i) argument quality scored by blind raters (number and quality of reasons, counterarguments addressed, fallacy avoidance), (ii) performance on an unrelated argument-evaluation task, and (iii) engagement: completion rate, number of turns, and self-reported willingness to continue. Also measure perceived adversarialness. If A does not significantly outperform B and C on (i)–(iii), or shows higher dropout, the causal premise in §4 is not supported and the design recommendation needs empirical revision.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central normative claim depends on an empirical causal premise stated in §4: 'reasonable parrots are meant to trigger improved reasoning skills in their interlocutor, regardless of their performance.' For this premise to hold, users must stay engaged when challenged, perceive the challenge as informative rather than adversarial, and improve in reasoning skills that transfer beyond the conversation. The paper offers no direct evidence for any of these. Costello et al. (2024) demonstrates belief change for conspiracy beliefs via tailored counterarguments, not acquisition of general critical-thinking skills; Ma et al. (2025) studies constrained deliberation for binary decisions, not the multi-parrot design. The prototype in Tables 2–4 shows only that GPT-4 Turbo, Claude 3.7, and Llama 3.1 can enact Socratic, Cynical, Eclectic, and Aristotelian moves; it does not measure user reasoning, retention, or disengagement. The prompt in Table 1 lets the user 'end the conversation anytime'—a feature for autonomy, but also a plausible failure mode: if challenge triggers reactance or overload, users may exit, making the 'argue by design' mechanism counterproductive. The paper itself notes that Costello et al. 'overlooks the role of individuals' perceptions of AI as a discussant,' yet the prototype likewise does not measure perceived adversarialness or trust. Thus the key reason to prefer reasonable parrots over ordinary Q&A is currently unsupported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This position paper argues that large language models (LLMs) should be designed not as providers of final answers but as interlocutors that engage users in argumentative dialogue, thereby exercising rather than replacing human critical thinking. The authors introduce the concept of 'reasonable parrots' grounded in pragma-dialectical argumentation theory, organized around the principles of relevance, responsibility, and freedom, and instantiated through four parrot personas: Socratic, Cynical, Eclectic, and Aristotelian. After critiquing current LLMs as 'unreasonable' on the basis of two illustrative ChatGPT responses, the paper proposes a multi-parrot dialogue design and demonstrates it with system prompts and short user–parrot transcripts generated by GPT-4 Turbo, Claude 3.7, and Llama 3.1. The paper concludes by calling for a shift from argumentative products to argumentative processes in LLM-based conversational technology.","tokens_in":9532,"tokens_out":2961,"duration_ms":33282,"significance":"The proposal is timely and well-motivated: it brings a rich tradition of argumentation theory into the design of conversational AI and offers a concrete, implementable sketch (the multi-parrot personas) that could inform future HCI research. The paper is honest about the illustrative nature of its examples and does not claim to have solved the problem. Its main contribution is conceptual—reframing LLMs as tools for fostering deliberative skills—and that framing is worth taking seriously. The significance is conditional, however, on empirical validation of the central causal premise that argumentative dialogical moves improve users' critical thinking, and on evidence that users remain engaged rather than disengaging when challenged. The paper currently provides neither, so its value lies primarily in defining a research agenda rather than demonstrating an effective technology.","major_comments":[{"comment":"The central claim that 'reasonable parrots are meant to trigger improved reasoning skills in their interlocutor, regardless of their performance' rests on an empirical causal premise that is asserted without supporting evidence. The prototype in Tables 2–4 demonstrates only that GPT-4 Turbo, Claude 3.7, and Llama 3.1 can follow a prompt instructing them to enact Socratic, Cynical, Eclectic, and Aristotelian moves; it does not measure any effect on the user's reasoning, learning, engagement, trust, or subsequent behavior. The cited studies do not bridge this gap: Costello et al. (2024) measures belief change about conspiracy theories, not acquisition of general critical-thinking skills, and Ma et al. (2025) evaluates constrained binary decisions rather than an open multi-parrot dialogue. As it stands, the reason to prefer 'argue by design' over ordinary Q&A is unsupported.","section":"Section 4, paragraph beginning 'As a caveat' through 'Prototypical Realization'"},{"comment":"The diagnosis of current LLM unreasonableness relies on two hand-picked queries answered by ChatGPT on a single date. The paper explicitly acknowledges that the example is 'not claimed to generalize,' which is appropriate, but the critique is nevertheless assessed against the very ideal critical discussion framework (van Eemeren and Grootendorst, 2003) that is later used to define the proposed design. This makes the evaluation partly circular: current LLMs are judged unreasonable by a standard that the paper itself selects, and the prototype is then judged reasonable by the same standard. To make the argument more robust, the paper should either sample a broader set of queries or state more clearly that the example serves only as an indexical illustration, not as an empirical diagnosis.","section":"Section 3, Query 1 and Response 1"},{"comment":"The multi-parrot dialogues are generated with a system prompt that explicitly instructs the parrots to challenge starting points, rebut arguments, offer alternatives, and point out fallacies. The resulting transcripts therefore show that the models can follow this instruction, not that the reasonable-parrots design is effective or that the observed behavior would arise without such explicit prompting. The claim that 'all models show notable similarities in their approach to user interaction' is a statement about prompt compliance and surface behavior, not about whether the design achieves its goal of improving user critical thinking. The paper should not present these transcripts as evidence for the proposal's benefits; at most they illustrate a feasible interaction pattern.","section":"Section 4, Table 1 and Tables 2–4"},{"comment":"A plausible failure mode of the proposed design is user disengagement: the prompt allows the user to 'end the conversation anytime,' and if challenge is perceived as adversarial, condescending, or cognitively overloading, users may exit exactly when the system is trying to provoke reflection. The paper gives no evidence about user perceptions of the parrots' tone, trustworthiness, or perceived adversarialness, nor about whether the dialogical moves are experienced as informative rather than annoying. This is not a fatal objection in a position paper, but it is a load-bearing uncertainty for the central claim and should be addressed either by explicit acknowledgment as an open research question or by pilot data.","section":"Section 4, Table 1; Section 5 conclusion"}],"minor_comments":[{"comment":"The text contains garbled fragments: 'for developinghci! evaluation metrics' and 'more reasonable hci! (hci!)' appear to be missing spaces or formatting errors; these should be corrected to 'HCI' with proper spacing.","section":"Section 5, conclusion"},{"comment":"There is a punctuation error in the Cynical parrot's turn: the transcript shows 'Cynical parrot:. Ah' with an extra period after the colon; this should be cleaned up.","section":"Table 4"},{"comment":"The phrase 'as an possible way' should read 'as a possible way'; this is a minor grammatical error.","section":"Section 2, sentence about Kiesel et al."},{"comment":"The figure is referenced in Section 1 but has no caption and its labels ('Socratic Eclectic Cynical Aristotelian') are compressed; adding a brief caption and explaining the arrows would improve readability.","section":"Figure 1"}],"recommendation":"major_revision","confidential_remarks":"This is a genuinely useful position paper for a workshop or a venue that welcomes normative design arguments. The main gap is not the lack of a full empirical study—that is reasonable in a position paper—but the fact that the central causal premise is presented as settled ('they are meant to trigger improved reasoning skills') while no evidence or even a detailed falsifiable evaluation plan is offered. I would advise the editor to request that the authors reframe the claim as a hypothesis or research agenda, add a dedicated limitations section, and moderate the abstract's categorical tone. The paper's title and abstract overclaim relative to the evidence, but the underlying idea is defensible and the argumentation-theoretic grounding is a strength."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nHere's my quick take on Musi et al. The paper is worth a look: it offers the most concrete attempt I've seen to turn pragma-dialectics into a design specification for LLM assistants. The four-parrot pattern—Socratic, Cynical, Eclectic, Aristotelian—is genuinely useful as a starting point for prototyping, and the opening critique of ChatGPT's smartphone advice, mapped onto the four stages of a critical discussion, is a nice pedagogical move. The authors are honest that it's a sketch.\n\nThe novelty isn't in adversarial multi-agent debate; that exists. What's new is the normative framing: LLMs should be designed to exercise users' critical thinking rather than to produce conclusions, and the grounding of that in argumentation theory's notion of reasonableness. That's a legitimate contribution to a design conversation.\n\nThe soft spot is exactly where the stress-test note lands. The load-bearing premise—that engaging users with Socratic, Cynical, Eclectic, and Aristotelian moves improves their reasoning skills—is unsupported. The prototype shows that GPT-4 Turbo, Claude, and Llama can enact the moves; it says nothing about whether users benefit, disengage, or learn anything that transfers. The prompt's \"user can end anytime\" feature is a plausible escape hatch when challenge gets uncomfortable. The cited work by Costello et al. and Ma et al. shows adjacent effects, not this design. I don't think the circularity issue is serious: of course a normative framework judges the world by its own standards. The paper's own criteria are the product being proposed.\n\nOne minor editorial thing: the conclusion contains a repeated 'hci!' artifact that should be cleaned up.\n\nBottom line: this is a position paper that does what position papers should do—it frames a problem and offers a testable design hypothesis. The evidence gap is real but not fatal; the paper does not overclaim. I'd send it to peer review: a serious referee will push on the empirical premise and on operationalizing the dialogical moves, which is the right conversation to have. I'd bring it to a reading group on human-AI interaction. For my own work, I'd likely cite it as a reference point for argumentation-aware LLM design.\n\nRecommendation: engage with it—revise, don't reject.","headline":"A solid, well-grounded position paper whose four-parrot design pattern is a real contribution, but whose central empirical premise about improving users' critical thinking remains untested.","tokens_in":9936,"tokens_out":2764,"would_cite":true,"duration_ms":29045,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that large language models should be redesigned as 'reasonable parrots' that argue with users through dialogical moves, to enhance rather than replace critical thinking.","keywords":["large language models","argumentation theory","critical thinking","reasonable parrots","pragma-dialectics","multi-agent dialogue","fallacies","conversational AI"],"falsifier":"Run a randomized experiment in which one group discusses a contested issue with a reasonable-parrot multi-agent system, a control group with a standard answer-giving LLM, and a third group with no AI assistance, measuring critical-thinking disposition and argument quality before and after; if the reasonable-parrot group shows no greater gain than the control group, the central claim is falsified.","tokens_in":8914,"feed_emoji":"🦜","tokens_out":3733,"duration_ms":34561,"temperature":0.7,"pith_summary":"This position paper argues that large language models should be designed to argue with users rather than merely answer them. The authors claim that current LLMs are 'stochastic parrots' that instantiate the ad populum fallacy, echoing popular training data, and that this makes them inadequate for fostering critical thinking. They propose a new design ideal, the 'reasonable parrot', grounded in argumentation theory, that embodies relevance, responsibility, and freedom and interacts through dialogical moves such as doubting, rebutting, and offering alternatives. If the proposal is right, the payoff is conversational technology that trains users' individual and social critical thinking skills instead of replacing them.","feed_headline":"Design LLMs to argue with us, not just answer us","feed_subtitle":"Proposal: LLMs as 'reasonable parrots' that challenge, rebut, and offer alternatives to sharpen thinking.","key_machinery":"The central object is the 'reasonable parrot', a conversational agent guided by the principles of relevance, responsibility, and freedom and by argumentative dialogical moves from pragma-dialectics. The concrete mechanism is a multi-parrot environment: four personas—Socratic (challenges starting points and beliefs), Cynical (rebuts standpoints and arguments), Eclectic (offers alternative perspectives), and Aristotelian (points out fallacies)—interact with the user and with each other to open up space for agreement and disagreement, fostering critical reflection rather than delivering a finished conclusion.","core_discovery":"The central claim is that conversational AI should externalize reasoning by confronting users with diverse argumentative viewpoints, through a multi-parrot system where each parrot embodies a distinct critical role. The paper demonstrates with prototype dialogues that current LLMs can be prompted to play Socratic, Cynical, Eclectic, and Aristotelian personas, and it argues that this process-oriented design, rooted in pragma-dialectical rules for critical discussion, would enhance critical thinking. The contribution is a design principle and an architectural sketch, not an empirical result: LLMs should be judged by whether they improve their interlocutor's reasoning, regardless of the parrot's own performance.","pith_inferences":["The paper's core premise—that being challenged by argumentative dialogical moves improves a user's critical thinking—remains untested; a controlled study measuring critical thinking before and after reasonable-parrot versus standard LLM interaction would settle it.","If the premise holds, reasonable parrots could complement explainable AI by making model reasoning contestable rather than merely transparent, which may increase user trust and scrutiny.","The multi-parrot approach might also improve LLM self-consistency by externalizing disagreement across personas instead of relying on internal chain-of-thought, though the paper does not test this.","A risk the authors leave implicit is user disengagement: if challenging questions feel adversarial, people may abandon the conversation, so the design likely needs to balance challenge with perceived helpfulness."],"forward_implications":["LLM design goals would shift from producing persuasive answers to facilitating an argumentative process, with evaluation metrics based on users' critical thinking gains.","The multi-parrot persona structure can be instantiated in existing LLMs through prompting, as shown across three different models in the paper.","Such technology could explicitly counteract fallacies like the appeal to popularity by questioning whether popularity is a valid reason for belief or action.","The proposed principles give a concrete starting point for building deliberation-support tools in high-stakes domains such as medicine, finance, and human resources.","The design would externalize reasoning by putting diverse viewpoints in front of the user, instead of hiding deliberation inside the model."],"supporting_citations":[{"why":"Supplies the 'stochastic parrots' critique of LLMs that the paper builds on and reframes.","marker":"Bender et al. 2021"},{"why":"Provides the pragma-dialectical model of a critical discussion and its stages, which structure the proposed dialogical moves.","marker":"van Eemeren and Grootendorst 2003"},{"why":"Defines the ad populum fallacy that the paper claims LLMs inherently instantiate.","marker":"Walton 1980"},{"why":"Source of the three principles—relevance, responsibility, and freedom—that characterize reasonable parrots.","marker":"Danesi and Rocci 2009"},{"why":"Frames the goal of augmenting human intellect with hybrid intelligence, which the reasonable-parrot proposal aligns with.","marker":"Akata et al. 2020"},{"why":"Provides empirical evidence that AI-led dialogues can durably change beliefs, motivating the argumentative design.","marker":"Costello et al. 2024"},{"why":"Shows that constraining AI interaction with deliberation principles improves decision accuracy, supporting the proposed interaction design.","marker":"Ma et al. 2025"},{"why":"Represents a prior attempt at AI debate, used to argue that argumentation lies outside the comfort zone of AI.","marker":"Slonim et al. 2021"}],"fun_headline_variants":["LLMs as reasonable parrots that argue to sharpen thought","Design LLMs to challenge, rebut, and offer alternatives","Why LLMs should argue with us, not just answer","Multi-parrot LLMs for your critical thinking","Reasonable parrots: AI that argues to improve reasoning"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The proposal stands on the untested premise that being challenged by argumentative dialogical moves improves a user's critical thinking; if users disengage or learn nothing, the reasonable-parrot design loses its purpose.","fun_headline_variants_meta":{"raw":{"variants":["LLMs as reasonable parrots that argue to sharpen thought","Design LLMs to challenge, rebut, and offer alternatives","Why LLMs should argue with us, not just answer","Multi-parrot LLMs for your critical thinking","Reasonable parrots: AI that argues to improve reasoning"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000808,"raw_usage":{"total_tokens":3472,"prompt_tokens":797,"completion_tokens":2675,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":413,"completion_tokens_details":{"reasoning_tokens":2596}},"tokens_in":413,"tokens_out":2675,"duration_ms":18725,"temperature":1.0,"reasoning_tokens":2596,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T23:06:27.293729+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a randomized experiment in which one group discusses a contested issue with a reasonable-parrot multi-agent system, a control group with a standard answer-giving LLM, and a third group with no AI assistance, measuring critical-thinking disposition and argument quality before and after; if the reasonable-parrot group shows no greater gain than the control group, the central claim is falsified.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Source of the three principles—relevance, responsibility, and freedom—that characterize reasonable parrots."}],"review_version":1}