{"id":"120d961a-a618-447f-a406-589fde9e3ca7","arxiv_id":"2411.09969","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A slider-controlled AI tool that personalizes science articles for individual readers improved self-reported understanding in a small study, with preferences varying by motivation.","lead":"The authors built TranSlider, a tool that lets readers adjust a slider to control how personally an AI rewrites science articles using their hobbies, location, and background. In a 15-person study, readers who saw multiple personalized versions reported understanding the science better, though they disagreed on how much personalization they wanted.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No fact-check of the 268 study translations: the 3-scientist pilot used authors' own papers, so hallucination risk in the actual study materials is unmeasured and the comprehension benefit claim is not yet supported.","rationale":"Good-faith read: the paper is an honest exploratory study with a clear interaction design, rich qualitative data, and appropriately cautious conclusions. The strongest claim about understanding is, however, conditional on output reliability. The reader's chosen weakest assumption—that GPT-4o translations are accurate—is the most load-bearing because accuracy is a prerequisite for the entire science-communication value proposition, and the evidence for it is thin: three scientists reviewed only their own papers, with no check on the two articles actually shown to the 15 participants. Other concerns (no baseline, self-reported understanding, small N) affect effect size and generalizability, but they would not invalidate the direction of the claim if the content were accurate. In contrast, undetected hallucinations would make the tool actively harmful, not merely under-proven. The proposed test directly addresses this gap and could be run post hoc on the existing logs, since all 268 translations were recorded. The paper's own discussion (Section 7.2.2) proposes author verification of a few translations as a mitigation, acknowledging the gap. Hence the reader's CONDITIONAL verdict stands, and the accuracy check is a reasonable condition for acceptance. The critique is about missing evidence, not about author conduct.","tokens_in":28581,"tokens_out":4643,"duration_ms":48590,"concrete_test":"Recruit the corresponding authors of the two source articles (Cell 187:2269 and Nat. Commun. 15:5548), or two independent domain experts with relevant expertise, to review a stratified sample of the 268 logged translations (e.g., 25 per article spanning slider degrees 0/25/50/75/100). Ask them to flag any statement that is factually wrong, overstates, or implies a causal claim not in the abstract, and to rate whether each error would mislead a lay reader. Report the error rate per article and per degree. If any sampled translation contains a misleading factual error, the 'compounding understanding' finding becomes unreliable. If none are found, the accuracy concern is retired.","verdict_should_be":"UNCHANGED","load_bearing_attack":"TranSlider's central value—helping non-experts understand scientific articles via personalized translations—depends on the translations being factually correct. The only accuracy check (Section 4.3) was a pilot in which three scientists reviewed AI translations of their own publications. They detected no hallucinations, but this does not cover the two articles used in the main study (health [55] and environment [38], Section 5.1). The main study generated 268 translations (Section 6.1.1), yet none were verified against the source abstracts by domain experts or the original authors. If any of those translations contain subtle factual errors, participants' reported 'compounding understanding' could instead be a compounding of confidently held misconceptions, and the tool would be unsafe for science communication. This is not merely hypothetical: P2's quote in Section 6.3.1 shows a participant who initially believed 'the sugar was BRCA2' before later translations corrected it, illustrating that translations can mislead before being corrected. The pilot's scope, described in Section 4.3.2, is authors verifying 'their recently published work'; the study articles are different, so the grounding is absent. Without a systematic fact-check of the actual materials, the central claim that the tool helps general audiences understand scientific articles is not established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents TranSlider, an LLM-powered reading interface that generates personalized translations of scientific abstracts for general audiences, with a slider (0–100) that lets users steer the degree of personalization grounded in a user profile (background, hobbies, location, food). The authors report an exploratory user study with 15 non-expert participants who read two scientific articles (health and environment) and explored multiple translations, followed by post-session interviews and log analysis. The main findings are that participants differed in their preferred degree of personalization (extrinsic vs. intrinsic motivation correlated with lower vs. higher degrees), that reading multiple translations produced what participants described as a 'compounding understanding' of the content, and that the slider offered a simple and enjoyable but sometimes overwhelming form of control. The paper also contributes design implications for steerable human-AI alignment and for science communication, as well as a discussion of trust and verification concerns.","tokens_in":28794,"tokens_out":2798,"duration_ms":32042,"significance":"The work addresses a timely and important problem: making scientific text accessible to heterogeneous general audiences at scale while preserving user agency over AI-generated personalization. Its strengths include a clearly described tool and prompt template (Figure 3), a transparently reported exploratory study with participant quotes and behavioral logs, and a limitations section that names the small sample and lack of diversity. The idea of presenting multiple personalized translations as complementary 'puzzle pieces' is a useful framing for CSCW/HCI research on science communication and human-AI alignment. If the claims are appropriately qualified, the paper can inform future designs for scalable, user-controllable science communication tools.","major_comments":[{"comment":"The accuracy of the AI-generated translations is load-bearing for the paper's central value proposition, but the only verification was a pilot in which three scientists reviewed translations of their own papers; the 268 translations actually used in the main study (from the two articles in §5.1) were never fact-checked against the source abstracts. Given that the tool is framed as a trustworthy science-communication intermediary (Section 3.2) and that the P2 quote in §6.3.1 shows a participant initially forming a wrong mental model ('the sugar was BRCA2'), the comprehension benefit claims are not fully supported unless the actual study materials were free of hallucinations or factual errors. I recommend a post-hoc verification of the study translations by domain experts or the original authors, or a careful softening of the claim that the tool 'helps general audiences understand' scientific content.","section":"§4.3.2 and §6.1.1"},{"comment":"The 'compounding understanding' claim rests entirely on participants' self-reports and their written takeaways; no objective comprehension measure (e.g., a pre/post test or an expert assessment of the takeaways) was administered. RQ2.1 specifically asks about comprehension, and the study design does not distinguish between perceived understanding and accurate understanding. The finding is suggestive and valuable for an exploratory study, but the abstract and §6.3.1 should either be reframed as perceived or reported understanding or be supported with a more direct measure.","section":"§5.4 and §6.3.1"}],"minor_comments":[{"comment":"The thematic analysis was coded by the first author and then discussed with the second author and the team, but no inter-rater reliability or independent coding is reported; given the qualitative nature of the main findings, a brief note on how disagreements were resolved would strengthen the methodology.","section":"§5.4.2"},{"comment":"The Pearson correlation (r = 0.36) between personalization degree and translation length is reported without a confidence interval or p-value; even though the paper does not claim inferential statistics, adding the 95% CI would help readers gauge the strength of the observed trend.","section":"§6.1.1"},{"comment":"The 'Finish' button is described as a proxy for interest in the original paper, but the paper does not report whether participants actually requested the PDF; reporting these counts would make the proxy more interpretable.","section":"§4.1.4"},{"comment":"The categorization into 'lower degree' and 'higher degree' preference groups (0–45 vs. 51–100) is introduced without a stated cutoff rule; please clarify how this threshold was determined.","section":"§6.2.2"},{"comment":"The pilot study's sentence 'All three scientists found the translations surprisingly understandable' is vivid but informal; consider rewording for a formal report.","section":"§4.3.2"}],"recommendation":"major_revision","confidential_remarks":"This is a solid exploratory CSCW paper with a well-described tool and an honest limitations section. The main revision needed is to address the accuracy-verification gap and to temper the comprehension claim, which are both load-bearing for the central contribution. If the authors can add a post-hoc fact-check of the actual study articles' translations (or explicitly reframe the contributions as an exploration of perceived understanding and user experience), the paper would likely be publishable. I do not see a circularity problem or any issue with the paper's internal consistency."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know about this paper: it's a genuinely new interaction idea—a slider that steers the degree of LLM personalization for science text—and the user study is honestly conducted, but the evidence for the 'compounding understanding' claim is weaker than the abstract suggests.\n\nWhat's new: TranSlider combines a 0-100 personalization slider with editable user profiles and a translation history, so readers can compare multiple AI-generated analogies. Prior work on templated personalization (Persalog, spatial analogies) or LLM simplification didn't offer this direct control over personalization level. The authors also contribute a descriptive finding that users with extrinsic motivation prefer lower degrees, intrinsic higher.\n\nWhat it does well: the study is careful. They counterbalanced article order, recruited non-experts across fields, logged 268 translations, and the quotes support the themes. The limitations section is candid: small N, homogeneous sample, no robust stats. They don't claim more than exploratory.\n\nSoft spots, in proportion: first, the accuracy pilot is thin. Three scientists checked translations of their own papers, not the two articles used in the main study. So we have no systematic check on the 268 translations participants actually read. The P2 quote—initially thinking 'the sugar was BRCA2'—shows a translation produced a misconception that later translations corrected. That's actually evidence for the compounding benefit, but it also underscores that a single translation can mislead. If the tool is meant for science communication, hallucination risk needs measurement. Second, there is no baseline condition (e.g., original abstract or generic simplification), so we can't say whether the slider and multiple translations beat a simpler alternative. Third, comprehension is self-reported; no objective test. These are real gaps but not fatal for an exploratory study—the authors frame their claims as perceived utility.\n\nSomething the stress-test note overweights: the 'unsafe' framing. The tool is a research probe, not a deployed product, and the paper explicitly discusses reliability concerns and suggests author verification. The lack of fact-checking is a quality issue, not evidence of actual misinformation.\n\nWho this is for: HCI/CSCW folks working on human-AI alignment, personalization, or science communication. It deserves a serious referee—the idea is novel, the study is usable, and the limitations are addressable in future work.\n\nRecommendation: send it out. With a request for a baseline condition and a fact-check of study materials, it would be much stronger. As is, accept if the venue values exploratory qualitative contributions, with revisions.","headline":"A novel slider-based interaction for steering LLM personalization of science text, with an honest but small user study; the compounding-understanding claim is intriguing but rests on self-report and unverified translations.","tokens_in":29306,"tokens_out":1863,"would_cite":true,"duration_ms":21613,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A slider-based AI translation tool can help non-expert readers understand scientific articles, and reading multiple personalized translations produces a compounding understanding.","keywords":["science communication","personalized translation","large language models","analogy generation","human-AI alignment","interactive slider","science literacy","user study"],"falsifier":"Have domain experts fact-check each of the 268 logged translations against the two source articles for unsupported mechanisms, numbers, or causal claims; any confirmed hallucination would break the paper's accuracy assumption. Separately, a randomized comprehension study, one group reading several translations of the BRCA2 article and another reading a single translation, with a factual quiz that asks whether the sugar molecule is the BRCA2 protein, would directly test whether the compounding effect is real or just an artifact of more exposure.","tokens_in":28375,"feed_emoji":"🎚️","tokens_out":9315,"duration_ms":91428,"temperature":0.7,"pith_summary":"TranSlider is an interactive reading tool that generates analogy-based translations of scientific articles from a reader's profile (hobbies, location, education, favorite food), with a slider that sets the personalization strength from 0 to 100. The paper argues that this makes science communication scalable to diverse audiences, because readers can choose, or explore, the level of contextualization that works for them. In a 15-person study, all participants reported benefits, but they split: some preferred strongly personalized, detailed translations, while others wanted concise, lightly contextualized ones. The central finding is that reading several translations of the same article produced a compounding understanding, as readers assembled a complete picture from partial pieces and corrected initial misunderstandings. If this holds, one article could serve many readers at once through steered translations rather than a single plain-language summary.","feed_headline":"Reading several AI translations compounds science understanding","feed_subtitle":"15 readers built fuller understanding from multiple personalized translations, each tuned to their own context.","key_machinery":"The central mechanism is TranSlider itself, an interactive reading interface with three load-bearing parts: a 0-100 personalization slider, an editable user profile (background, age, hobbies, location, favorite food), and a history box that stacks past translations for comparison. The other central object is the prompt template: it combines the profile, the current slider value, a three-point personalization spectrum (degrees 0, 50, and 100), in-context examples of generic versus personalized analogies, and a chain-of-thought instruction that asks the model to reason about how the slider value should change the output. The prompt directs the model to produce two paragraphs, an analogy-based explanation of the research and a statement of its implications for the reader, so the slider value scales how much the translation leans on specific personal details rather than inferred high-level context. This design is what makes personalization steerable without free-form prompting.","core_discovery":"On the paper's own terms, the discovery is that AI-personalized translation is not a one-size-fits-all output but a spectrum, and that the spectrum itself is useful. TranSlider generates, for each article, many translations across the 0-100 personalization slider; readers used them to build understanding cumulatively. Participants generated 268 translations in total, and their favorite translations averaged 53.14 on the personalization scale with a wide standard deviation (33.24), showing that people diverged in what they wanted. Descriptively, extrinsically motivated readers favored lower personalization (mean 40.52) while intrinsically motivated readers favored higher personalization (mean 62.75). The paper also reports that reading multiple translations provided a compounding understanding, with participants describing assembling a complete picture from partial pieces and one reader correcting the misunderstanding that the sugar molecule in the health article was the BRCA2 protein. At the same time, participants worried about the reliability of AI-generated translations and about readers over-relying on them instead of the original articles.","pith_inferences":["My inference beyond the paper: because the accuracy pilot checked only scientists' own papers, the first test of TranSlider's promise is a systematic fact-check of the two study articles' translations.","My inference: the compounding effect is probably not specific to science communication; deliberately showing learners multiple paraphrases or analogies of the same concept may help them repair misunderstandings in other learning contexts.","My inference: the paper's observed positive correlation between personalization degree and translation length ($r = 0.36$) suggests part of what readers experience as 'personalization' may be elaboration; a length-controlled comparison could separate the two.","My inference: the slider could become a feedback mechanism instead of a one-shot control, letting readers mark the analogy that worked and feeding that signal back into the prompt to make the system learn from the reader without collecting more data."],"forward_implications":["Science blogs and news sites could offer readers a range of steered translations instead of a single plain-language summary, letting each reader choose the personalization level they want.","The compounding-understanding result implies that reading tools should keep a history of generated translations and encourage comparison, because multiple versions helped readers catch and fix misunderstandings.","The descriptive split between extrinsic and intrinsic readers suggests that personalization systems could set different default degrees depending on why someone is reading, not just on their profile.","The controlled user-initiative design offers a template for human-AI alignment that reduces data collection: let users steer the output rather than require the model to infer personalization preferences."],"supporting_citations":[{"why":"Establishes analogies as an effective science-communication strategy that TranSlider's translations are built on.","marker":"[9]"},{"why":"Shows the benefits and pitfalls of generating plain-language summaries for specific audiences, motivating user-specific translations.","marker":"[10]"},{"why":"Supplies the thematic-analysis method used to derive themes from the study's interview transcripts.","marker":"[17]"},{"why":"Demonstrates LLMs' capacity to generate personalized analogies, the capability TranSlider's prompt exploits.","marker":"[30]"},{"why":"Provides prior template-based concrete re-expression work that the paper extends with free-form LLM generation.","marker":"[44]"},{"why":"Introduces AI-bridged scalable personalization and authors' attitudes toward it, framing the design rationale and the misinformation caution.","marker":"[50]"},{"why":"The health article used as study material; the compounding-understanding finding includes readers' misconceptions about it.","marker":"[55]"},{"why":"The environment article used as study material, chosen to test whether personalization also helps on less personally relevant topics.","marker":"[38]"},{"why":"Chain-of-thought prompting, the technique the TranSlider prompt template adapts to steer the personalization degree.","marker":"[91]"}],"fun_headline_variants":["Slider-tuned AI translations build compound science understanding","Personalized AI translations: spectrum beats one-size-fits-all","Multiple AI translations deepen science comprehension","AI translation slider helps readers assemble full picture","Varied AI translations compound readers' science insights"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that GPT-4o's personalized translations are scientifically accurate and free of misinformation; the pilot verified this only by having three scientists review their own papers, not by checking the two articles actually used in the user study.","fun_headline_variants_meta":{"raw":{"variants":["Slider-tuned AI translations build compound science understanding","Personalized AI translations: spectrum beats one-size-fits-all","Multiple AI translations deepen science comprehension","AI translation slider helps readers assemble full picture","Varied AI translations compound readers' science insights"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000162,"raw_usage":{"total_tokens":1241,"prompt_tokens":949,"completion_tokens":292,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":565,"completion_tokens_details":{"reasoning_tokens":222}},"tokens_in":565,"tokens_out":292,"duration_ms":3672,"temperature":1.0,"reasoning_tokens":222,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T20:05:25.694759+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have domain experts fact-check each of the 268 logged translations against the two source articles for unsupported mechanisms, numbers, or causal claims; any confirmed hallucination would break the paper's accuracy assumption. Separately, a randomized comprehension study, one group reading several translations of the BRCA2 article and another reading a single translation, with a factual quiz that asks whether the sugar molecule is the BRCA2 protein, would directly test whether the compounding effect is real or just an artifact of more exposure.","supporting_citations":[{"cited_title":"Guelfo, P","cited_arxiv_id":null,"evidence_quote":"The environment article used as study material, chosen to test whether personalization also helps on less personally relevant topics."}],"review_version":1}