{"id":"6e8298c5-b5bb-49fd-bbe1-58f6db9b10a1","arxiv_id":"2606.13187","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"The authors release BioStance, a dataset of 39,600 context-preserving Reddit post-comment pairs annotated for favor, against, or none stance on six bioethical controversy targets, with Krippendorff's alpha of 0.82.","lead":"This paper creates BioStance, a dataset with 39,600 Reddit post and comment pairs labeled for stance on bioethical topics, maintaining conversation context and showing good annotator agreement. General readers might find it relevant for understanding how data resources support AI in analyzing complex social debates on ethics and science.","discovery_kind":"unclear","skeptic_critique":{"model":"grok-4.3","headline":"No significant objection identified","rationale":"Reader correctly notes that full methods, sampling, and release details cannot be assessed from abstract alone. With the placeholder indicating full text access, the presented claim still holds as a standard dataset contribution; no load-bearing flaw in the argument itself is apparent.","tokens_in":1674,"tokens_out":192,"duration_ms":15873,"concrete_test":"Confirm public release of the dataset (or code to reconstruct it) and verify that a random sample of 100 instances matches the described three-class labels and context structure.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is the release and description of a new annotated dataset with reported properties (size, annotation reliability, structure). The abstract and described construction are internally consistent for a dataset paper; no internal contradiction, unsupported calculation, or missing prerequisite for the stated properties is visible from the provided material.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript presents BioStance, a dataset of 39,600 annotated Post-Comment pairs drawn from Reddit discussions on bioethical topics. It covers six targets across three dimensions (fundamental value conflicts, individual liberty vs. collective responsibility, technological uncertainty), preserves hierarchical conversational context in each instance, and reports annotations by three independent annotators using a three-class scheme (Favor, Against, None) with mean Krippendorff's α = 0.82.","tokens_in":1691,"tokens_out":475,"duration_ms":21387,"significance":"If the construction details hold, the dataset would fill a documented gap in domain-specific, context-rich resources for stance detection and argument mining in bioethics on social media. The reported inter-annotator agreement and conversational structure are strengths that could support modeling of context-dependent discourse.","major_comments":[{"comment":"Dataset Construction section: The manuscript provides the final size (39,600 pairs) and thematic coverage but does not detail the subreddit selection criteria, search terms, or sampling procedure used to identify bioethical discussions. This information is required to evaluate selection bias and the claim that the targets adequately represent bioethical controversies on Reddit.","section":"Dataset Construction"},{"comment":"Annotation section: While the mean Krippendorff's α of 0.82 is reported, the paper does not describe the annotation guidelines given to annotators, the label distribution across the three dimensions, or any adjudication process. These details are load-bearing for the claim of 'high-quality human annotation' and the dataset's utility for downstream modeling.","section":"Annotation"}],"minor_comments":[{"comment":"Abstract: The three dimensions are named but not illustrated with example targets; adding one concrete example per dimension would improve immediate readability.","section":"Abstract"},{"comment":"Related Work: The positioning against existing stance datasets would be strengthened by a brief quantitative comparison (e.g., size, domain, context preservation) rather than only qualitative statements.","section":"Related Work"}],"recommendation":"major_revision","confidential_remarks":"This is a dataset-release paper; confirm whether the target journal routinely publishes such contributions or prefers them in a data-track format."},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their constructive feedback and positive assessment of BioStance's potential contribution. We address each major comment below and will revise the manuscript to improve transparency and reproducibility.","responses":[{"response":"We agree that these procedural details are necessary for assessing selection bias and representativeness. In the revised manuscript we will expand the Dataset Construction section with the specific subreddits chosen, the exact search terms and filters applied per target, and the sampling procedure (including how threads were filtered for sufficient context and how the final 39,600 pairs were obtained).","revision_made":"yes","referee_comment":"[Dataset Construction] Dataset Construction section: The manuscript provides the final size (39,600 pairs) and thematic coverage but does not detail the subreddit selection criteria, search terms, or sampling procedure used to identify bioethical discussions. This information is required to evaluate selection bias and the claim that the targets adequately represent bioethical controversies on Reddit."},{"response":"We concur that these elements are essential. The revised version will include the full annotation guidelines (as an appendix), a breakdown of label distributions (Favor/Against/None) by dimension and target, and a description of the adjudication procedure used when the three annotators disagreed.","revision_made":"yes","referee_comment":"[Annotation] Annotation section: While the mean Krippendorff's α of 0.82 is reported, the paper does not describe the annotation guidelines given to annotators, the label distribution across the three dimensions, or any adjudication process. These details are load-bearing for the claim of 'high-quality human annotation' and the dataset's utility for downstream modeling."}],"tokens_in":1290,"tokens_out":370,"duration_ms":14377,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"This paper's main point is the release of BioStance: 39,600 Reddit post-comment pairs annotated for stance on six bioethical targets across three dimensions, with conversational context kept intact and a mean Krippendorff's alpha of 0.82.\n\nWhat it does well is target a narrow but real gap. Most stance datasets stay broad or political; this one narrows to value conflicts, liberty versus collective responsibility, and technological uncertainty, using actual Reddit threads. The three-annotator setup and reported agreement give it a usable foundation for people who want context-aware models.\n\nThe soft spots are the usual ones for a dataset paper and stay minor. The abstract gives size and alpha but skips sourcing details, subreddit selection, and exact guidelines. Without those in the full text, it's hard to judge selection effects or how well the labels match the claimed dimensions. Reddit data always carries platform bias, and the paper does not test whether the preserved context actually helps downstream models. That is fine for a data release but caps immediate impact.\n\nThis is for NLP researchers who need specialized stance resources or computational social scientists working on bioethics. If the full paper shows transparent collection, public data, and clear guidelines, it is worth a serious referee. I would send it to review rather than desk reject.","headline":"BioStance is a clean new dataset release for stance detection on bioethics with good annotation numbers, but its real value hinges on the full methods and public release.","tokens_in":2191,"tokens_out":344,"would_cite":false,"duration_ms":12056,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"BioStance supplies 39,600 context-preserved Reddit pairs labeled for stance in bioethical controversies.","keywords":["stance detection","bioethical controversies","Reddit dataset","context-aware annotation","social media discourse","argument mining","Krippendorff's alpha"],"falsifier":"Re-annotating a random sample of the pairs with new annotators and obtaining agreement below 0.67 on Krippendorff's alpha would undermine the reliability of the resource.","tokens_in":2553,"feed_emoji":"📊","tokens_out":627,"duration_ms":15570,"temperature":0.7,"pith_summary":"The paper introduces BioStance, a dataset of 39,600 Post-Comment pairs from Reddit discussions on bioethical topics. Each pair includes hierarchical context and receives labels of Favor, Against, or None from three annotators. The annotations show substantial agreement with a mean Krippendorff's alpha of 0.82. This resource targets six controversial issues across three dimensions of bioethics to aid modeling of context-dependent stance in social media discourse.","feed_headline":"BioStance supplies 39,600 labeled Reddit pairs for bioethics stance","feed_subtitle":"Context from full threads plus three-class labels across value, liberty and uncertainty dimensions achieve 0.82 agreement.","key_machinery":"BioStance dataset, which supplies hierarchical conversational context for each post-comment pair to enable context-aware stance detection.","core_discovery":"We present BioStance, a context-aware dataset of 39,600 annotated Post-Comment pairs from Reddit bioethical discussions. BioStance covers six controversial targets across three dimensions of bioethical controversy: fundamental value conflicts, individual liberty versus collective responsibility, and technological uncertainty. Each instance preserves hierarchical conversational context and is labeled by three independent annotators using a three-class stance scheme: Favor, Against, and None. The annotations achieve a mean Krippendorff's α of 0.82, indicating substantial reliability.","pith_inferences":["Similar context-preserving collection methods could be applied to stance datasets in other polarized domains such as climate policy or economic regulation.","The annotation reliability suggests the three-class scheme may transfer to related tasks like detecting neutrality in ethical debates.","Downstream models trained on BioStance might reveal patterns in how context shifts stance that single-post datasets miss."],"forward_implications":["The dataset enables training and evaluation of context-aware stance detection models on bioethical topics.","It supports argument mining research by providing structured conversational threads.","It facilitates computational analysis of how bioethical discourse unfolds on social media platforms.","The three dimensions and six targets allow systematic comparison across different types of controversy."],"fun_headline_variants":["BioStance: 39,600 Reddit pairs for stance detection in bioethics","39,600 annotated pairs in context-aware BioStance dataset from Reddit","BioStance dataset annotates 39,600 pairs for bioethics stance on Reddit","Reddit BioStance has 39,600 labeled context-aware bioethics pairs"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The hierarchical conversational context preserved in each instance meaningfully supports context-aware stance detection modeling, and the chosen targets and dimensions adequately represent bioethical controversies on Reddit.","fun_headline_variants_meta":{"raw":{"variants":["BioStance: 39,600 Reddit pairs for stance detection in bioethics","39,600 annotated pairs in context-aware BioStance dataset from Reddit","BioStance dataset annotates 39,600 pairs for bioethics stance on Reddit","Reddit BioStance has 39,600 labeled context-aware bioethics pairs"]},"model":"grok-4.3","cost_usd":0.009891,"raw_usage":{"total_tokens":4378,"prompt_tokens":628,"num_sources_used":0,"completion_tokens":83,"cost_in_usd_ticks":98912000,"prompt_tokens_details":{"text_tokens":628,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":3667,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":628,"tokens_out":83,"duration_ms":22586,"temperature":1.0,"reasoning_tokens":3667,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-27T06:46:51.279322+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Re-annotating a random sample of the pairs with new annotators and obtaining agreement below 0.67 on Krippendorff's alpha would undermine the reliability of the resource.","supporting_citations":[],"review_version":1}