{"id":"d56466c2-c1aa-40a6-978d-961385f5c1d0","arxiv_id":"2506.09221","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Non-expert crowd workers can judge online information truthfulness about as well as experts, and an interpretable model can automate fact-checking with explanations.","lead":"This PhD thesis compiles 13 published studies on using crowdsourced non-expert judgments to assess online misinformation. It reports that crowd workers' truthfulness judgments often match expert fact-checkers, and introduces E-BART, a model that predicts truthfulness and generates explanations.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim that non-experts reliably assess truthfulness rests on treating expert fact-checker labels as unbiased ground truth; the thesis itself notes expert fact-checking is human judgment susceptible to cognitive biases, so crowd–expert agreement may only show alignment with a biased…","rationale":"The reader's weakest assumption identifies exactly the concern I consider most load-bearing: expert fact-checker labels are treated as ground truth, and all agreement analyses inherit that choice. The thesis itself provides the strongest evidence for this concern in Section 1.2, where it notes that the fact-checking organizations considered rely exclusively on human judgment and are therefore potentially susceptible to cognitive biases. That admission means the measured crowd–expert agreement is evidence about alignment with a particular human-produced standard, not direct evidence about objective truthfulness. The central claim in Section 1.10.1 is worded more strongly than the experiments can support, because 'reliably assess online (mis)information' requires a valid gold standard, while the experiments only demonstrate agreement with one or two expert-labeling processes. I do not see an internal inconsistency in the statistical work; the issue is external validity of the gold standard. The thesis deserves credit for releasing the data and for building on thirteen peer-reviewed articles, so the concern does not warrant rejection. It does, however, justify keeping the reader's CONDITIONAL verdict: the central claim should be accepted only if the ground-truth assumption is defended or the claim is rephrased as alignment with expert fact-checkers. Since the reader already made this condition explicit, my read does not change the verdict.","tokens_in":55375,"tokens_out":4297,"duration_ms":49854,"concrete_test":"Re-analyze the released crowdsourced judgments (OSF DOI 10.17605/OSF.IO/JR6VC) using only statements for which at least two independent fact-checking organizations (e.g., PolitiFact, FactCheck.org, Snopes) agree on a normalized truth label as gold, and recompute the reported agreement metrics (CEM_ORD, MAE, correlations) for Chapters 4, 5, and 7. If crowd–multi-source-gold agreement is substantially lower than crowd–single-source agreement, the ground-truth assumption is doing real work and the central claim should be weakened to 'crowds align with individual fact-checking organizations' rather than 'crowds reliably assess truthfulness.' If agreement is unchanged, the concern is mitigated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The thesis's central claim—'non-expert judges can reliably assess online (mis)information' (Section 1.10.1)—is operationalized throughout as agreement with expert labels from PolitiFact and RMIT ABC Fact Check (e.g., RQ1, RQ5, RQ16). This operationalization only supports the claim if those labels are a valid gold standard for truthfulness. The thesis does not establish that. On the contrary, Section 1.2 states that all three fact-checking organizations 'rely exclusively on human judgment for their evaluations' and are 'potentially susceptible to the systematic errors due to the limits of human cognition, known as cognitive biases.' Expert fact-checking is itself a human process involving selection, evidence retrieval, and consensus, any stage of which can embed systematic error. If expert labels are biased (e.g., by partisanship, sampling of check-worthy claims, or reliance on the same public evidence the crowd uses), then high crowd–expert agreement is evidence of alignment with that biased standard, not of reliable truthfulness assessment. The conclusion in Section 1.10.1 that crowds 'can objectively assess and categorize' information therefore goes beyond what the experiments can show. This is the load-bearing assumption: all main results (Chapters 4, 5, 7, 9, 10) compare crowd judgments to expert labels, so if the gold standard is invalid, the central claim loses its warrant.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The thesis investigates whether non-expert crowd workers can reliably assess the truthfulness of online information. It is organized around three meta-research questions: crowdsourced truthfulness assessment (including judgment scales, longitudinal designs, and multidimensional judgments), cognitive biases in fact-checking, and automated joint prediction and explanation of truthfulness. The empirical core consists of large-scale crowdsourcing experiments on political and COVID-19 statements, with expert labels from PolitiFact and RMIT ABC Fact Check used as ground truth, plus a neural model called E-BART evaluated on FEVER, e-FEVER, and e-SNLI. The central claim is that non-expert judgments often align with expert assessments and that multidimensional crowd judgments are reliable across truthfulness dimensions.","tokens_in":55649,"tokens_out":5296,"duration_ms":60342,"significance":"If the central claim holds, the thesis offers several useful contributions: large released datasets of crowd truthfulness judgments, systematic evidence on judgment scales and longitudinal participation, a PRISMA-based review of cognitive biases relevant to fact-checking, an open-source crowdsourcing framework (Crowd_Frame), and an explainable model that jointly predicts truthfulness and generates explanations. Strengths include the scale of the data collection, the use of statistical testing, the grounding in peer-reviewed publications, and the public release of datasets and software. The main caveat is that the inferential target throughout is agreement with expert fact-checker labels, so the strength of the conclusions depends on what those labels are taken to represent.","major_comments":[{"comment":"","section":"§1.2, §1.10.1, Chapters 4–5, 7, 9"},{"comment":"","section":"§1.10.1, Chapters 4–5, 7"},{"comment":"","section":"§10.4.2 (RQ30), §1.10.2"}],"minor_comments":[{"comment":"","section":"§1.9, publication list"},{"comment":"","section":"§4.2.2"},{"comment":"","section":"§4.2.1, §4.4.1"},{"comment":"","section":"§5.4.1, Figures 5.1–5.4"}],"recommendation":"major_revision","confidential_remarks":"This is a dissertation-style compilation of thirteen published papers, which may be appropriate for a repository/archival venue but is unusual for a journal submission. The main risk is overclaiming objectivity from expert-label agreement; a revision that scopes the claims and adds a validation or sensitivity analysis for the gold standard would make the central contribution defensible. The published nature of most chapters should be disclosed clearly to the editor, along with the specific new synthesis provided by the thesis."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear [Colleague],\n\nThe thing to know: this is a PhD thesis, not a new research paper. It compiles 13 peer-reviewed articles by the candidate and collaborators into one coherent narrative. If you expect novel experimental results, you won't find them; if you want a thorough, well-organized synthesis of crowdsourced fact-checking work, this is it. The data release (OSF) and the detailed appendices (Crowd_Frame infrastructure) are genuinely useful assets.\n\nWhat the thesis does well: it lays out three research lines — judgment scales, cognitive biases, automated fact-checking — and ties them together through a set of meta-questions. The experiments are large-scale (thousands of judgments, multiple platforms, longitudinal designs), and the statistical reporting is careful. The E-BART model for joint prediction and explanation is a sensible contribution, and the evaluation includes human-centered checks. The treatment of cognitive biases, with 39 biases and 11 countermeasures, is a reasonable synthesis even if the categories borrow from prior work.\n\nThe soft spots are real but not fatal. First, the central claim — that non-experts 'can objectively assess and categorize' truthfulness — is too strong. The thesis itself acknowledges in Section 1.2 that expert fact-checkers are human and susceptible to cognitive biases. If expert labels are the gold standard, then demonstrating crowd-expert agreement shows alignment with that standard, not objectivity in any absolute sense. The author could have hedged to 'can reliably approximate expert judgments' and the thesis would be on firmer ground. Second, because each chapter reproduces published findings, the novel contribution is the integration, not the evidence. That's fine for a thesis, but the arXiv version should be explicit that it's a compilation. Third, the stress-test's concern about expert bias is not addressed head-on; the limitations section discusses some issues, but the ground-truth assumption is left undefended.\n\nWho should read it: people entering human-computation or fact-checking research, and practitioners who want an overview of the crowdsourcing landscape. It's a good reference for experimental design and known pitfalls.\n\nRecommendation: I would send it to a serious referee, but conditional. The referee should push on the ground-truth assumption and on the 'objectivity' framing. The empirical core is solid and the data release is valuable. So: engage with it, but ask for a revised framing.","headline":"A solid, well-organized thesis compiling 13 peer-reviewed papers; the empirical work is careful, but the 'objectivity' claim overreaches relative to the expert-label gold standard.","tokens_in":56139,"tokens_out":2119,"would_cite":false,"duration_ms":22170,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Non-expert crowd workers can reliably assess online truthfulness, this thesis shows.","keywords":["misinformation","crowdsourcing","fact-checking","truthfulness assessment","cognitive biases","longitudinal studies","explainable AI","E-BART"],"falsifier":"Gather statements that two professional fact-checking organizations rate differently, or cases where later evidence shows the original expert rating was wrong, and run the same crowdsourcing protocol. If aggregated crowd judgments follow the expert labels rather than the independently verifiable facts, the core claim fails; crowd judgments that track documented facts where experts erred would strengthen it. A simpler check is temporal: the thesis predicts judgment quality falls as statements age, so a longitudinal dataset showing no such decline would count against the timing claim.","tokens_in":55164,"feed_emoji":"🧠","tokens_out":8945,"duration_ms":83508,"temperature":0.7,"pith_summary":"This thesis sets out to show that non-expert crowd workers can assess the truthfulness of online statements well enough to support fact-checking at scale. It compares thousands of crowd judgments against expert labels from two public fact-checking archives, across three judgment scales, a longitudinal study on COVID-19 misinformation, and a seven-dimension truthfulness instrument. The thesis reports that aggregated non-expert judgments often align with expert assessments, that judgment quality improves with worker experience and with the recency of the statement, and that a multidimensional view of truthfulness captures distinct facets that help interpret judgments. It also identifies cognitive biases that distort fact-checking and proposes countermeasures, and it introduces E-BART, a neural model that jointly predicts truthfulness and generates human-readable explanations.","feed_headline":"Crowd judgments align with expert fact-checkers, thesis shows","feed_subtitle":"Non-expert judgments match experts, and one model explains its own verdicts.","key_machinery":"The argument is carried by a pipeline of large-scale crowdsourcing experiments in which non-experts judge statements whose truthfulness has already been labeled by professional fact-checkers, so that agreement can be measured. Three judgment scales are used—a three-level scale, a six-level scale, and a 0-to-100 slider scale—and judgments are aggregated by mean or median before comparison. The longitudinal design repeats the task in multiple batches to isolate the effects of timing and experience. The multidimensional instrument asks workers to rate each statement on seven truthfulness dimensions. For bias, the machinery is a systematic literature review following standard reporting guidelines, plus controlled hypothesis-testing experiments with worker-trait questionnaires. For automation, the central object is the E-BART architecture, which places a Joint Prediction Head on the BART language model so that truthfulness classification and explanation generation share one training objective, with model calibration performed by temperature scaling.","core_discovery":"On its own terms, the thesis establishes that the crowd can serve as a reliable source of truthfulness judgments. The central discovery is a converging body of evidence from controlled crowdsourcing experiments: agreement between aggregated non-expert judgments and expert fact-checker labels is high enough to make crowd-based fact-checking viable, especially when judgments are aggregated, when statements are assessed soon after publication, and when workers have prior experience. A second discovery is that truthfulness is not one-dimensional: seven dimensions—Correctness, Neutrality, Comprehensibility, Precision, Completeness, Speaker's Trustworthiness, and Informativeness—are largely non-redundant and jointly more informative than a single overall score. A third is that cognitive biases are systematic and measurable in crowd judgments, with 39 biases identified as relevant to fact-checking and specific worker traits associated with biased behavior. Finally, the E-BART model shows that a single architecture can jointly predict a truthfulness label and produce a coherent explanation, and that these machine-generated explanations help human judges detect misinformation.","pith_inferences":["Inference: if the agreement results generalize beyond political and COVID-19 claims, the same experimental pipeline could serve as real-time triage in which crowd labels prioritize claims for expert fact-checking; the thesis does not itself test that deployment.","Inference: because biases act differently on different truthfulness dimensions, debiasing interventions could be targeted per dimension and evaluated by tracking dimension-level error; this is a natural but untested extension.","Inference: the paradoxical link between reported belief in science and lower accuracy suggests domain confidence may be a moderating factor, so testing the protocol across scientific and health claims with varying controversy levels would clarify the scope.","Inference: the thesis's core claim is relative to expert-defined truthfulness; validating crowd judgment against independently verified facts would be a stronger test the thesis does not perform."],"forward_implications":["A crowd-based first-pass fact-checking layer becomes practical: non-expert judgments can flag suspicious claims for professional review, cutting the volume experts must inspect.","Longitudinal studies with repeated participation from the same workers are not only feasible but beneficial, since returning workers produce higher-quality judgments.","Truthfulness should be measured and modeled as multidimensional rather than as a single score, because the seven dimensions carry distinct information.","Automated fact-checking systems that generate explanations as part of the prediction, as E-BART does, can improve human skepticism and detection of misinformation.","Cognitive-bias countermeasures, selected from the 39 identified biases, can be embedded in task design to reduce systematic errors in crowd work."],"supporting_citations":[{"why":"Reports the three-scale crowdsourcing experiment that supplies the core comparison between non-expert and expert judgments.","marker":"[437]"},{"why":"Provides the longitudinal COVID-19 study showing timing and worker experience affect judgment quality.","marker":"[441]"},{"why":"Introduces the seven-dimension truthfulness instrument and shows the dimensions are non-redundant.","marker":"[480]"},{"why":"Tests hypotheses linking worker traits to biased judgment in fact-checking tasks.","marker":"[140]"},{"why":"Presents the joint prediction-and-explanation model evaluated on fact-checking and natural-language-inference datasets.","marker":"[57]"},{"why":"Systematic review that identifies the 39 fact-checking-relevant cognitive biases and countermeasures.","marker":"[483]"},{"why":"Survey of barriers to longitudinal studies that grounds the recommendations for sustained participation.","marker":"[482]"},{"why":"Earlier result, extended here, that aggregated crowdsourced truthfulness judgments agree with expert labels on a larger sample.","marker":"[284]"}],"fun_headline_variants":["Crowd fact-check scores match expert verdicts, study finds","Non-experts align with experts in truthfulness ratings","One model predicts truth labels and explains its own reasoning","Seven dimensions define truthfulness better than one score"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the labels produced by professional fact-checking organizations—the six-level U.S. archive and the three-level Australian archive used throughout—are the correct ground truth for truthfulness; if those labels are biased or wrong, then crowd agreement with them does not establish that non-experts can judge truthfulness.","fun_headline_variants_meta":{"raw":{"variants":["Crowd fact-check scores match expert verdicts, study finds","Non-experts align with experts in truthfulness ratings","One model predicts truth labels and explains its own reasoning","Seven dimensions define truthfulness better than one score"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000278,"raw_usage":{"total_tokens":1639,"prompt_tokens":913,"completion_tokens":726,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":529,"completion_tokens_details":{"reasoning_tokens":661}},"tokens_in":529,"tokens_out":726,"duration_ms":8017,"temperature":1.0,"reasoning_tokens":661,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T04:53:18.007755+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Gather statements that two professional fact-checking organizations rate differently, or cases where later evidence shows the original expert rating was wrong, and run the same crowdsourcing protocol. If aggregated crowd judgments follow the expert labels rather than the independently verifiable facts, the core claim fails; crowd judgments that track documented facts where experts erred would strengthen it. A simpler check is temporal: the thesis predicts judgment quality falls as statements age, so a longitudinal dataset showing no such decline would count against the timing claim.","supporting_citations":[],"review_version":1}