REVIEW 5 major objections 4 minor 64 references
Efficiency and Effectiveness of LLM-Based Summarization of Evidence in Crowdsourced Fact-Checking
T0 review · 5 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read LLM-generated summaries of evidence let crowdsourced fact-checkers match the accuracy of full-length webpages while completing significantly more assessments in less time.
desk verdict Useful A/B result on LLM summaries for crowd truthfulness, but the 'comparable accuracy' claim needs equivalence testing before you trust it. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying mechanism is the Summary modality: each evidence webpage is condensed by Meta-Llama-3-8B-Instruct into a concise bullet-point list, guided by a prompt (Figure 1) that instructs the model to describe the document, focus on query-related points, stay accurate, and return JSON. The argument proceeds by comparing this modality against the full-length Standard condition on the same 120 PolitiFact statements, using accuracy, MSE, MAE, task time, Krippendorff's alpha, and self-reported need and usefulness as outcome measures. Summary fidelity is assumed rather than measured, so the comparison tests the summary as a presentation format, not the quality of the summarization itself.
What would settle it
Compare worker accuracy on a set of claims for which the LLM summary demonstrably omits a decisive detail, such as a quoted 'not' or a conditional caveat; if workers given those summaries systematically misjudge those claims compared to workers given full pages, then the comparable-effectiveness result depends on this particular summary quality rather than holding for summarized evidence generally.
Extended reading notes
Core claim
The central claim is that presenting crowd workers with LLM-generated summaries of evidence, rather than full-length webpages, yields comparable truthfulness assessments while significantly improving efficiency. In their experiment, individual accuracy is 0.31 in both modalities, and aggregated mean accuracy is 0.33 vs 0.31 for Standard vs Summary; bootstrap confidence intervals for the differences in MAE and MSE include zero, so no statistically significant gap appears. Workers in the Summary modality show higher agreement (Krippendorff's alpha) at p<0.05, and they complete roughly 15% more assessments per unit time. The paper reads these results as evidence that summarization preserves the decision-relevant content of the evidence while reducing cognitive load, making it a viable substitute for full-length evidence in large-scale truthfulness annotation.
Load-bearing premise
The load-bearing premise is that the LLM summaries preserve the factual content needed to judge truthfulness; the paper asserts this design goal in the prompt but never validates it against source documents or another summarizer.
Editorial extensions
If this is right
- Crowdsourced fact-checking pipelines can swap full evidence pages for LLM summaries without losing accuracy, shrinking per-task time and cost.
- Higher internal agreement in the Summary condition implies that condensed evidence reduces interpretational variance, which should make aggregated verdicts more stable.
- Comparable over- and under-estimation error patterns across modalities indicate summaries do not systematically bias truthfulness judgments in one direction.
- Efficiency gains compound with scale: the paper estimates 6,000 assessments cost about £1,716 with full pages but £1,486 with summaries, a saving that grows with dataset size.
- Worker reliance on and perceived usefulness of evidence stay high with summaries, so the format does not cause workers to skip the evidence.
Reading between the lines
- Because summary fidelity is never validated, the comparable-accuracy result may be specific to this Llama-3 prompt and model on PolitiFact-style webpages; a summarizer that drops caveats or flips negations could break the equivalence. A testable extension: run the same A/B with an adversarial summarizer that deletes key qualifiers and check whether accuracy diverges.
- The ~15% throughput gain likely understates real-world gains, since the Summary condition also had lower initial abandonment (38% vs 49%); combining summaries with per-judgment payment could further cut effective cost per completed label.
- Higher internal agreement in the Summary condition could reflect reduced information rather than better judgment: if summaries make workers converge on the same (possibly wrong) reading, agreement rises without accuracy improving. A follow-up could compare agreement on summaries known to omit a decisive fact.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports an A/B crowdsourcing study comparing two evidence-presentation modalities for truthfulness assessment: a Standard condition showing full-length webpages and a Summary condition showing LLM-generated (Meta-Llama-3-8B-Instruct) summaries of those webpages. Using 120 PolitiFact statements and 100 Prolific workers per condition, the authors measure agreement with expert labels (accuracy, MAE, MSE), internal agreement, task completion time, and worker-reported reliance on and usefulness of evidence. The paper claims that the Summary modality achieves accuracy and error metrics comparable to the Standard modality while significantly improving efficiency and lowering cost, and that it also increases internal agreement among workers.
Significance. If the central claim were established, the paper would make a useful practical contribution: LLM-generated summaries could reduce the time and cost of large-scale crowdsourced fact-checking without degrading judgment quality. The study has genuine strengths: it uses a real crowdsourcing platform, a controlled A/B design, external expert labels from PolitiFact, and it makes the collected data publicly available. The efficiency result is supported by a significant Mann-Whitney test, and the cost simulation, once corrected, gives a plausible order-of-magnitude estimate. However, the headline effectiveness claim rests on interpreting non-significant differences as equivalence, with wide confidence intervals and no pre-specified equivalence margin. The paper also does not validate whether the summaries preserve the factual content needed for truthfulness judgments. These issues are load-bearing for the paper's main conclusion, so the contribution is currently conditional on additional statistical and fidelity analysis.
major comments (5)
- [4.1] In Section 4.1 the authors state that because bootstrapped 95% confidence intervals for the MAE difference (−0.167 to 0.112) and MSE difference (−0.767 to 0.340) include zero, there is no statistically significant difference, and then conclude that the modalities are comparable. This is the central claim of RQ1, but non-significance is not evidence of equivalence; on the six-level truthfulness scale, an MAE difference of 0.112 is a meaningful degradation. The paper needs a pre-specified equivalence margin, a two one-sided test (TOST), or, at minimum, an effect size with a power analysis to justify 'comparable.' Without this, the headline effectiveness conclusion is underdetermined.
- [5.2] Section 5.2 repeats the same logical gap: Table 4 and the simulation in Figure 6 are used to conclude comparable performance under equal time constraints because 'no statistical significance is detected.' The same equivalence-margin requirement applies. Additionally, Section 5.2's judgment multiplier is +15%, yet Section 7 states that workers 'complete nearly double the number of judgments within the same time frame.' These statements are inconsistent, and the conclusions section should be reconciled with Table 4.
- [3.2] The summary-fidelity premise is asserted but never validated. The prompt in Figure 1 instructs the LLM to produce query-focused summaries, and Section 3.2 states the design goal of preserving factuality and stance, but the paper provides no evaluation of whether the summaries omit key caveats or introduce distortions, no comparison with another summarization model or prompt, and no error analysis of summaries. Because the general claim is that LLM summaries can replace full evidence, the absence of any fidelity check leaves open that the observed comparability, if real, is an artifact of this particular model and prompt. At minimum, a sample-level manual fidelity audit or a summary-evidence entailment check is needed.
- [5.1] The cost simulation in Section 5.1 contains a unit inconsistency. The authors state that 600 assessments require 23.67 hours at a U.S. minimum wage of $7.25/hour (≈ £6.08), and report costs of £171.60 (Standard) and £148.63 (Summary). These figures equal 23.67 × 7.25 and 20.50 × 7.25, i.e., dollar amounts, not pounds. The correct pound figures at £6.08/hour would be approximately £143.9 and £124.6. This error affects the quantitative cost-savings claim and must be corrected.
- [4.3] Table 3 contains internally inconsistent count/percentage pairs. In the Standard False row, the counts 39, 44, 35 sum to 118, but the percentages sum to 100% and correspond to different denominators; in the Half True rows the counts sum to 100 but the percentages (48.5%, 30.1%, 20.6%) do not match those counts. The table therefore cannot be used to support the claim that error-type distributions are similar across modalities, and it needs to be recomputed or clearly explained.
minor comments (4)
- [4.2] The Mann-Whitney U test comparing Krippendorff's α values is performed on a very small set of aggregate values (six ground-truth levels plus overall); the paper should state the sample size and justify treating these values as independent observations, or use a worker-level bootstrap for agreement.
- [6.3] The notation 'σ2 = 0.67' and 'σ2 = 0.99' is inconsistent with the text's use of standard deviation elsewhere; please clarify whether these are variances or standard deviations.
- [6.1] The sentence 'In theStandard modality, there is a notable pattern where evidence is frequently considered at least moderately useful, which aligns with the reduced amount of information presented to the workers' seems to describe the Summary condition rather than the Standard condition; please rephrase.
- [3.3] The minimum time requirement of 3 seconds per statement is extremely low and could admit speeders; please report how many workers were excluded by this check and by the gold-standard questions.
Circularity Check
No significant circularity: the summary-versus-standard comparison is an empirical A/B study benchmarked against external PolitiFact expert labels, and the central claim is not assumed by the inputs.
full rationale
The paper is an empirical A/B comparison in which the central claim—that Summary evidence yields comparable accuracy and error metrics to Standard evidence while improving efficiency—is evaluated against external PolitiFact expert labels and observed worker judgments. No parameter is fitted to the outcome and then re-reported as a prediction; the comparison metrics (accuracy, MAE, MSE, time) are computed directly from crowd responses and ground truth. The reuse of the La Barbera et al. dataset and task design is inheritance of instrumentation, not circular reasoning: the prior work supplies the task layout and six-level scale, but the target result (summary versus full evidence) is not assumed by those inputs. The summarization prompt in Figure 1 is an LLM configuration choice, not a derivation step, and the paper explicitly acknowledges in the Conclusions that automatically generated summaries may omit critical nuances. Concerns about interpreting non-significant bootstrapped confidence intervals as equivalence (Sections 4.1 and 5.2) are statistical-evidence concerns, not circularity: 'no statistically significant difference' is not an input that produces the conclusion by construction. No self-citation chain forces the main finding, and no known result is merely renamed or repackaged. Hence no significant circularity is present.
Assumptions & free parameters
free parameters (3)
- Minimum time threshold =
3 seconds per statement
- Compensation rates =
£2.00 Standard, £1.80 Summary
- Summarization prompt wording =
Prompt in Figure 1
assumptions (5)
- domain assumption PolitiFact expert labels are treated as ground truth for accuracy and error metrics.
- domain assumption LLM-generated summaries preserve the factual content needed for truthfulness assessment.
- ad hoc to paper Absence of a statistically significant difference is interpreted as comparable effectiveness.
- domain assumption Workers passing gold-standard and 3-second checks are attentive and comparable across the two arms.
- domain assumption Bing-retrieved webpages are relevant evidence for each claim.
Cite this review
Pith. "Pith review of Efficiency and Effectiveness of LLM-Based Summarization of Evidence in Crowdsourced Fact-Checking." pith.science (2026). https://pith.science/paper/JMVADAU6
@misc{pith2026250118265,
author = {Pith},
title = {Pith review of: Efficiency and Effectiveness of LLM-Based Summarization of Evidence in Crowdsourced Fact-Checking},
year = {2026},
howpublished = {\url{https://pith.science/paper/JMVADAU6}},
note = {Machine review of arXiv:2501.18265}
}
read the original abstract
Evaluating the truthfulness of online content is critical for combating misinformation. This study examines the efficiency and effectiveness of crowdsourced truthfulness assessments through a comparative analysis of two approaches: one involving full-length webpages as evidence for each claim, and another using summaries for each evidence document generated with a large language model. Using an A/B testing setting, we engage a diverse pool of participants tasked with evaluating the truthfulness of statements under these conditions. Our analysis explores both the quality of assessments and the behavioral patterns of participants. The results reveal that relying on summarized evidence offers comparable accuracy and error metrics to the Standard modality while significantly improving efficiency. Workers in the Summary setting complete a significantly higher number of assessments, reducing task duration and costs. Additionally, the Summary modality maximizes internal agreement and maintains consistent reliance on and perceived usefulness of evidence, demonstrating its potential to streamline large-scale truthfulness evaluations.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
Alim Al Ayub Ahmed, Ayman Aljabouh, Praveen Kumar Donepudi, and Myung Suh Choi. 2021. Detecting Fake News Using Machine Learning: A Systematic Literature Review.Psychology and Education Journal58, 1 (2021), 10 pages. https://doi.org/10.17762/pae.v58i1.1046
-
[2]
Jennifer Allen, Cameron Martel, and David G Rand. 2022. Birds of a Feather Don’t Fact-Check Each Other: Partisanship and the Evaluation of News in Twitter’s Birdwatch Crowdsourced Fact- Checking Program. InProceedings of the 2022 CHI Conference on Human Factors in Computing Systems (New Orleans, LA, USA)(CHI ’22). Association for Computing Machinery, New ...
-
[3]
Pepa Atanasova, Jakob Grue Simonsen, Christina Lioma, and Isabelle Augenstein. 2020. Generating Fact Checking Explanations. InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (Eds.). Association for Computational Linguistics, Online, 7352–7364.https://do...
doi:10.18653/v1/ 2020
-
[4]
Pepa Atanasova, Jakob Grue Simonsen, Christina Lioma, and Isabelle Augenstein. 2022. Fact Checking with Insufficient Evidence.Transactions of the Association for Computational Linguistics 10 (2022), 746–763. https://doi.org/10.1162/tacl_a_00486
-
[5]
Isabelle Augenstein, Timothy Baldwin, Meeyoung Cha, Tanmoy Chakraborty, Giovanni Luca Ciampaglia, David Corney, Renee DiResta, Emilio Ferrara, Scott Hale, Alon Halevy, Eduard Hovy, Heng Ji, Filippo Menczer, Ruben Miguez, Preslav Nakov, Dietram Scheufele, Shivam Sharma, and Giovanni Zagni. 2024. Factuality challenges in the era of large language models and...
-
[6]
Isabelle Augenstein, Christina Lioma, Dongsheng Wang, Lucas Chaves Lima, Casper Hansen, Christian Hansen, and Jakob Grue Simonsen. 2019. MultiFC: A Real-World Multi-Domain Dataset for Evidence-Based Fact Checking of Claims. InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference o...
2019
-
[7]
Nadav Borenstein, Greta Warren, Desmond Elliott, and Isabelle Augenstein. 2025. Can Community Notes Replace Professional Fact-Checkers? arXiv. https://doi.org/10.48550/arXiv.2502. 14132 arXiv:2502.14132 [cs.CL]
-
[8]
Erik Brand, Kevin Roitero, Michael Soprano, Afshin Rahimi, and Gianluca Demartini. 2022. A Neural Model to Jointly Predict and Explain Truthfulness of Statements.Journal of Data and Information Quality 15, 1, Article 4 (12 2022), 19 pages.https://doi.org/10.1145/3546917
doi:10.1145/3546917 2022
Show all 64 references
-
[9]
Alessandro Checco, Kevin Roitero, Eddy Maddalena, Stefano Mizzaro, and Gianluca Demartini
-
[10]
Botambu Collins, Dinh Tuyen Hoang, Ngoc Thanh Nguyen, and Dosam Hwang. 2021. Trends in Combating Fake News On Social Media – A Survey.Journal of Information and Telecommunication 5, 2 (2021), 247–266. https://doi.org/10.1080/24751839.2020.1847379
2021
-
[11]
Alphaeus Dmonte, Roland Oruche, Marcos Zampieri, Prasad Calyam, and Isabelle Augenstein
-
[12]
Xishuang Dong, Shouvon Sarker, and Lijun Qian. 2022. Integrating Human-in-the-loop into Swarm Learning for Decentralized Fake News Detection. InInternational Conference on Intelligent Data Science Technologies and Applications (IDSTA). IEEE, San Antonio, TX, USA, 46–53. https:...
2022
-
[13]
Tim Draws, David La Barbera, Michael Soprano, Kevin Roitero, Davide Ceolin, Alessandro Checco, and Stefano Mizzaro. 2022. The Effects of Crowd Worker Biases in Fact-Checking Tasks. In 2022 ACM Conference on Fairness, Accountability, and Transparency (FAccT ’22). Association fo...
2022
-
[14]
Tim Draws, Alisa Rieger, Oana Inel, Ujwal Gadiraju, and Nava Tintarev. 2021. A Checklist to Combat Cognitive Biases in Crowdsourcing. InProceedings of the Ninth AAAI Conference on Human Computation and Crowdsourcing. AAAI Press, Palo Alto, California, USA, 48–59. https://doi.o...
2021 doi
- [15]
-
[16]
Tibshirani
Bradley Efron and Robert J. Tibshirani. 1993.An Introduction to the Bootstrap. Monographs on Statistics and Applied Probability, Vol. 57. Chapman & Hall/CRC, New York. https: //doi.org/10.1201/9780429246593
1993 doi
-
[17]
Carsten Eickhoff. 2018. Cognitive Biases in Crowdsourcing. In Proceedings of the Eleventh ACM International Conference on Web Search and Data Mining(Marina Del Rey, CA, USA) (WSDM ’18). Association for Computing Machinery, New York, NY, USA, 162–170.https: //doi.org/10.1145/31...
2018
-
[18]
Nicola Ferro, Yubin Kim, and Mark Sanderson. 2019. Using Collection Shards to Study Retrieval Performance Effect Sizes.ACM Transactions on Information Systems37, 3, Article 30 (March 2019), 40 pages. https://doi.org/10.1145/3310364
2019 doi
-
[19]
Nicola Ferro and Gianmaria Silvello. 2016. A General Linear Mixed Models Approach to Study System Component Effects. InProceedings of the 39th International ACM SIGIR Conference on Research and Development in Information Retrieval(Pisa, Italy)(SIGIR ’16). Association for Compu...
2016 doi
-
[20]
Nicola Ferro and Gianmaria Silvello. 2018. Toward an anatomy of IR system component perfor- mances. Journal of the Association for Information Science and Technology69, 2 (2018), 187–200. https://doi.org/10.1002/asi.23910
2018 doi
-
[21]
Freedman
David A. Freedman. 2009.Statistical Models: Theory and Practice(2 ed.). Cambridge University Press, New York, NY, USA. https://doi.org/10.1017/CBO9780511815867
2009 doi
-
[22]
Meric Altug Gemalmaz and Ming Yin. 2021. Accounting for Confirmation Bias in Crowdsourced Label Aggregation. In Proceedings of the Thirtieth International Joint Conference on Artifi- cial Intelligence (IJCAI). International Joint Conferences on Artificial Intelligence Organiza...
2021 doi
-
[23]
Lei Han, Kevin Roitero, Ujwal Gadiraju, Cristina Sarasua, Alessandro Checco, Eddy Maddalena, and Gianluca Demartini. 2019. All Those Wasted Hours: On Task Abandonment in Crowdsourcing. In Proceedings of the Twelfth ACM International Conference on Web Search and Data Mining (Me...
2019
-
[24]
Linmei Hu, Siqi Wei, Ziwang Zhao, and Bin Wu. 2022. Deep Learning For Fake News Detection: A Comprehensive Survey.AI Open3 (2022), 133–155. https://doi.org/10.1016/j.aiopen. 2022.09.001
2022 doi
-
[25]
Christoph Hube, Besnik Fetahu, and Ujwal Gadiraju. 2019. Understanding and Mitigating Worker Biases in the Crowdsourced Collection of Subjective Judgments. InProceedings of the 2019 CHI Conference on Human Factors in Computing Systems(Glasgow, Scotland Uk)(CHI ’19). Associatio...
2019
-
[26]
Klaus Krippendorff. 2011. Computing Krippendorff’s Alpha-Reliability.https://repository. upenn.edu/asc_papers/43. Technical Report, University of Pennsylvania
2011
-
[27]
David La Barbera, Eddy Maddalena, Michael Soprano, Kevin Roitero, Gianluca Demartini, Davide Ceolin, Damiano Spina, and Stefano Mizzaro. 2024. Crowdsourced Fact-checking: Does It Actually Work? Information Processing & Management 61, 5 (2024), 103792. https: //doi.org/10.1016/...
2024
-
[28]
David La Barbera, Kevin Roitero, Gianluca Demartini, Stefano Mizzaro, and Damiano Spina
-
[29]
David M. J. Lazer, Matthew A. Baum, Yochai Benkler, Adam J. Berinsky, Kelly M. Greenhill, Filippo Menczer, Miriam J. Metzger, Brendan Nyhan, Gordon Pennycook, David Rothschild, Michael Schudson, Steven A. Sloman, Cass R. Sunstein, Emily A. Thorson, Duncan J. Watts, and Jonatha...
2018 doi
-
[30]
Huiru Li, Liangxiao Jiang, and Siqing Xue. 2023. Neighborhood Weighted Voting-Based Noise Correction for Crowdsourcing. ACM Transactions on Knowledge Discovery from Data17, 7, Article 96 (4 2023), 18 pages. https://doi.org/10.1145/3586998
2023 doi
-
[31]
Houjiang Liu, Anubrata Das, Alexander Boltz, Didi Zhou, Daisy Pinaroc, Matthew Lease, and Min Kyung Lee. 2024. Human-centered NLP Fact-checking: Co-Designing with Fact-checkers using Matchmaking for AI.Proceedings of the ACM on Human-Computer Interaction8, CSCW2, Article 423 (...
2024 doi
-
[32]
Houjiang Liu, Jacek Gwizdka, and Matthew Lease. 2024. Exploring Multidimensional Check- worthiness: Designing AI-assisted Claim Prioritization for Human Fact-checkers. arXiv.https: //doi.org/10.48550/arXiv.2412.08185
2024 doi
-
[33]
Mann and Donald R
Henry B. Mann and Donald R. Whitney. 1947. On a Test of Whether One of Two Random Variables is Stochastically Larger than the Other.Annals of Mathematical Statistics18, 1 (1947), 50–60. https://doi.org/10.1214/aoms/1177730491
1947
-
[34]
Syed Ishfaq Manzoor, Jimmy Singla, and Nikita. 2019. Fake News Detection Using Machine Learning approaches: A systematic Review. InProceedings of the 3rd International Conference on Trends in Electronics and Informatics (ICOEI). IEEE, Tirunelveli, India, 230–234. https: //doi....
2019
-
[35]
Cameron Martel, Jennifer Allen, Gordon Pennycook, and David G. Rand. 2024. Crowds Can Effectively Identify Misinformation at Scale.Perspectives on Psychological Science19, 2 (2024), 477–488. https://doi.org/10.1177/17456916231190388
2024 doi
-
[36]
Preslav Nakov, Alberto Barrón-Cedeño, Giovanni da San Martino, Firoj Alam, Julia Maria Struß, Thomas Mandl, Rubén Míguez, Tommaso Caselli, Mucahid Kutlu, Wajdi Zaghouani, Chengkai Li, Shaden Shaar, Gautam Kishore Shahi, Hamdy Mubarak, Alex Nikolov, Nikolay Babulkov, Yavuz Seli...
2022
-
[37]
Preslav Nakov, Giovanni Da San Martino, Tamer Elsayed, Alberto Barrón-Cedeño, Rubén Míguez, Shaden Shaar, Firoj Alam, Fatima Haouari, Maram Hasanain, Nikolay Babulkov, Alex Nikolov, Gautam Kishore Shahi, Julia Maria Struß, and Thomas Mandl. 2021. The CLEF-2021 CheckThat! Lab o...
2021
-
[38]
Thanh Tam Nguyen, Matthias Weidlich, Hongzhi Yin, Bolong Zheng, Quang Huy Nguyen, and Quoc Viet Hung Nguyen. 2020. FactCatch: Incremental Pay-as-You-Go Fact Checking with Minimal User Effort. InProceedings of the 43rd International ACM SIGIR Conference on Research and Developm...
2020 doi
-
[39]
Olejnik and James Algina
Stephen F. Olejnik and James Algina. 2003. Generalized Eta and Omega Squared Statistics: Measures of Effect Size for Some Common Research Designs.Psychological Methods8, 4 (2003), 434–447. https://doi.org/10.1037/1082-989X.8.4.434
2003 doi
-
[40]
Gordon Pennycook and David G. Rand. 2019. Fighting misinformation on social media using crowdsourced judgments of news source quality.Proceedings of the National Academy of Sciences 116, 7 (2019), 2521–2526. https://doi.org/10.1073/pnas.1806781116
2019 doi
-
[41]
Gordon Pennycook and David G. Rand. 2021. The Psychology of Fake News.Trends in Cognitive Sciences 25, 5 (2021), 388–402. https://doi.org/10.1016/j.tics.2021.02.007
2021 doi
-
[42]
Yunke Qu, Kevin Roitero, David La Barbera, Damiano Spina, Stefano Mizzaro, and Gianluca Demartini. 2022. Combining Human and Machine Confidence in Truthfulness Assessment.Journal of Data and Information Quality15, 1, Article 5 (dec 2022), 17 pages.https://doi.org/10. 1145/3546916
2022
- [43]
-
[44]
LeveragingBehavioral Heterogeneity Across Markets for Cross-Market Training of Recommender Systems
KevinRoitero, BenCarterette, RishabhMehrotra, andMouniaLalmas.2020. LeveragingBehavioral Heterogeneity Across Markets for Cross-Market Training of Recommender Systems. InCompanion Proceedings of the Web Conference 2020(Taipei, Taiwan)(WWW ’20). Association for Computing Machin...
2020
-
[45]
Kevin Roitero, Michael Soprano, Shaoyang Fan, Damiano Spina, Stefano Mizzaro, and Gianluca Demartini. 2020. Can The Crowd Identify Misinformation Objectively? The Effects of Judgment Scale and Assessor’s Background. InProceedings of the 43rd International ACM SIGIR Conference ...
2020
-
[46]
Mohammed Saeed, Nicolas Traub, Maelle Nicolas, Gianluca Demartini, and Paolo Papotti. 2022. Crowdsourced Fact-Checking at Twitter: How Does the Crowd Compare With Experts?. In Proceedings of the 31st ACM International Conference on Information and Knowledge Management (Atlanta...
2022
-
[47]
Michael Schlichtkrull, Yulong Chen, Chenxi Whitehouse, Zhenyun Deng, Mubashara Akhtar, Rami Aly, Zhijiang Guo, Christos Christodoulopoulos, Oana Cocarascu, Arpit Mittal, James Thorne, and Andreas Vlachos. 2024. The Automated Verification of Textual Claims (AVeriTeC) Shared Tas...
2024
-
[48]
Vinay Setty. 2024. Surprising Efficacy of Fine-Tuned Transformers for Fact-Checking over Larger Language Models. InProceedings of the 47th International ACM SIGIR Conference on Research 17 and Development in Information Retrieval(Washington DC, USA)(SIGIR ’24). Association for...
2024 doi
-
[49]
Michael Soprano, Kevin Roitero, David La Barbera, Davide Ceolin, Damiano Spina, Gian- luca Demartini, and Stefano Mizzaro. 2024. Cognitive Biases in Fact-Checking and Their Countermeasures: A Review. Information Processing & Management 61, 3 (2024), 103672. https://doi.org/10....
2024
-
[50]
Michael Soprano, Kevin Roitero, David La Barbera, Davide Ceolin, Damiano Spina, Stefano Mizzaro, and Gianluca Demartini. 2021. The Many Dimensions of Truthfulness: Crowdsourcing Misinformation Assessments on a Multidimensional Scale.Information Processing & Management 58, 6 (2...
2021
-
[51]
All Abdullah Tanvir, Ehesas Mia Mahir, Saima Akhter, and Mohammad Rezwanul Huq. 2019. Detecting Fake News using Machine Learning and Deep Learning Algorithms. InProceedings of the 7th International Conference on Smart Computing & Communications (ICSCC). IEEE, Piscataway, NJ, U...
2019
-
[52]
James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal. 2018. FEVER: a Large-scale Dataset for Fact Extraction and VERification. InProceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Lan...
2018 doi
-
[53]
James Thorne, Andreas Vlachos, Oana Cocarascu, Christos Christodoulopoulos, and Arpit Mittal
-
[54]
Nikhita Vedula and Srinivasan Parthasarathy. 2021. FACE-KEG: Fact Checking Explained using KnowledgE Graphs. InProceedings of the 14th ACM International Conference on Web Search and Data Mining (Virtual Event, Israel)(WSDM ’21). Association for Computing Machinery, New York, N...
2021
-
[56]
Yuxia Wang, Revanth Gangi Reddy, Zain Muhammad Mujahid, Arnav Arora, Aleksandr Ruba- shevskii, Jiahui Geng, Osama Mohammed Afzal, Liangming Pan, Nadav Borenstein, Aditya Pillai, Isabelle Augenstein, Iryna Gurevych, and Preslav Nakov. 2024. Factcheck-Bench: Fine- Grained Evalua...
2024
-
[57]
Greta Warren, Irina Shklovski, and Isabelle Augenstein. 2025. Show Me the Work: Fact-Checkers’ RequirementsforExplainableAutomatedFact-Checking.In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI ’25). ACM, 1–21. https://doi.org/10.1145/ 370659...
2025
-
[58]
Jing Yang, Didier Vega-Oliveros, Tais Seibt, and Anderson Rocha. 2021. Scalable Fact-checking with Human-in-the-Loop. In2021 IEEE International Workshop on Information Forensics and Security (WIFS). IEEE, Montpellier, France, 1–6.https://doi.org/10.1109/WIFS53200.2021.9648388 18
2021
-
[59]
Barry Menglong Yao, Aditya Shah, Lichao Sun, Jin-Hee Cho, and Lifu Huang. 2023. End-to-End Multimodal Fact-Checking and Explanation Generation: A Challenging Dataset and Models. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Inform...
2023
-
[60]
Shane Culpepper, Oren Kurland, and Stefano Mizzaro
Fabio Zampieri, Kevin Roitero, J. Shane Culpepper, Oren Kurland, and Stefano Mizzaro. 2019. On Topic Difficulty in IR Evaluation: The Effect of Systems, Corpora, and System Components. In Proceedings of the 42nd International ACM SIGIR Conference on Research and Development in...
2019
-
[61]
Andy Zhao and Mor Naaman. 2023. Variety, Velocity, Veracity, and Viability: Evaluating the Contributions of Crowdsourced and Professional Fact-checking.https://doi.org/10.31235/ osf.io/yfxd3 19
2023
-
[2017]
Proceedings of the AAAI Conference on Human Computation and Crowdsourcing5, 1 (Sep
Let’s Agree to Disagree: Fixing Agreement Measures for Crowdsourcing. Proceedings of the AAAI Conference on Human Computation and Crowdsourcing5, 1 (Sep. 2017), 11–20. https://doi.org/10.1609/hcomp.v5i1.13306
2017 doi
-
[2019]
The FEVER2.0 Shared Task. In Proceedings of the Second Workshop on Fact Extrac- tion and VERification (FEVER), James Thorne, Andreas Vlachos, Oana Cocarascu, Christos Christodoulopoulos, and Arpit Mittal (Eds.). Association for Computational Linguistics, Hong Kong, China, 1–6....
-
[2020]
InAdvances in Information Retrieval, Joemon M
Crowdsourcing Truthfulness: The Impact of Judgment Scale and Assessor Bias. InAdvances in Information Retrieval, Joemon M. Jose, Emine Yilmaz, João Magalhães, Pablo Castells, Nicola Ferro, Mário J. Silva, and Flávio Martins (Eds.). Springer International Publishing, Cham, 207–...
- [2025]
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.