REVIEW 3 major objections 5 minor 77 references
Detecting Toxicity in News Articles: Application to Bulgarian
T0 review · 3 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read A Bulgarian news toxicity detector reaches 59% accuracy across nine labels.
desk verdict A useful new dataset for Bulgarian toxicity detection, but the evaluation is confounded by source selection, so the headline accuracy likely overstates content-based detection. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the meta-classifier: a logistic regression model whose inputs are the posterior class probabilities produced by several base classifiers, rather than raw text. Because the dataset is small (317 articles), the authors train one model per feature type and then let the meta-classifier learn how to weight those models' predictions. The base models cover complementary views of the text: BERT and XLM transformer representations, ElMo contextual word vectors, Universal Sentence Encoder embeddings, LSA low-rank vectors, NELA style and content features, stylometric counts, and six metadata features about the news medium. This stacking step is what lifts accuracy from 30.3% at the baseline to 59.0%.
What would settle it
Ask independent Bulgarian-speaking annotators to label the 96 non-toxic articles, or to re-label all 317; if a large share of the assumed-clean articles turn out to be toxic, retraining on corrected labels would likely drop the reported 59.0% accuracy and 39.7 macro-F1. More narrowly, the confusion matrix's heavy misclassification of fake news as conspiracy can be checked against a larger sample of those two classes.
Extended reading notes
Core claim
The paper's central claim is that a nine-way classifier can distinguish Bulgarian news articles into non-toxic and eight toxic categories—fake news, sensationalism, hate speech, conspiracy theories, anti-democratic, pro-authoritarian, defamation, and delusion—using only a few hundred labeled examples. The evidence is a 317-article dataset whose toxic labels come from a five-year human-curated archive of Bulgarian media, supplemented by 96 articles taken from media with no listed toxic items. The best model is not any single text representation but a meta-classifier: a logistic regression model trained on the posterior probabilities of several base classifiers built from BERT, XLM, ElMo, Universal Sentence Encoder, LSA, stylometric, NELA, and media features. It reports 59.0% accuracy and 39.7 macro-F1, compared with 30.3% accuracy and 5.2 macro-F1 for the majority-class baseline; the best single representation, LSA, reaches 55.6% accuracy and 42.1 macro-F1. The paper also reports that oversampling and a feed-forward neural network did not improve over logistic regression on this small dataset.
Load-bearing premise
The non-toxic training articles are assumed to be non-toxic only because they come from media that have no toxic articles listed in the monitoring archive; no one independently verified that these 96 articles are actually clean, so the non-toxic class and the headline accuracy figures could be built on mislabeled examples.
Editorial extensions
If this is right
- A nine-way toxic/non-toxic classifier for Bulgarian reaches 59.0% accuracy and 39.7 macro-F1, roughly doubling the majority baseline's F1, on a 317-article dataset.
- The meta-classifier outperforms every single representation, including LSA (55.6% accuracy, 42.1 macro-F1), so combining diverse feature models appears to be the most reliable path on small in-language datasets.
- English-translation features (BERT, USE, ElMo, NELA) perform close to or better than Bulgarian-native features, meaning English resources can be reused for low-resource toxicity detection.
- The confusion matrix shows the model is weakest on the rare toxic classes: the three smallest classes together cover less than 18% of the dataset and are rarely predicted.
- Because the dataset and code are released, the reported numbers can be reproduced and the classifier can be extended or rebalanced in future work.
Reading between the lines
- The paper's treatment of the non-toxic class is the least verified part of the data; an obvious extension is to have those 96 articles independently annotated, which would test whether the accuracy gain is real or an artifact of a clean-by-assumption class.
- Since the source archive allows multiple toxicity labels per article, reformulating the task as multi-label classification could better match the data and may improve recall on rare classes like delusion and anti-democratic.
- The finding that English-translation features nearly match Bulgarian-native ones suggests a cheap recipe for other low-resource languages: translate, apply English transfer models, and stack the probabilities; this recipe is testable on a second language.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents a multi-class toxicity detection system for Bulgarian news articles. The authors construct a dataset of 317 articles, of which 221 are labelled into eight toxic categories based on the Media Scan repository and 96 are labelled non-toxic from media without Media Scan toxicity listings. They compare feature sets including LSA, BERT, XLM, USE, ElMo, NELA, stylometric features, and media-level features, and then build a meta-classifier over posterior probabilities of the individual models. The best system (setup 14) is reported to reach 59.06% accuracy and 39.70% macro-F1, which the authors claim are sizable improvements over the majority-class baseline (30.30% accuracy, 5.17% macro-F1).
Significance. If validated, the paper would make a useful contribution as one of the few resources for Bulgarian toxicity detection: it introduces a new multi-label dataset, systematically compares several representations, and releases code and data. The multi-class setup with eight toxicity types plus a non-toxic class is more differentiated than much prior English-focused work. However, the central empirical claim is currently undermined by a source-label confound in the construction of the non-toxic set and by insufficient detail about potential information leakage in the LSA component. The significance of the contribution is therefore contingent on addressing these methodological issues.
major comments (3)
- [Section 3 (Data), Section 4.4 (Media Features), Section 5.3 (Meta Classifier), Table 3] The dataset construction introduces a confound between article-level toxicity and source reputation that prevents the headline results from being interpreted as evidence of article-level toxicity detection. Section 3 states that toxic articles come from media listed in Media Scan, while the 96 non-toxic articles were "fetched from media without toxicity examples in Media Scan." This makes the toxic/non-toxic label almost perfectly correlated with whether the source medium has ever been flagged, rather than with the content of the individual article. The media features introduced in Section 4.4 (editor, responsible person, bg server, popularity, domain person, days existing) are then used both as a standalone model (setup 12 in Table 3) and as inputs to the meta-classifier, whose input list in Section 5.3 explicitly includes the posteriors of setup 12. Setup 12 alone reaches 42.04% accuracy, a 12-point gain over the majority baseline, showing that source-level information carries substantial predictive signal. Consequently, the reported 59.06% accuracy and 39.70% macro-F1 may reflect a source-reputation shortcut rather than detection of toxic content. To support the article-level claim, the authors should re-collect or re-annotate non-toxic articles from the same media that also have toxic articles, or otherwise control for source in the evaluation, and should report results without media features.
- [Section 4.1 (LSA)] The description of the LSA model is ambiguous about whether the SVD is fitted inside each cross-validation fold. The text says "We trained a Latent Semantic Analysis (LSA) model on our data," but it does not state that the TF-IDF and SVD transformations are recomputed on the training portion of each fold in the 5-fold cross-validation described in Section 5.1. If the SVD is fit on all 317 articles before splitting, test-fold articles contribute to the learned low-dimensional space, which is a form of information leakage. This matters because LSA (setup 5) is the best individual text representation and feeds into the meta-classifier. The authors should clarify the exact fitting procedure; if the model is fit on the full data, the LSA results and the meta-classifier results must be recomputed with nested fitting.
- [Section 5.1, Table 3] No measures of variability or statistical significance are reported for any of the cross-validation results. With only 317 articles and nine classes, differences such as the 3.5% absolute accuracy gain of the meta-classifier over the best individual LSA model (59.06% vs. 55.59%) could easily fall within fold-to-fold variance. The paper repeatedly emphasizes "sizable improvements" over the baseline, but this claim is not supported without per-fold results, standard deviations, or a paired significance test between setup 14 and the relevant baselines and individual models.
minor comments (5)
- [Section 5.3] The construction of the meta-classifier is described only as "we made sure that we do not leak information about the labels when training the meta classifier." Please specify exactly how the posterior probabilities are produced (e.g., out-of-fold predictions from the inner cross-validation) so that the procedure is reproducible.
- [Section 4.4] There is a typo in the example: "Januarty 1, 2005" should read "January 1, 2005."
- [Section 4.5] The word "1024-demnsional" should be "1024-dimensional."
- [Section 7 (Conclusion)] The phrase "ElMo, BERT, xand XLM" contains a typo: "xand" should be "and."
- [Section 5.1] The statement that 15,000 additional experiments were run for fine-tuning is not accompanied by the hyperparameter search space or the selection criterion; adding this information would improve reproducibility.
Circularity Check
No circularity: evaluation is an external benchmark against manually curated labels; self-citations are contextual and non-load-bearing.
full rationale
The central derivation chain is empirical: labels come from Media Scan, a five-year manually curated external database, and the classifier's predictions are compared against these labels in 5-fold cross-validation. No parameter is fitted to the test labels, and no 'prediction' is defined in terms of the labels it claims to predict. The meta-classifier stacks posterior probabilities of individual models, which is a standard ensemble construction and not circular. Self-citations (e.g., Hardalov et al. 2016, Karadzhov et al. 2017a, Dinkov et al. 2019) appear only in related-work surveys and are not used to justify the reported 59.06% accuracy or the dataset's toxicity assignments. The source-label confound noted by the skeptic (non-toxic articles sourced from media with no Media Scan toxic listings, plus media-only features) is a validity threat to the benchmark, but it is not a circular derivation: the non-toxic label is not defined as a function of the media features, and the media-only model's 42.04% accuracy is an empirical result, not an identity. Thus the paper's claims, while possibly over-optimistic due to the confound, are not circular by construction.
Assumptions & free parameters
free parameters (4)
- LSA dimensions =
15 (title), 200 (body)
- BERT pooling strategies =
REDUCE_MAX (title), CLS_TOKEN (body)
- Logistic regression hyperparameters =
not reported; tuned via internal cross-validation
- Meta-classifier input models =
posteriors from setups 2-5, 7-10, 12 (153 dims)
assumptions (4)
- domain assumption Media Scan's manual toxicity labels are accurate.
- domain assumption Articles from media without toxicity listings are non-toxic.
- domain assumption English translations preserve toxicity-relevant content.
- standard math Standard statistical assumptions of cross-validation and logistic regression hold.
Cite this review
Pith. "Pith review of Detecting Toxicity in News Articles: Application to Bulgarian." pith.science (2026). https://pith.science/paper/PWL3J2G6
@misc{pith2026190809785,
author = {Pith},
title = {Pith review of: Detecting Toxicity in News Articles: Application to Bulgarian},
year = {2026},
howpublished = {\url{https://pith.science/paper/PWL3J2G6}},
note = {Machine review of arXiv:1908.09785}
}
read the original abstract
Online media aim for reaching ever bigger audience and for attracting ever longer attention span. This competition creates an environment that rewards sensational, fake, and toxic news. To help limit their spread and impact, we propose and develop a news toxicity detector that can recognize various types of toxic content. While previous research primarily focused on English, here we target Bulgarian. We created a new dataset by crawling a website that for five years has been collecting Bulgarian news articles that were manually categorized into eight toxicity groups. Then we trained a multi-class classifier with nine categories: eight toxic and one non-toxic. We experimented with different representations based on ElMo, BERT, and XLM, as well as with a variety of domain-specific features. Due to the small size of our dataset, we created a separate model for each feature type, and we ultimately combined these models into a meta-classifier. The evaluation results show an accuracy of 59.0% and a macro-F1 score of 39.7%, which represent sizable improvements over the majority-class baseline (Acc=30.3%, macro-F1=5.2%).
Figures
Reference graph
Works this paper leans on
-
[1]
Pepa Atanasova, Llu\' i s M\` a rquez, Alberto Barr\' o n-Cede\ n o, Tamer Elsayed, Reem Suwaileh, Wajdi Zaghouani, Spas Kyuchukov, Giovanni Da San Martino, and Preslav Nakov. 2018. Overview of the CLEF-2018 CheckThat! lab on automatic identification and verification of political claims, T ask 1: Check-worthiness. In CLEF 2018 Working Notes\/ . CEUR-WS.or...
work page 2018
-
[2]
Pepa Atanasova, Preslav Nakov, Georgi Karadzhov, Mitra Mohtarami, and Giovanni Da San Martino. 2019. Overview of the CLEF-2019 CheckThat! Lab on Automatic Identification and Verification of Claims. Task 1: Check-Worthiness . In CLEF 2019 Working Notes\/ . CEUR-WS.org, Lugano, Switzerland
work page 2019
-
[3]
Mouhamadou Lamine Ba, Laure Berti-Equille, Kushal Shah, and Hossam M. Hammady. 2016. VERA : A platform for veracity estimation over web data. In Proceedings of the 25th International Conference Companion on World Wide Web\/ . Montr \'e al, Qu \'e bec, Canada, WWW '16, pages 159--162
work page 2016
-
[4]
Ramy Baly, Georgi Karadzhov, Dimitar Alexandrov, James Glass, and Preslav Nakov. 2018 a . Predicting factuality of reporting and bias of news media sources. In Proceedings of the Conference on Empirical Methods in Natural Language Processing\/ . Brussels, Belgium, EMNLP '18, pages 3528--3539
work page 2018
-
[5]
Ramy Baly, Georgi Karadzhov, Abdelrhman Saleh, James Glass, and Preslav Nakov. 2019. Multi-task ordinal regression for jointly predicting the trustworthiness and the leading political ideology of news media. In Proceedings of the 17th Annual Conference of the North American Chapter of the Association for Computational Linguistics\/ . Minneapolis, MN, USA,...
work page 2019
-
[6]
Ramy Baly, Mitra Mohtarami, James Glass, Llu\' i s M\` a rquez, Alessandro Moschitti, and Preslav Nakov. 2018 b . Integrating stance detection and fact checking in a unified corpus. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics\/ . New Orleans, LA, USA, NAACL-HLT '18, pages 21--27
work page 2018
-
[7]
Alberto Barr\' o n-Cede\ no, Giovanni Da San Martino, Israa Jaradat, and Preslav Nakov. 2019. Proppy: Organizing the news based on their propagandistic content . Information Processing & Management\/ 56(5):1849 -- 1864
work page 2019
-
[8]
Alberto Barr\' o n-Cede\ n o, Tamer Elsayed, Reem Suwaileh, Llu\' i s M\` a rquez, Pepa Atanasova, Wajdi Zaghouani, Spas Kyuchukov, Giovanni Da San Martino, and Preslav Nakov. 2018. Overview of the CLEF-2018 CheckThat! lab on automatic identification and verification of political claims, T ask 2: Factuality. In CLEF 2018 Working Notes\/ . CEUR-WS.org, Avi...
work page 2018
Show all 77 references
-
[9]
Ann M Brill. 2001. Online journalists embrace new marketing function. Newspaper Research Journal\/ 22(2):28--40
2001
-
[10]
Canini, Bongwon Suh, and Peter L
Kevin R. Canini, Bongwon Suh, and Peter L. Pirolli. 2011. Finding credible information sources in social networks based on content and social structure. In Proceedings of the IEEE International Conference on Privacy, Security, Risk, and Trust, and the IEEE International Confer...
2011
-
[11]
Carlos Castillo, Marcelo Mendoza, and Barbara Poblete. 2011. Information credibility on T witter. In Proceedings of the 20th International Conference on World Wide Web\/ . Hyderabad, India, WWW '11, pages 675--684
2011
-
[12]
Daniel Cer, Yinfei Yang, Sheng-yi Kong, Nan Hua, Nicole Limtiaco, Rhomni St John, Noah Constant, Mario Guajardo-Cespedes, Steve Yuan, Chris Tar, et al. 2018. Universal sentence encoder. arXiv preprint arXiv:1803.11175\/
2018 arXiv
-
[13]
Abhijnan Chakraborty, Bhargavi Paranjape, Kakarla Kakarla, and Niloy Ganguly. 2016. Stop clickbait: Detecting and preventing clickbaits in online news media. In Proceedings of the 2016 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining\/ . San...
2016
-
[14]
Nitesh V Chawla, Kevin W Bowyer, Lawrence O Hall, and W Philip Kegelmeyer. 2002. SMOTE : synthetic minority over-sampling technique. Journal of artificial intelligence research\/ 16:321--357
2002
-
[15]
Cheng Chen, Kui Wu, Venkatesh Srinivasan, and Xudong Zhang. 2013. Battling the I nternet W ater A rmy: detection of hidden paid posters. In Proceedings of the 2013 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining\/ . Niagara, Canada, ASONAM ...
2013
-
[16]
Giovanni Da San Martino, Seunghak Yu, Alberto Barron-Cedeno, Rostislav Petrov, and Preslav Nakov. 2019. Fine-grained analysis of propaganda in news articles. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing\/ . Hong Kong, China, EMNLP '19
2019
-
[17]
Kareem Darwish, Dimitar Alexandrov, Preslav Nakov, and Yelena Mejova. 2017 a . Seminar users in the A rabic T witter sphere. In Proceedings of the 9th International Conference on Social Informatics\/ . Oxford, UK, SocInfo '17, pages 91--108
2017
-
[18]
Kareem Darwish, Walid Magdy, and Tahar Zanouda. 2017 b . Improved stance prediction in a user similarity feature space. In Proceedings of the Conference on Advances in Social Networks Analysis and Mining\/ . Sydney, Australia, ASONAM '17, pages 145--148
2017
-
[19]
Sohan De Sarkar, Fan Yang, and Arjun Mukherjee. 2018. Attending sentences to detect satirical fake news. In Proceedings of the 27th International Conference on Computational Linguistics\/ . Santa Fe, NM, USA, COLING '18, pages 3371--3380
2018
-
[20]
Leon Derczynski, Kalina Bontcheva, Maria Liakata, Rob Procter, Geraldine Wong Sak Hoi, and Arkaitz Zubiaga. 2017. SemEval-2017 Task 8: RumourEval : Determining rumour veracity and support for rumours. In Proceedings of the 11th International Workshop on Semantic Evaluation\/ ....
2017
-
[21]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT : Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North A merican Chapter of the Association for Computational Linguistics\/ . ...
2019
-
[22]
Yoan Dinkov, Ahmed Ali, Ivan Koychev, and Preslav Nakov. 2019. Predicting the leading political ideology of Y outube channels using acoustic, textual and metadata information. In Proceedings of the 20th Annual Conference of the International Speech Communication Association\/ ...
2019
-
[23]
Xin Luna Dong, Evgeniy Gabrilovich, Kevin Murphy, Van Dang, Wilko Horn, Camillo Lugaresi, Shaohua Sun, and Wei Zhang. 2015. Knowledge-based trust: Estimating the trustworthiness of web sources. Proc. VLDB Endow.\/ 8(9):938--949
2015
-
[24]
Sebastian Dungs, Ahmet Aker, Norbert Fuhr, and Kalina Bontcheva. 2018. Can rumour stance alone predict veracity? In Proceedings of the 27th International Conference on Computational Linguistics\/ . Santa Fe, NM, USA, COLING '18, pages 3360--3370
2018
-
[25]
Tamer Elsayed, Preslav Nakov, Alberto Barr\' o n-Cede\ n o, Maram Hasanain, Reem Suwaileh, Pepa Atanasova, and Giovanni Da San Martino. 2019 a . CheckThat ! at CLEF 2019: Automatic identification and verification of claims. In Proceedings of the 41st European Conference on Inf...
2019
-
[26]
Tamer Elsayed, Preslav Nakov, Alberto Barr\' o n-Cede \ n o, Maram Hasanain, Reem Suwaileh, Giovanni Da San Martino , and Pepa Atanasova. 2019 b . Overview of the CLEF-2019 CheckThat! : Automatic identification and verification of claims. In Experimental IR Meets Multilinguali...
2019
-
[27]
Genevieve Gorrell, Elena Kochkina, Maria Liakata, Ahmet Aker, Arkaitz Zubiaga, Kalina Bontcheva, and Leon Derczynski. 2019. S em E val-2019 task 7: R umour E val, determining rumour veracity and support for rumours. In Proceedings of the 13th International Workshop on Semantic...
2019
-
[28]
Jesse Graham, Jonathan Haidt, and Brian A Nosek. 2009. Liberals and conservatives rely on different sets of moral foundations. Journal of personality and social psychology\/ 96(5):1029
2009
-
[29]
Meyer, and Iryna Gurevych
Andreas Hanselowski, Avinesh PVS, Benjamin Schiller, Felix Caspelherr, Debanjan Chaudhuri, Christian M. Meyer, and Iryna Gurevych. 2018. A retrospective analysis of the fake news challenge stance-detection task. In Proceedings of the 27th International Conference on Computatio...
2018
-
[30]
Momchil Hardalov, Ivan Koychev, and Preslav Nakov. 2016. In search of credible news. In Proceedings of the 17th International Conference on Artificial Intelligence: Methodology, Systems, and Applications\/ . Varna, Bulgaria, AIMSA '16, pages 172--180
2016
-
[31]
Maram Hasanain, Reem Suwaileh, Tamer Elsayed, Alberto Barr\' o n-Cede \ n o, and Preslav Nakov. 2019. Overview of the CLEF-2019 CheckThat! Lab on Automatic Identification and Verification of Claims. Task 2: Evidence and Factuality . In CLEF 2019 Working Notes\/ . CEUR-WS.org, ...
2019
-
[32]
Benjamin Horne and Sibel Adali. 2017. This just in: Fake news packs a lot in title, uses simpler, repetitive content in text body, more similar to satire than real news. CoRR\/ abs/1703.09398
2017 arXiv
-
[33]
Benjamin Horne, Sibel Adali, and Sujoy Sikdar. 2017. Identifying the social signals that drive online discussions: A case study of Reddit communities. In Proceedings of the 26th IEEE International Conference on Computer Communication and Networks\/ . Vancouver, Canada, ICCCN '...
2017
-
[34]
Horne, William Dron, Sara Khedr, and Sibel Adali
Benjamin D. Horne, William Dron, Sara Khedr, and Sibel Adali. 2018 a . Assessing the news landscape: A multi-module toolkit for evaluating the credibility of news. In Proceedings of the The Web Conference\/ . Lyon, France, WWW '18, pages 235--238
2018
-
[35]
Horne, Sara Khedr, and Sibel Adali
Benjamin D. Horne, Sara Khedr, and Sibel Adali. 2018 b . Sampling the news producers: A large news and feature data set for the study of the complex media landscape. In Proceedings of the Twelfth International Conference on Web and Social Media\/ . Stanford, CA, USA, ICWSM '18...
2018
-
[36]
Clayton Hutto and Eric Gilbert. 2014. VADER : A parsimonious rule-based model for sentiment analysis of social media text. In Proceedings of the 8th International Conference on Weblogs and Social Media\/ . Ann Arbor, MI, USA, ICWSM '14
2014
-
[37]
Georgi Karadzhov, Pepa Gencheva, Preslav Nakov, and Ivan Koychev. 2017 a . We built a fake news & click-bait filter: What happened next will blow your mind! In Proceedings of the International Conference on Recent Advances in Natural Language Processing\/ . Varna, Bulgaria, RA...
2017
-
[38]
Georgi Karadzhov, Preslav Nakov, Llu\' i s M\` a rquez, Alberto Barr\'on-Cede \ n o, and Ivan Koychev. 2017 b . Fully automated fact checking using external sources. In Proceedings of the International Conference on Recent Advances in Natural Language Processing\/ . Varna, Bul...
2017
-
[39]
Elena Kochkina, Maria Liakata, and Arkaitz Zubiaga. 2018. All-in-one: Multi-task learning for rumour verification. In Proceedings of the International Conference on Computational Linguistics\/ . Santa Fe, NM, USA, COLING '18, pages 3402--3413
2018
-
[40]
Vivek Kulkarni, Junting Ye, Steve Skiena, and William Yang Wang. 2018. Multi-view models for political ideology detection of news articles. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing\/ . Brussels, Belgium, EMNLP '18, pages 3518--3527
2018
-
[41]
Guillaume Lample and Alexis Conneau. 2019. Cross-lingual language model pretraining. CoRR\/ abs/1901.07291
2019 arXiv
-
[42]
Lazer, Matthew A
David M.J. Lazer, Matthew A. Baum, Yochai Benkler, Adam J. Berinsky, Kelly M. Greenhill, Filippo Menczer, Miriam J. Metzger, Brendan Nyhan, Gordon Pennycook, David Rothschild, Michael Schudson, Steven A. Sloman, Cass R. Sunstein, Emily A. Thorson, Duncan J. Watts, and Jonathan...
2018
-
[43]
Yaliang Li, Jing Gao, Chuishi Meng, Qi Li, Lu Su, Bo Zhao, Wei Fan, and Jiawei Han. 2016. A survey on truth discovery. SIGKDD Explor. Newsl.\/ 17(2):1--16
2016
-
[44]
Y. Lin, J. Hoover, G. Portillo-Wightman, C. Park, M. Dehghani, and H. Ji. 2018. Acquiring background knowledge to improve moral value prediction. In Proceedings of the 2018 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining\/ . Los Alamitos, C...
2018
-
[45]
Jansen, Kam-Fai Wong, and Meeyoung Cha
Jing Ma, Wei Gao, Prasenjit Mitra, Sejeong Kwon, Bernard J. Jansen, Kam-Fai Wong, and Meeyoung Cha. 2016. Detecting rumors from microblogs with recurrent neural networks. In Proceedings of the 25th International Joint Conference on Artificial Intelligence\/ . New York, NY, USA...
2016
-
[46]
Jing Ma, Wei Gao, and Kam-Fai Wong. 2017. Detect rumors in microblog posts using propagation structure via kernel learning. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics\/ . Vancouver, Canada, ACL '17, pages 708--717
2017
-
[47]
Suman Kalyan Maity, Aishik Chakraborty, Pawan Goyal, and Animesh Mukherjee. 2017. D etection of sockpuppets in social media. In Proceedings of the ACM Conference on Computer Supported Cooperative Work and Social Computing\/ . Portland, OR, USA, CSCW '17, pages 243--246
2017
-
[48]
Todor Mihaylov, Georgi Georgiev, and Preslav Nakov. 2015 a . F inding opinion manipulation trolls in news community forums. In Proceedings of the Nineteenth Conference on Computational Natural Language Learning\/ . Beijing, China, CoNLL '15, pages 310--314
2015
-
[49]
Todor Mihaylov, Ivan Koychev, Georgi Georgiev, and Preslav Nakov. 2015 b . E xposing paid opinion manipulation trolls. In Proceedings of the International Conference Recent Advances in Natural Language Processing\/ . Hissar, Bulgaria, RANLP '15, pages 443--450
2015
-
[50]
Todor Mihaylov, Tsvetomila Mihaylova, Preslav Nakov, Llu\' i s M\` a rquez, Georgi Georgiev, and Ivan Koychev. 2018. The dark side of news community forums: Opinion manipulation trolls. Internet Research\/ 28(5):1292--1312
2018
-
[51]
Todor Mihaylov and Preslav Nakov. 2016. H unting for troll comments in news community forums. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics\/ . Berlin, Germany, ACL '16, pages 399--405
2016
-
[52]
Tsvetomila Mihaylova, Georgi Karadzhov, Pepa Atanasova, Ramy Baly, Mitra Mohtarami, and Preslav Nakov. 2019. S em E val-2019 task 8: Fact checking in community question answering forums. In Proceedings of the 13th International Workshop on Semantic Evaluation\/ . Minneapolis, ...
2019
-
[53]
Tsvetomila Mihaylova, Preslav Nakov, Llu\' i s M\` a rquez, Alberto Barr\'on-Cede \ n o, Mitra Mohtarami, Georgi Karadjov, and James Glass. 2018. Fact checking in community forums. In Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence\/ . New Orleans, ...
2018
-
[54]
Lewis Mitchell, Morgan R Frank, Kameron Decker Harris, Peter Sheridan Dodds, and Christopher M Danforth. 2013. The geography of happiness: Connecting T witter sentiment and expression, demographics, and objective characteristics of place. PloS one\/ 8(5):e64417
2013
-
[55]
Mitra Mohtarami, Ramy Baly, James Glass, Preslav Nakov, Llu\' i s M\` a rquez, and Alessandro Moschitti. 2018. Automatic stance detection using end-to-end memory networks. In Proceedings of the Annual Conference of the North American Chapter of the Association for Computationa...
2018
-
[56]
Mitra Mohtarami, James Glass, and Preslav Nakov. 2019. Contrastive language adaptation for cross-lingual stance detection. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing\/ . Hong Kong, China, EMNLP '19
2019
-
[57]
Subhabrata Mukherjee and Gerhard Weikum. 2015. Leveraging joint interactions for credibility analysis in news communities. In Proceedings of the 24th ACM International on Conference on Information and Knowledge Management\/ . Melbourne, Australia, CIKM '15, pages 353--362
2015
-
[58]
Preslav Nakov, Alberto Barr\' o n-Cede\ n o, Tamer Elsayed, Reem Suwaileh, Llu\' i s M\` a rquez, Wajdi Zaghouani, Pepa Atanasova, Spas Kyuchukov, and Giovanni Da San Martino. 2018. Overview of the CLEF-2018 CheckThat! lab on automatic identification and verification of politi...
2018
-
[59]
Nguyen, Aditya Kharosekar, Matthew Lease, and Byron C
An T. Nguyen, Aditya Kharosekar, Matthew Lease, and Byron C. Wallace. 2018. An interpretable joint graphical model for fact-checking from crowds. In Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence\/ . New Orleans, LA, USA, AAAI '18, pages 1511--1518
2018
-
[60]
Ver\' o nica P\' e rez-Rosas, Bennett Kleinberg, Alexandra Lefevre, and Rada Mihalcea. 2018. Automatic detection of fake news. In Proceedings of the 27th International Conference on Computational Linguistics\/ . Santa Fe, NM, USA, COLING '18, pages 3391--3401
2018
-
[61]
Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018. Deep contextualized word representations. In Proceedings of the Conference of the North A merican Chapter of the Association for Computational Linguistics\/ . Ne...
2018
-
[62]
Kashyap Popat, Subhabrata Mukherjee, Jannik Str\" o tgen, and Gerhard Weikum. 2016. Credibility assessment of textual claims on the web. In Proceedings of the 25th ACM International on Conference on Information and Knowledge Management\/ . Indianapolis, IN, USA, CIKM '16, page...
2016
-
[63]
Kashyap Popat, Subhabrata Mukherjee, Jannik Str\" o tgen, and Gerhard Weikum. 2017. Where the truth lies: Explaining the credibility of emerging claims on the W eb and social media. In Proceedings of the 26th International Conference on World Wide Web Companion\/ . Perth, Aust...
2017
-
[64]
Kashyap Popat, Subhabrata Mukherjee, Jannik Str\" o tgen, and Gerhard Weikum. 2018. CredEye : A credibility lens for analyzing and explaining misinformation. In Proceedings of The Web Conference 2018\/ . Lyon, France, WWW '18, pages 155--158
2018
-
[65]
Martin Potthast, Johannes Kiesel, Kevin Reinartz, Janek Bevendorff, and Benno Stein. 2018. A stylometric inquiry into hyperpartisan and fake news. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics\/ . Melbourne, Australia, ACL '18, page...
2018
-
[66]
Hannah Rashkin, Eunsol Choi, Jin Yea Jang, Svitlana Volkova, and Yejin Choi. 2017. Truth of varying shades: Analyzing language in fake news and political fact-checking. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing\/ . Copenhagen, De...
2017
-
[67]
Marta Recasens, Cristian Danescu-Niculescu-Mizil, and Dan Jurafsky. 2013. Linguistic models for analyzing and detecting biased language. In Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics\/ . Sofia, Bulgaria, ACL '13, pages 1650--1659
2013
-
[68]
Benjamin Riedel, Isabelle Augenstein, Georgios P Spithourakis, and Sebastian Riedel. 2017. A simple but tough-to-beat baseline for the Fake News Challenge stance detection task. ArXiv:1707.03264\/
2017 arXiv
-
[69]
Kai Shu, Amy Sliva, Suhang Wang, Jiliang Tang, and Huan Liu. 2017. Fake news detection on social media: A data mining perspective. SIGKDD Explor. Newsl.\/ 19(1):22--36
2017
-
[70]
James Thorne, Mingjie Chen, Giorgos Myrianthous, Jiashu Pu, Xiaoxuan Wang, and Andreas Vlachos. 2017. Fake news stance detection using stacked ensemble of classifiers. In Proceedings of the EMNLP Workshop on Natural Language Processing meets Journalism\/ . Copenhagen, Denmark,...
2017
-
[71]
James Thorne and Andreas Vlachos. 2018. Automated fact checking: Task formulations, methods and future directions. In Proceedings of the International Conference on Computational Linguistics\/ . Santa Fe, NM, USA, COLING '18, pages 3346--3359
2018
-
[72]
James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal. 2018. FEVER : a large-scale dataset for fact extraction and VERification . In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics\/ . New Orle...
2018
-
[73]
Soroush Vosoughi, Deb Roy, and Sinan Aral. 2018. The spread of true and false news online. Science\/ 359(6380):1146--1151
2018
-
[74]
Yifan Zhang, Giovanni Da San Martino, Alberto Barrón-Cedeño, Salvatore Romeo, Jisun An, Haewoon Kwak, Todor Staykovski, Israa Jaradat, Georgi Karadzhov, Ramy Baly, Kareem Darwish, and Preslav Nakov James Glass. 2019. Tanbih: Get to know what you are reading. In Proceedings of ...
2019
-
[75]
Dimitrina Zlatkova, Preslav Nakov, and Ivan Koychev. 2019. Fact-checking meets fauxtography: Verifying claims about images. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing\/ . Hong Kong, China, EMNLP '19
2019
-
[76]
, " * write output.state after.block = add.period write newline
ENTRY address author booktitle chapter doi edition editor howpublished institution journal key month note number organization pages publisher school series title type url volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.se...
-
[77]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.