REVIEW 3 major objections 5 minor 1 cited by
EDTok: A Dataset for Eating Disorder Content on TikTok
T0 review · 3 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read This paper presents EDTok, a curated dataset of 43,040 TikTok videos related to eating disorders, spanning January 2019 to June 2024, together with 577,071 comments and replies and a collection framework.
desk verdict Useful dataset, but the text-only Gemini filter keeps 'comprehensive' from being true; send to review with revisions. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the EDTok dataset and its curation pipeline. The pipeline starts with a manually assembled list of eating-disorder keywords and hashtags, including deliberate misspellings such as 'edrec0very' used to evade moderation, queried through TikTok's Research API; the PykTok module downloads the video files; then a two-stage filter removes unrelated posts, first by dropping videos whose metadata matches irrelevant keywords, then by running a large language model prompt on each video description that labels the video as eating-disorder-related or not. The classifier was validated on 200 randomly sampled videos with reported 99% accuracy and then applied to the full collection, followed by a 300-video manual check of the filtered result. This mechanism carries the argument because the dataset's research value depends entirely on the filter separating eating-disorder content from the roughly 20% unrelated videos that the initial keyword query pulls in.
What would settle it
Have two independent human annotators review a random sample of 1,000 videos labeled relevant by the pipeline, drawn across all collection years, and compare their labels to the classifier's; if the human-classifier agreement falls well below 99%, especially on videos without explicit eating-disorder keywords, the filter does not generalize as claimed.
Extended reading notes
Core claim
The central discovery is the dataset itself: a curated corpus of 43,040 TikTok videos on eating-disorder topics, plus 577,071 comments and replies, collected from a query set of keywords and hashtags and filtered through a two-stage pipeline that first removes obviously unrelated posts and then uses an off-the-shelf large language model classifier on video descriptions to keep only eating-disorder-relevant content. The paper reports that a manual review of 200 videos found the classifier about 99% accurate, and that a later random check of 300 videos in the final set found all of them relevant. The dataset spans January 1, 2019 to June 28, 2024, and records engagement totals of about 80 million likes, 537 million views, 577 thousand comments, and 962 thousand shares from 10,897 unique users. In the authors' framing, the contribution is not a finding about eating disorders so much as an instrument: a reproducible collection method and a public resource of video IDs and metadata that can support research on content spread, user engagement, moderation, and the pandemic's influence on online eating-disorder discourse.
Load-bearing premise
The curation pipeline assumes that a single text classifier, written with a manually authored prompt and checked on only 200 videos, correctly identifies eating-disorder relevance for every one of the more than 56,000 descriptions it then labels.
Editorial extensions
If this is right
- Researchers can trace monthly video volume and engagement before, during, and after the pandemic, testing whether COVID-19 increased eating-disorder content on TikTok.
- The dataset's combination of video files, descriptions, and comments enables multimodal studies linking visual content, caption sentiment, and community responses.
- The keyword list and filtering framework can be reused or extended to collect other sensitive health topics, and the video IDs permit replication without redistributing copyrighted videos.
- Analysis of engagement and moderation-evading hashtags can inform whether TikTok's moderation efforts are curbing harmful content or inadvertently suppressing recovery-oriented posts.
- The included comments and replies allow studying how audiences react emotionally to recovery narratives versus harmful content, as the paper's emotion analysis begins to do.
Reading between the lines
- Inference: The dataset's dependence on a single untuned classifier means its validity could be checked by re-annotating a stratified sample across years, engagement levels, and hashtag families; a large disagreement with human labels would not overturn the corpus but would narrow the claims that can be drawn from it.
- Inference: Because the query relies on known hashtags and manual curation, the dataset is likely biased toward recovery and awareness content and may under-represent communities using newer or more obscure jargon; researchers comparing prevalence over time should treat volume changes cautiously.
- Inference: The same collection framework could be pointed at other health conditions or at platform policy changes to create comparable datasets, effectively turning the pipeline into a template for studying moderation dynamics on short-video platforms.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents EDTok, a curated dataset of 43,040 TikTok videos related to eating disorders, collected via a set of hashtags and keywords in Table 1, spanning January 2019 to June 2024, alongside 577,071 comments and replies. The curation pipeline (Section 'Data Filtering') first removes videos with manually identified irrelevant keywords, then applies a Google Gemini text classifier to each video description, validating on a random sample of 200 videos (reporting ~99% accuracy) and on 300 post-filter videos (all confirmed relevant). The paper then provides descriptive analyses of engagement statistics, temporal trends, topic modeling of video descriptions and comments, emotion analysis, and a multimodal analysis of a 500-video subsample. The central claim is that EDTok is a comprehensive, multimodal resource for studying eating disorder content on TikTok, enabling analyses of content spread, moderation, engagement, and pandemic-related trends.
Significance. If the dataset's curation is valid, EDTok would be a valuable public resource for the study of eating disorder discourse on TikTok, combining video files, metadata, comments, and temporal coverage across the COVID-19 pandemic. The paper also contributes a reproducible metadata collection framework built on the TikTok Research API, and it is transparent about several limitations. However, the dataset's central validity claim rests on a single classifier applied to text descriptions, with unmeasured recall and incomplete reproducibility details; if those issues are addressed, the dataset could support meaningful research in computational social science and digital health. The descriptive analyses (engagement, temporal, topic, emotion) are straightforward and mostly illustrative, but they do not independently validate the curation.
major comments (3)
- [Data Filtering] The curation pipeline applies the Gemini classifier to video descriptions only (Table 5), yet the abstract and Discussion claim a multimodal dataset and a 'comprehensive view' of eating disorder content. Because recall on the excluded set is never measured, videos whose eating-disorder relevance is carried by visuals or audio rather than by text (for example, posts captioned 'day 1' or with emoji-only descriptions) are silently discarded, biasing the final 43,040 videos toward textually self-describing posts. Please estimate the false-negative rate by sampling videos excluded by the classifier and annotating them for eating-disorder relevance, and report this estimate alongside the precision figures. Without a recall estimate, the temporal trend and moderation analyses in Figures 2 and 3 cannot be interpreted as describing the full population of eating-disorder videos on TikTok.
- [Data Filtering] Neither the Gemini model version nor the date of execution is provided, and the manual keyword/hashtag removal list that reduced the dataset from 56,472 to 43,040 videos is not disclosed. This makes the curation pipeline non-reproducible and non-auditable, which is a load-bearing limitation for a dataset paper. Please specify the exact model identifier and query date, and include the full exclusion keyword list in an appendix or repository file.
- [Data Filtering] The validation of the classifier is based on 200 hand-checked examples with a self-reported 99% accuracy and a 300-video post-filter check with all confirmed related, but no inter-annotator agreement, confidence intervals, or stratified sampling are reported. The prompt in Table 5 explicitly lists terms such as 'anorexia', 'dieting', and 'weight loss', so the 99% figure may not generalize to the hard cases (misspelled hashtags, oblique references, recovery jargon) that motivated the study. I recommend adding a second annotator for a subset of the validation samples, reporting Cohen's kappa, and providing stratified validation across time periods and hashtag families.
minor comments (5)
- [Figure 1] The dataset flowchart is central to understanding the pipeline, but the exact counts at each stage (e.g., how many videos were removed by the keyword step versus the Gemini step) are not given in the text; please add these numbers so readers can audit the 56,472-to-43,040 reduction.
- [Table 3] The word 'residental' is misspelled and should be 'residential'; the top-words column would also benefit from consistent punctuation (e.g., listing each word separated by commas).
- [Table 4] The 'Fighting Battle' topic lists 'fighting, battle and fighter', which reads as a phrase containing 'and' rather than a list of keywords; please use a consistent separator (e.g., 'fighting, battle, fighter').
- [Text Analysis] The paragraph describing the Challenges topic contains 'signficant' instead of 'significant'; please correct the typo.
- [Text Analysis] The BERTopic results are presented without any quantitative quality metrics (e.g., topic coherence or topic diversity); adding such measures would strengthen confidence in the interpretability of the topics.
Circularity Check
No significant circularity: the dataset is assembled from external TikTok data via stated keywords and an external classifier, and the analyses are descriptive rather than predictions derived from the selection criteria.
full rationale
The paper's contribution is a collected and filtered dataset, not a derived prediction. The collection query terms are stated in Table 1, the classifier prompt is reproduced in Table 5, and the pipeline is documented in Figure 1. The analyses (engagement totals, temporal counts, topic models, emotion scores) are descriptive summaries of the resulting corpus; none is derived from a model whose parameters were fit to the same data and then called a prediction. The temporal 'increase after March 2020' is an empirical count of videos captured by the API, not a model output equivalent to the collection criteria. The only self-citations (Bickham et al. 2024; Lerman et al. 2023) supply previously used hashtags that are themselves listed in Table 1, so the argument does not reduce to an unverified self-cited theorem. The acknowledged limitations (keyword-selection bias, untuned Gemini classifier, API constraints) concern recall and representativeness, not circularity: the selection criteria are inputs, and the paper's claims about the selected corpus follow from descriptive statistics rather than from the criteria by construction.
Assumptions & free parameters
assumptions (3)
- domain assumption TikTok Research API data is representative of TikTok content.
- domain assumption Google Gemini classification of video descriptions accurately identifies eating-disorder-related videos.
- domain assumption The predefined keyword and hashtag set covers the eating-disorder content space.
Cite this review
Pith. "Pith review of EDTok: A Dataset for Eating Disorder Content on TikTok." pith.science (2026). https://pith.science/paper/JLJT2STG
@misc{pith2026250502250,
author = {Pith},
title = {Pith review of: EDTok: A Dataset for Eating Disorder Content on TikTok},
year = {2026},
howpublished = {\url{https://pith.science/paper/JLJT2STG}},
note = {Machine review of arXiv:2505.02250}
}
read the original abstract
Eating disorders, which include anorexia nervosa and bulimia nervosa, have been exacerbated by the COVID-19 pandemic, with increased diagnoses linked to heightened exposure to idealized body images online. TikTok, a platform with over a billion predominantly adolescent users, has become a key space where eating disorder content is shared, raising concerns about its impact on vulnerable populations. In response, we present a curated dataset of 43,040 TikTok videos, collected using keywords and hashtags related to eating disorders. Spanning from January 2019 to June 2024, this dataset, offers a comprehensive view of eating disorder-related content on TikTok. Our dataset has the potential to address significant research gaps, enabling analysis of content spread and moderation, user engagement, and the pandemic's influence on eating disorder trends. This work aims to inform strategies for mitigating risks associated with harmful content, contributing valuable insights to the study of digital health and social media's role in shaping mental health.
Figures
Figures from the paper (3 more)
Forward citations
Cited by 1 Pith paper
-
WhichTok? Comparing Three TikTok Data Acquisition Tools
Three TikTok data collection tools retrieve largely non-overlapping datasets for identical queries, especially for hashtag and keyword searches.
Reference graph
Works this paper leans on
-
[1]
, " * write output.state after.block = add.period write newline
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts #0 'before.a...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Alhuzali, H.; and Ananiadou, S. 2021. SpanEmo: Casting multi-label emotion classification as span-prediction. arXiv preprint arXiv:2101.10038
work page Pith review arXiv 2021
-
[4]
Bickham, C.; Kazemi-Nia, K.; Luceri, L.; Lerman, K.; and Ferrara, E. 2024. Hidden in Plain Sight: Exploring the Intersections of Mental Health, Eating Disorders, and Content Moderation on TikTok. In Workshop Proceedings of the 18th International AAAI Conference on Web and Social Media
work page 2024
-
[5]
Chochlakis, G.; Mahajan, G.; Baruah, S.; Burghardt, K.; Lerman, K.; and Narayanan, S. 2023. Leveraging label correlations in a multi-label setting: A case study in emotion. In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 1--5. IEEE
work page 2023
-
[6]
Grootendorst, M. 2022. BERTopic: Neural topic modeling with a class-based TF-IDF procedure. arXiv preprint arXiv:2203.05794
arXiv 2022
-
[7]
Herrick, S. S.; Hallward, L.; and Duncan, L. R. 2021. “This is just how I cope”: An inductive thematic analysis of eating disorder recovery content created and shared on TikTok using\# EDrecovery. International journal of eating disorders, 54(4): 516--526
work page 2021
-
[8]
Herzog, D. B.; Greenwood, D. N.; Dorer, D. J.; Flores, A. T.; Ekeblad, E. R.; Richards, A.; Blais, M. A.; and Keller, M. B. 2000. Mortality in eating disorders: a descriptive study. International Journal of Eating Disorders, 28(1): 20--26
work page 2000
Show all 28 references
-
[9]
Iqbal, M. 2021. TikTok revenue and usage statistics (2021). Business of apps, 1(1)
2021
-
[10]
K.; Patten, S
J Devoe, D.; Han, A.; Anderson, A.; Katzman, D. K.; Patten, S. B.; Soumbasis, A.; Flanagan, J.; Paslakis, G.; Vyver, E.; Marcoux, G.; et al. 2023. The impact of the COVID-19 pandemic on eating disorders: A systematic review. International Journal of Eating Disorders, 56(1): 5--25
2023
-
[11]
L.; Garc \' a, M
Jordan, G. L.; Garc \' a, M. D.; Esteban, P. S.; Mart \' n, A. S.; S \'a nchez, P. M.; Cubo, M. J.; Cano, M. M.; Del Barrio, J. G.; and Ayesa-Arriola, R. 2021. Tiktok, a vehicle for Pro-Ana and Pro-Mia content boosted by the COVID-19 pandemic. European Psychiatry, 64(S1): S703--S703
2021
-
[12]
Lerman, K.; Karnati, A.; Zhou, S.; Chen, S.; Kumar, S.; He, Z.; Yau, J.; and Horn, A. 2023. Radicalized by Thinness: Using a Model of Radicalization to Understand Pro-Anorexia Communities on Twitter. arXiv preprint arXiv:2305.11316
2023 arXiv
-
[13]
McCashin, D.; and Murphy, C. M. 2023. Using TikTok for public and youth mental health--A systematic review and content analysis. Clinical Child Psychology and Psychiatry, 28(1): 279--306
2023
-
[14]
P.; Utpala, R.; and Sharp, G
McLean, C. P.; Utpala, R.; and Sharp, G. 2022. The impacts of COVID-19 on eating disorders and disordered eating: A mixed studies systematic review and implications. Frontiers in Psychology, 13: 926709
2022
-
[15]
Montag, C.; Yang, H.; and Elhai, J. D. 2021. On the psychology of TikTok use: A first glimpse from empirical findings. Frontiers in public health, 9: 641673
2021
-
[16]
Monteleone, P. 2021. Eating disorders in the era of the COVID-19 pandemic: what have we learned?
2021
-
[17]
Polivy, J.; and Herman, C. P. 2002. Causes of eating disorders. Annual review of psychology, 53(1): 187--213
2002
-
[18]
Pruccoli, J.; De Rosa, M.; Chiasso, L.; Perrone, A.; and Parmeggiani, A. 2022. The use of TikTok among children and adolescents with Eating Disorders: experience in a third-level public Italian center during the SARS-CoV-2 pandemic. Italian Journal of Pediatrics, 48(1): 138
2022
-
[19]
Rando-Cueto, D.; De Las Heras-Pedrosa, C.; and Paniagua-Rojano, F. J. 2023. Health communication strategies via TikTok for the prevention of eating disorders. Systems, 11(6): 274
2023
-
[20]
Reed, J.; and Ort, K. 2022. The rise of eating disorders during COVID-19 and the impact on treatment. Journal of the American Academy of Child and Adolescent Psychiatry, 61(3): 349
2022
-
[21]
F.; Lombardo, C.; Cerolini, S.; Franko, D
Rodgers, R. F.; Lombardo, C.; Cerolini, S.; Franko, D. L.; Omori, M.; Fuller-Tyszkiewicz, M.; Linardon, J.; Courtet, P.; and Guillaume, S. 2020. The impact of the COVID-19 pandemic on eating disorder risk and symptoms. International Journal of Eating Disorders, 53(7): 1166--1170
2020
-
[22]
Schlegl, S.; Maier, J.; Meule, A.; and Voderholzer, U. 2020. Eating disorders in times of the COVID-19 pandemic—Results from an online survey of patients with anorexia nervosa. International Journal of Eating Disorders, 53(11): 1791--1800
2020
-
[23]
R.; Zuromski, K
Smith, A. R.; Zuromski, K. L.; and Dodd, D. R. 2018. Eating disorders and suicidality: what we know, what we don’t know, and suggestions for future research. Current opinion in psychology, 22: 63--67
2018
-
[24]
F.; et al
Sullivan, P. F.; et al. 1995. Mortality in anorexia nervosa. American Journal of Psychiatry, 152(7): 1073--1074
1995
-
[25]
R.; Luciano, S.; and Harrison, P
Taquet, M.; Geddes, J. R.; Luciano, S.; and Harrison, P. J. 2022. Incidence and outcomes of eating disorders during the COVID-19 pandemic. The British Journal of Psychiatry, 220(5): 262--264
2022
-
[26]
Tavolacci, M.-P.; Ladner, J.; and D \'e chelotte, P. 2021. Sharp increase in eating disorders among university students since the COVID-19 pandemic. Nutrients, 13(10): 3415
2021
-
[27]
L.; Deter, H.-C.; and Herzog, W
Zipfel, S.; L \"o we, B.; Reas, D. L.; Deter, H.-C.; and Herzog, W. 2000. Long-term prognosis in anorexia nervosa: lessons from a 21-year follow-up study. The Lancet, 355(9205): 721--722
2000
-
[28]
Zipfel, S.; Schmidt, U.; and Giel, K. E. 2022. The hidden burden of eating disorders during the COVID-19 pandemic. The Lancet. Psychiatry, 9(1): 9
2022
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.