Pith. sign in

REVIEW 3 major objections 5 minor 1 cited by

EDTok: A Dataset for Eating Disorder Content on TikTok

T0 review · 3 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read This paper presents EDTok, a curated dataset of 43,040 TikTok videos related to eating disorders, spanning January 2019 to June 2024, together with 577,071 comments and replies and a collection framework.

desk verdict Useful dataset, but the text-only Gemini filter keeps 'comprehensive' from being true; send to review with revisions. read the letter →

arxiv 2505.02250 v1 pith:JLJT2STG submitted 2025-05-04 cs.SI

classification cs.SI
keywords EDTokeatingdisordersTikToksocialmediadatasetcontentmoderationCOVID-19pandemictopicmodelingmultimodalanalysis
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper's aim is to give researchers a large, multimodal public dataset for studying how eating disorder content moves through TikTok. It reports collecting 43,040 videos posted between January 2019 and June 2024, along with their metadata, audio transcripts, 577,071 comments and replies, and the video files, using eating-disorder keywords and hashtags. The authors argue the dataset fills a gap left by text-only or self-report studies, because it captures both the visual and textual dimensions of eating disorder discourse and spans the COVID-19 pandemic. If the collection is as reliable as claimed, it would let researchers track content spread, measure engagement, evaluate moderation evasion such as misspelled hashtags like #edr3c0very, and test whether the pandemic changed the volume and tone of such content. This matters for informing content moderation policy and mental-health interventions on a platform dominated by adolescents.

What carries the argument

The load-bearing object is the EDTok dataset and its curation pipeline. The pipeline starts with a manually assembled list of eating-disorder keywords and hashtags, including deliberate misspellings such as 'edrec0very' used to evade moderation, queried through TikTok's Research API; the PykTok module downloads the video files; then a two-stage filter removes unrelated posts, first by dropping videos whose metadata matches irrelevant keywords, then by running a large language model prompt on each video description that labels the video as eating-disorder-related or not. The classifier was validated on 200 randomly sampled videos with reported 99% accuracy and then applied to the full collection, followed by a 300-video manual check of the filtered result. This mechanism carries the argument because the dataset's research value depends entirely on the filter separating eating-disorder content from the roughly 20% unrelated videos that the initial keyword query pulls in.

What would settle it

Have two independent human annotators review a random sample of 1,000 videos labeled relevant by the pipeline, drawn across all collection years, and compare their labels to the classifier's; if the human-classifier agreement falls well below 99%, especially on videos without explicit eating-disorder keywords, the filter does not generalize as claimed.

Watch

Extended reading notes

Core claim

The central discovery is the dataset itself: a curated corpus of 43,040 TikTok videos on eating-disorder topics, plus 577,071 comments and replies, collected from a query set of keywords and hashtags and filtered through a two-stage pipeline that first removes obviously unrelated posts and then uses an off-the-shelf large language model classifier on video descriptions to keep only eating-disorder-relevant content. The paper reports that a manual review of 200 videos found the classifier about 99% accurate, and that a later random check of 300 videos in the final set found all of them relevant. The dataset spans January 1, 2019 to June 28, 2024, and records engagement totals of about 80 million likes, 537 million views, 577 thousand comments, and 962 thousand shares from 10,897 unique users. In the authors' framing, the contribution is not a finding about eating disorders so much as an instrument: a reproducible collection method and a public resource of video IDs and metadata that can support research on content spread, user engagement, moderation, and the pandemic's influence on online eating-disorder discourse.

Load-bearing premise

The curation pipeline assumes that a single text classifier, written with a manually authored prompt and checked on only 200 videos, correctly identifies eating-disorder relevance for every one of the more than 56,000 descriptions it then labels.

Editorial extensions

If this is right

  • Researchers can trace monthly video volume and engagement before, during, and after the pandemic, testing whether COVID-19 increased eating-disorder content on TikTok.
  • The dataset's combination of video files, descriptions, and comments enables multimodal studies linking visual content, caption sentiment, and community responses.
  • The keyword list and filtering framework can be reused or extended to collect other sensitive health topics, and the video IDs permit replication without redistributing copyrighted videos.
  • Analysis of engagement and moderation-evading hashtags can inform whether TikTok's moderation efforts are curbing harmful content or inadvertently suppressing recovery-oriented posts.
  • The included comments and replies allow studying how audiences react emotionally to recovery narratives versus harmful content, as the paper's emotion analysis begins to do.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Inference: The dataset's dependence on a single untuned classifier means its validity could be checked by re-annotating a stratified sample across years, engagement levels, and hashtag families; a large disagreement with human labels would not overturn the corpus but would narrow the claims that can be drawn from it.
  • Inference: Because the query relies on known hashtags and manual curation, the dataset is likely biased toward recovery and awareness content and may under-represent communities using newer or more obscure jargon; researchers comparing prevalence over time should treat volume changes cautiously.
  • Inference: The same collection framework could be pointed at other health conditions or at platform policy changes to create comparable datasets, effectively turning the pipeline into a template for studying moderation dynamics on short-video platforms.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper presents EDTok, a curated dataset of 43,040 TikTok videos related to eating disorders, collected via a set of hashtags and keywords in Table 1, spanning January 2019 to June 2024, alongside 577,071 comments and replies. The curation pipeline (Section 'Data Filtering') first removes videos with manually identified irrelevant keywords, then applies a Google Gemini text classifier to each video description, validating on a random sample of 200 videos (reporting ~99% accuracy) and on 300 post-filter videos (all confirmed relevant). The paper then provides descriptive analyses of engagement statistics, temporal trends, topic modeling of video descriptions and comments, emotion analysis, and a multimodal analysis of a 500-video subsample. The central claim is that EDTok is a comprehensive, multimodal resource for studying eating disorder content on TikTok, enabling analyses of content spread, moderation, engagement, and pandemic-related trends.

Significance. If the dataset's curation is valid, EDTok would be a valuable public resource for the study of eating disorder discourse on TikTok, combining video files, metadata, comments, and temporal coverage across the COVID-19 pandemic. The paper also contributes a reproducible metadata collection framework built on the TikTok Research API, and it is transparent about several limitations. However, the dataset's central validity claim rests on a single classifier applied to text descriptions, with unmeasured recall and incomplete reproducibility details; if those issues are addressed, the dataset could support meaningful research in computational social science and digital health. The descriptive analyses (engagement, temporal, topic, emotion) are straightforward and mostly illustrative, but they do not independently validate the curation.

major comments (3)
  1. [Data Filtering] The curation pipeline applies the Gemini classifier to video descriptions only (Table 5), yet the abstract and Discussion claim a multimodal dataset and a 'comprehensive view' of eating disorder content. Because recall on the excluded set is never measured, videos whose eating-disorder relevance is carried by visuals or audio rather than by text (for example, posts captioned 'day 1' or with emoji-only descriptions) are silently discarded, biasing the final 43,040 videos toward textually self-describing posts. Please estimate the false-negative rate by sampling videos excluded by the classifier and annotating them for eating-disorder relevance, and report this estimate alongside the precision figures. Without a recall estimate, the temporal trend and moderation analyses in Figures 2 and 3 cannot be interpreted as describing the full population of eating-disorder videos on TikTok.
  2. [Data Filtering] Neither the Gemini model version nor the date of execution is provided, and the manual keyword/hashtag removal list that reduced the dataset from 56,472 to 43,040 videos is not disclosed. This makes the curation pipeline non-reproducible and non-auditable, which is a load-bearing limitation for a dataset paper. Please specify the exact model identifier and query date, and include the full exclusion keyword list in an appendix or repository file.
  3. [Data Filtering] The validation of the classifier is based on 200 hand-checked examples with a self-reported 99% accuracy and a 300-video post-filter check with all confirmed related, but no inter-annotator agreement, confidence intervals, or stratified sampling are reported. The prompt in Table 5 explicitly lists terms such as 'anorexia', 'dieting', and 'weight loss', so the 99% figure may not generalize to the hard cases (misspelled hashtags, oblique references, recovery jargon) that motivated the study. I recommend adding a second annotator for a subset of the validation samples, reporting Cohen's kappa, and providing stratified validation across time periods and hashtag families.
minor comments (5)
  1. [Figure 1] The dataset flowchart is central to understanding the pipeline, but the exact counts at each stage (e.g., how many videos were removed by the keyword step versus the Gemini step) are not given in the text; please add these numbers so readers can audit the 56,472-to-43,040 reduction.
  2. [Table 3] The word 'residental' is misspelled and should be 'residential'; the top-words column would also benefit from consistent punctuation (e.g., listing each word separated by commas).
  3. [Table 4] The 'Fighting Battle' topic lists 'fighting, battle and fighter', which reads as a phrase containing 'and' rather than a list of keywords; please use a consistent separator (e.g., 'fighting, battle, fighter').
  4. [Text Analysis] The paragraph describing the Challenges topic contains 'signficant' instead of 'significant'; please correct the typo.
  5. [Text Analysis] The BERTopic results are presented without any quantitative quality metrics (e.g., topic coherence or topic diversity); adding such measures would strengthen confidence in the interpretability of the topics.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the dataset is assembled from external TikTok data via stated keywords and an external classifier, and the analyses are descriptive rather than predictions derived from the selection criteria.

full rationale

The paper's contribution is a collected and filtered dataset, not a derived prediction. The collection query terms are stated in Table 1, the classifier prompt is reproduced in Table 5, and the pipeline is documented in Figure 1. The analyses (engagement totals, temporal counts, topic models, emotion scores) are descriptive summaries of the resulting corpus; none is derived from a model whose parameters were fit to the same data and then called a prediction. The temporal 'increase after March 2020' is an empirical count of videos captured by the API, not a model output equivalent to the collection criteria. The only self-citations (Bickham et al. 2024; Lerman et al. 2023) supply previously used hashtags that are themselves listed in Table 1, so the argument does not reduce to an unverified self-cited theorem. The acknowledged limitations (keyword-selection bias, untuned Gemini classifier, API constraints) concern recall and representativeness, not circularity: the selection criteria are inputs, and the paper's claims about the selected corpus follow from descriptive statistics rather than from the criteria by construction.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The dataset's validity rests on three domain assumptions: the TikTok API sample is representative, the Gemini-based text filter is accurate, and the keyword list spans eating-disorder content. None of these are machine-checked or externally benchmarked in the paper; the first two are only partially validated.

assumptions (3)
  • domain assumption TikTok Research API data is representative of TikTok content.
    The dataset's temporal and engagement analyses assume that API-returned videos and metadata are a valid sample of the platform; rate limits and token expiry may bias sampling.
  • domain assumption Google Gemini classification of video descriptions accurately identifies eating-disorder-related videos.
    The entire filtering step depends on the LLM's judgment under an author-written prompt; validated on 200 examples with self-reported 99 percent accuracy, with no model version pinned.
  • domain assumption The predefined keyword and hashtag set covers the eating-disorder content space.
    The authors acknowledge that the list, partly drawn from their own prior work, may fail to capture the full spectrum of eating-disorder content on TikTok.

how reviews work

0 comments
Cite this review

Pith. "Pith review of EDTok: A Dataset for Eating Disorder Content on TikTok." pith.science (2026). https://pith.science/paper/JLJT2STG

@misc{pith2026250502250,
  author       = {Pith},
  title        = {Pith review of: EDTok: A Dataset for Eating Disorder Content on TikTok},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/JLJT2STG}},
  note         = {Machine review of arXiv:2505.02250}
}
read the original abstract

Eating disorders, which include anorexia nervosa and bulimia nervosa, have been exacerbated by the COVID-19 pandemic, with increased diagnoses linked to heightened exposure to idealized body images online. TikTok, a platform with over a billion predominantly adolescent users, has become a key space where eating disorder content is shared, raising concerns about its impact on vulnerable populations. In response, we present a curated dataset of 43,040 TikTok videos, collected using keywords and hashtags related to eating disorders. Spanning from January 2019 to June 2024, this dataset, offers a comprehensive view of eating disorder-related content on TikTok. Our dataset has the potential to address significant research gaps, enabling analysis of content spread and moderation, user engagement, and the pandemic's influence on eating disorder trends. This work aims to inform strategies for mitigating risks associated with harmful content, contributing valuable insights to the study of digital health and social media's role in shaping mental health.

Figures

Figures reproduced from arXiv: 2505.02250 by the authors.

Figure 1
Figure 1. Dataset Flowchart Temporal Analysis In this section, we analyze the temporal trends in the dataset. These analyses provide insights into the changing volume of content related to eating disorders and its impact on user interaction over the dataset’s time span [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Total occurrences of videos per month. WHO: World Health Organization, NEDA: National Eating Disorder Aware [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. View and Like Counts of videos per month. WHO: World Health Organization, NEDA: National Eating Disorder [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Top 15 Most Frequent and Viewed Hashtags [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]
Figure 5
Figure 5. Figure 5: Comparative emotion analysis of video descriptions and comments and replies topics [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]
Figure 6
Figure 6. Figure 6: Word Cloud of Video Descriptions [PITH_FULL_IMAGE:figures/full_fig_p010_6.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. WhichTok? Comparing Three TikTok Data Acquisition Tools

    cs.SI 2026-08 conditional novelty 6.0 of 10

    Three TikTok data collection tools retrieve largely non-overlapping datasets for identical queries, especially for hashtag and keyword searches.

Reference graph

Works this paper leans on

28 extracted references · 25 canonical work pages · cited by 1 Pith paper

  1. [1]

    , " * write output.state after.block = add.period write newline

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts #0 'before.a...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    Alhuzali, H.; and Ananiadou, S. 2021. SpanEmo: Casting multi-label emotion classification as span-prediction. arXiv preprint arXiv:2101.10038

  4. [4]

    Bickham, C.; Kazemi-Nia, K.; Luceri, L.; Lerman, K.; and Ferrara, E. 2024. Hidden in Plain Sight: Exploring the Intersections of Mental Health, Eating Disorders, and Content Moderation on TikTok. In Workshop Proceedings of the 18th International AAAI Conference on Web and Social Media

  5. [5]

    Chochlakis, G.; Mahajan, G.; Baruah, S.; Burghardt, K.; Lerman, K.; and Narayanan, S. 2023. Leveraging label correlations in a multi-label setting: A case study in emotion. In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 1--5. IEEE

  6. [6]

    Grootendorst, M. 2022. BERTopic: Neural topic modeling with a class-based TF-IDF procedure. arXiv preprint arXiv:2203.05794

  7. [7]

    This is just how I cope

    Herrick, S. S.; Hallward, L.; and Duncan, L. R. 2021. “This is just how I cope”: An inductive thematic analysis of eating disorder recovery content created and shared on TikTok using\# EDrecovery. International journal of eating disorders, 54(4): 516--526

  8. [8]

    B.; Greenwood, D

    Herzog, D. B.; Greenwood, D. N.; Dorer, D. J.; Flores, A. T.; Ekeblad, E. R.; Richards, A.; Blais, M. A.; and Keller, M. B. 2000. Mortality in eating disorders: a descriptive study. International Journal of Eating Disorders, 28(1): 20--26

Show all 28 references
  1. [9]

    Iqbal, M. 2021. TikTok revenue and usage statistics (2021). Business of apps, 1(1)

  2. [10]

    K.; Patten, S

    J Devoe, D.; Han, A.; Anderson, A.; Katzman, D. K.; Patten, S. B.; Soumbasis, A.; Flanagan, J.; Paslakis, G.; Vyver, E.; Marcoux, G.; et al. 2023. The impact of the COVID-19 pandemic on eating disorders: A systematic review. International Journal of Eating Disorders, 56(1): 5--25

  3. [11]

    L.; Garc \' a, M

    Jordan, G. L.; Garc \' a, M. D.; Esteban, P. S.; Mart \' n, A. S.; S \'a nchez, P. M.; Cubo, M. J.; Cano, M. M.; Del Barrio, J. G.; and Ayesa-Arriola, R. 2021. Tiktok, a vehicle for Pro-Ana and Pro-Mia content boosted by the COVID-19 pandemic. European Psychiatry, 64(S1): S703--S703

  4. [12]

    Lerman, K.; Karnati, A.; Zhou, S.; Chen, S.; Kumar, S.; He, Z.; Yau, J.; and Horn, A. 2023. Radicalized by Thinness: Using a Model of Radicalization to Understand Pro-Anorexia Communities on Twitter. arXiv preprint arXiv:2305.11316

  5. [13]

    McCashin, D.; and Murphy, C. M. 2023. Using TikTok for public and youth mental health--A systematic review and content analysis. Clinical Child Psychology and Psychiatry, 28(1): 279--306

  6. [14]

    P.; Utpala, R.; and Sharp, G

    McLean, C. P.; Utpala, R.; and Sharp, G. 2022. The impacts of COVID-19 on eating disorders and disordered eating: A mixed studies systematic review and implications. Frontiers in Psychology, 13: 926709

  7. [15]

    Montag, C.; Yang, H.; and Elhai, J. D. 2021. On the psychology of TikTok use: A first glimpse from empirical findings. Frontiers in public health, 9: 641673

  8. [16]

    Monteleone, P. 2021. Eating disorders in the era of the COVID-19 pandemic: what have we learned?

  9. [17]

    Polivy, J.; and Herman, C. P. 2002. Causes of eating disorders. Annual review of psychology, 53(1): 187--213

  10. [18]

    Pruccoli, J.; De Rosa, M.; Chiasso, L.; Perrone, A.; and Parmeggiani, A. 2022. The use of TikTok among children and adolescents with Eating Disorders: experience in a third-level public Italian center during the SARS-CoV-2 pandemic. Italian Journal of Pediatrics, 48(1): 138

  11. [19]

    Rando-Cueto, D.; De Las Heras-Pedrosa, C.; and Paniagua-Rojano, F. J. 2023. Health communication strategies via TikTok for the prevention of eating disorders. Systems, 11(6): 274

  12. [20]

    Reed, J.; and Ort, K. 2022. The rise of eating disorders during COVID-19 and the impact on treatment. Journal of the American Academy of Child and Adolescent Psychiatry, 61(3): 349

  13. [21]

    F.; Lombardo, C.; Cerolini, S.; Franko, D

    Rodgers, R. F.; Lombardo, C.; Cerolini, S.; Franko, D. L.; Omori, M.; Fuller-Tyszkiewicz, M.; Linardon, J.; Courtet, P.; and Guillaume, S. 2020. The impact of the COVID-19 pandemic on eating disorder risk and symptoms. International Journal of Eating Disorders, 53(7): 1166--1170

  14. [22]

    Schlegl, S.; Maier, J.; Meule, A.; and Voderholzer, U. 2020. Eating disorders in times of the COVID-19 pandemic—Results from an online survey of patients with anorexia nervosa. International Journal of Eating Disorders, 53(11): 1791--1800

  15. [23]

    R.; Zuromski, K

    Smith, A. R.; Zuromski, K. L.; and Dodd, D. R. 2018. Eating disorders and suicidality: what we know, what we don’t know, and suggestions for future research. Current opinion in psychology, 22: 63--67

  16. [24]

    F.; et al

    Sullivan, P. F.; et al. 1995. Mortality in anorexia nervosa. American Journal of Psychiatry, 152(7): 1073--1074

  17. [25]

    R.; Luciano, S.; and Harrison, P

    Taquet, M.; Geddes, J. R.; Luciano, S.; and Harrison, P. J. 2022. Incidence and outcomes of eating disorders during the COVID-19 pandemic. The British Journal of Psychiatry, 220(5): 262--264

  18. [26]

    Tavolacci, M.-P.; Ladner, J.; and D \'e chelotte, P. 2021. Sharp increase in eating disorders among university students since the COVID-19 pandemic. Nutrients, 13(10): 3415

  19. [27]

    L.; Deter, H.-C.; and Herzog, W

    Zipfel, S.; L \"o we, B.; Reas, D. L.; Deter, H.-C.; and Herzog, W. 2000. Long-term prognosis in anorexia nervosa: lessons from a 21-year follow-up study. The Lancet, 355(9205): 721--722

  20. [28]

    Zipfel, S.; Schmidt, U.; and Giel, K. E. 2022. The hidden burden of eating disorders during the COVID-19 pandemic. The Lancet. Psychiatry, 9(1): 9

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.