{"id":"d9a6ba1a-3e85-444e-abbb-545fbe735f80","arxiv_id":"2509.15860","paper_version":2,"verdict":"REJECT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"PoliTok-DE is a proposed TikTok dataset with deletion tracking; the abstract's claimed scale and statistics are not supported by the paper's body.","lead":"The authors present PoliTok-DE, a multimodal dataset of TikTok posts from German elections with deletion tracking and a case study on intolerance and entertainment. The abstract claims over 930,000 posts from two elections, but the body describes a 195,000-post Saxony collection and contains inconsistent deletion figures.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The abstract's headline corpus (930k posts, 330k deletions, federal election rates) does not appear in the body, and the body's own deletion percentage is miscomputed; the central claim is currently unsubstantiated.","rationale":"The reader correctly flagged the Stage I/III assumptions and the annotation-reliability issues, but the most immediately falsifying problem is the document's internal inconsistency: the abstract's headline numbers (930k posts, 330k deletions, federal rates) are absent from the body, and the body's own deletion rate is miscalculated. No adjustment to the re-scraping schedule or API completeness assumption can salvage the central claim while these inconsistencies stand, because the dataset's value proposition is precisely its size and deletion statistics. I would still credit the authors for releasing post IDs, providing hydration code, and honestly acknowledging the annotation-reliability limitations; those are real positives. But the arithmetical error and the missing federal collection make the headline unsupported. The verdict remains REJECT. If the authors release a corrected dataset description that actually includes the federal corpus and recompute the deletion rates, a revised version could warrant a CONDITIONAL evaluation.","tokens_in":7766,"tokens_out":6741,"duration_ms":59796,"concrete_test":"Download the released Hugging Face dataset and verify the corpus partitions: (a) check whether any partition corresponds to the 2025 federal election with over 900,000 post IDs; (b) recompute the Saxony deletion rate as deleted IDs / total IDs using the provided code. If the federal partition is absent and/or the Saxony rate is 9.6% rather than 17.3%, the abstract's central claim is unsubstantiated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim, as stated in the abstract (p.1), is a two-election corpus with over 930,000 posts, over 330,000 deleted posts, 18.7% Saxony / 39.7% federal deletion rates, and a 13.0% platform-deletion rate said to be 14-19x above TikTok's reported rate. The full text never describes the federal-election collection: §2 and §3 document only a Saxony corpus of 195,373 posts (01.07.2024–30.11.2024) with 18,842 deleted. Even within the body, this deletion rate is arithmetically inconsistent: 18,842 / 195,373 = 9.6%, not the stated 17.3%. The 930k/330k/13.0% figures therefore cannot be derived from the reported methods, search queries, or results. Because the dataset contribution is precisely these numbers, the central claim is unsupported as written. Secondary concerns (e.g., the single re-scrape in Stage III and annotation α < 0.66) are acknowledged in the paper, but the abstract/body discrepancy and the internal arithmetic error are sufficient to invalidate the headline.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents PoliTok-DE, a multimodal TikTok dataset for the 2024 Saxony state election. The body describes a four-stage pipeline (research API, scraping, re-scraping to identify deletions, annotations), a corpus of 195,373 posts with 18,842 deletions, keyword distributions, and a case study of intolerance and entertainment in a subsample. The abstract supplied with the manuscript, however, advertises a two-election corpus of over 930,000 posts and over 330,000 deleted posts, with a 13.0% platform-deletion rate. These headline statistics do not appear in the full text; the body's own deletion percentage is also arithmetically inconsistent (18,842/195,373 ≈ 9.6%, not 17.3%). The paper also contains annotation-reliability problems (Krippendorff's α ≤ 0.55 for key constructs) and a codebook copy-paste error. If the Saxony-only dataset and pipeline are taken as the actual contribution, the paper is a useful data note, but the current abstract and text cannot both be right.","tokens_in":8087,"tokens_out":8851,"duration_ms":79220,"significance":"A well-documented multimodal dataset of political TikTok posts with deletion snapshots would be a valuable resource for computational social science. The release of post IDs and hydration code, the explicit treatment of API limitations in §2, and the detailed codebook in Appendix B (aside from the error) are strengths. The deletion snapshot is a useful feature, and the keyword prevalence findings (e.g., AfD overrepresentation in deleted content, Figure 2) are interesting. However, the current significance is undermined by the unsupported federal-election statistics and the arithmetic error. The annotation reliability is too low to support the case-study prevalence claims. The dataset's methodological contribution is potentially real, but the published claims must be corrected and substantiated before the paper can be evaluated.","major_comments":[{"comment":"The abstract claims a two-election corpus of over 930,000 posts, over 330,000 deletions, 18.7%/39.7% per-election deletion rates, and a 13.0% platform-deletion rate. Neither the federal-election collection nor any of the 930k/330k/13.0% figures appear in the full text; §2 and §3 describe only a Saxony corpus of 195,373 posts, and §8 concludes with the same Saxony-only numbers. The 13.0% figure and the 14–19x comparison are absent from the body. The central advertised contribution is therefore unsupported. The revised manuscript must either document the federal-election collection with full methodology, search queries, and statistics, or withdraw these claims and reframe the paper around the Saxony dataset.","section":"Abstract vs. §2–§3 and §8"},{"comment":"The deletion count is arithmetically inconsistent: 18,842/195,373 = 9.6%, not 17.3% as stated in §3 ('17.3% of the dataset'), in §8, and in the narrative of Stage III (§2). If the intended proportion is 17.3%, the deleted count should be about 33,800, not 18,842. This error affects every deletion-related statistic and comparison, including the claim that the deletion rate is 'surprising' relative to TikTok's reported rate. The numerator, denominator, or percentage must be corrected, and all dependent statements revised accordingly.","section":"§3 Dataset Statistics"},{"comment":"The inter-annotator agreement for the key constructs is below commonly accepted thresholds: Krippendorff's α = 0.48 for intolerance, 0.55 for hedonic entertainment, and 0.38 for eudaimonic entertainment. The authors acknowledge in the text that these do not meet the recommended α ≥ 0.66 (Krippendorff, 2018). Despite this, the results are presented as point estimates (20.5% intolerance, 62.9% hedonic) and are echoed in the abstract ('about one in five posts conveyed intolerance and a majority conveyed humor'). With these α values, the prevalence estimates are not reliable enough to support the case-study claims; the claims should be explicitly downgraded to exploratory, or the annotation instrument should be refined and the data re-coded before any substantive interpretation.","section":"§5.2 Results, Table 1"},{"comment":"The codebook for Intolerance reproduces the Politics section verbatim: the question reads 'Does the video refer to political topics (policy, politics or polity)?' with the same explanation and examples as B.1. If this is the codebook actually used by annotators, the intolerance annotations are not defined by the specified protocol; if it is a clerical error, the appendix must be corrected before the dataset is used or cited. As written, this undermines the reproducibility of the annotation procedure and the credibility of the intolerance results.","section":"Appendix B.3"},{"comment":"The deletion rate is estimated from a single re-scrape performed 10 days after the end of the collection window (§2 Stage III, §6). Additionally, daily API queries are run 96 hours after publication (§2 Stage I), and the paper acknowledges that content deleted before the API query 'never makes it to the research API.' Consequently, posts deleted within those 96 hours are systematically missing from the corpus, making the deletion count among collected posts a lower bound rather than an unbiased rate. The manuscript does not quantify this censoring or its impact on the headline deletion percentages. A sensitivity analysis or at least an explicit bounding statement is needed.","section":"§2 Stage I and §6 Limitations"}],"minor_comments":[{"comment":"The headline abstract and the abstract in the full text describe different corpora (930k two-election corpus vs. 195k Saxony-only corpus). Even after addressing the major data mismatch, the two versions must be harmonized.","section":"General"},{"comment":"The two panels ('Full' and 'Deleted') would benefit from explicit axis labels and a caption clarifying that percentages are computed within each subset. The current presentation is visually dense and hard to read.","section":"Figure 2"},{"comment":"The text reports that 'about 13% of the posts were annotated as conveying both hedonic and eudaimonic entertainment,' but Table 1 reports only marginal distributions. The joint distribution should be reported to support this statement.","section":"§5.2"},{"comment":"The phrase 'we collected the data for each day 96 hours (4 days) later' is ambiguous. Specify whether each day's query was run exactly 96 hours after publication date, after the end of the day, or at some other offset.","section":"§2 Stage I"},{"comment":"The citation 'TikTok [2025]' in §3 refers to an online transparency report; the reference list gives a URL but the citation format is inconsistent with the author-date style used elsewhere. Minor formatting issue.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"I debated between major_revision and reject. The headline abstract/body mismatch is severe: if the authors cannot substantiate the 930k/330k federal-election statistics, the paper's central contribution is unsupported and the manuscript should not be accepted in any form. I recommend major_revision because the Saxony-only component is a coherent dataset paper that could be salvaged by aligning the claims, correcting the arithmetic error (17.3% vs. 9.6%), fixing the codebook copy-paste error, and either improving annotation reliability or substantially downgrading the case-study claims. If the federal collection is not provided or the authors cannot produce the underlying data, the paper should be rejected."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The resource is real, but the paper has a serious presentation problem you should know before trusting any of the headline numbers. The arXiv abstract promises a two-election corpus of over 930,000 posts with 330,000 deletions and a 13.0% platform-deletion rate. The PDF body describes only a Saxony state-election corpus of 195,373 posts, and that body's own deletion percentage is miscomputed: 18,842 / 195,373 = 9.6%, not the stated 17.3%. The federal-election collection never appears in the methods, results, or appendix. That's not a minor inconsistency; the scale and deletion statistics are the paper's main selling point.\n\nWhat is genuinely useful: a 195k-post multimodal TikTok corpus around the 2024 Saxony election, with daily collection through the research API, full media scraping, a deletion re-check, and public post IDs plus hydration code. That's a solid contribution to political communication and platform-moderation research. The descriptive finding that AfD is overrepresented in deleted posts (70.2% vs 50.9%) is interesting. The case study is transparent: the authors report low Krippendorff's alpha values (0.48 for intolerance, 0.55/0.38 for entertainment) and discuss why multimodal implicit intolerance is hard to annotate. The limitations section honestly flags that the deletion check is a single 10-day-after snapshot.\n\nThe soft spots, though, are load-bearing. One is the abstract/body mismatch. Another is that even the smaller body's deletion rate is arithmetically impossible. And the annotation reliabilities fall below the tentative threshold, so the case study's quantitative claims are shaky. These are fixable, but as written the central claims are unsubstantiated. The authors also under-specify how the research API keyword queries and the 96-hour delay affect completeness; the related work they cite suggests the API misses posts, which they acknowledge but don't quantify.\n\nIf the authors reconcile the two abstracts, correct the arithmetic, and release reproducible code for the deletion check, the dataset paper could be worth publishing. I would not cite the 930k numbers in your own work until they appear in the body. For peer review: I'd send it out, but with a request for major revision, not desk-reject, because the underlying data collection and honest limitations suggest the authors know what they're doing. A serious referee can help them fix this.","headline":"A potentially valuable Saxony TikTok dataset, but the abstract and body disagree on scale and the body's deletion percentage is wrong; the headline numbers are currently unsubstantiated.","tokens_in":8498,"tokens_out":3794,"would_cite":false,"duration_ms":33299,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"German political TikTok posts are deleted at 14–19 times TikTok's reported platform-wide rate, according to a new multimodal dataset.","keywords":["TikTok","political communication","deleted content","multimodal dataset","German elections","Saxony 2024","federal election 2025","platform moderation"],"falsifier":"Replicate the collection with daily re-scrapes over at least 30 days instead of a single check at day 10, and cross-check the keyword-API query against a broader crawl of German election TikTok content. If the resulting platform-deletion rate drops below, say, 5%—or close to TikTok's reported 1%—the paper's headline rate of 13% is an artifact of the single-snapshot measurement; if it stays above 10%, the central claim holds.","tokens_in":7691,"feed_emoji":"🗳️","tokens_out":7389,"duration_ms":55763,"temperature":0.7,"pith_summary":"PoliTok-DE is a multimodal collection of German political TikTok posts—video, audio, images, and text—assembled from the TikTok research API and web scraping. Its central empirical finding is that election-related posts disappear from the platform at a much higher rate than TikTok's official figures suggest: about 17% of posts in the Saxony-election collection, roughly 40% in the extended federal-election collection, and a computed platform-deletion rate of 13.0%, which is 14–19 times the platform-wide rate TikTok reports. The paper argues that this deleted content is not a trivial tail but a substantial, analyzable slice of political communication, and it shows that deleted posts are disproportionately about the far-right AfD and often combine humor with intolerant messages. The authors release the post IDs and hydration code so other researchers can reproduce and extend the dataset.","feed_headline":"Election TikToks vanish at 14–19 times TikTok's official deletion rate","feed_subtitle":"Deleted German election TikToks are abundant, AfD-heavy, and often mix humor with intolerance.","key_machinery":"The load-bearing machinery is the four-stage collection pipeline: (I) keyword queries to TikTok's research API with a 96-hour delay to account for indexing lag; (II) direct web scraping to pull video, audio, images, and metadata; (III) a single re-scrape ten days after the collection window to mark posts as deleted; and (IV) human annotation of a sample for politics, Saxony-relevance, intolerance, and hedonic/eudaimonic entertainment. The deletion snapshot (Stage III) is what turns an ordinary scraped corpus into a dataset about platform removal, and the annotation codebook is what makes the deleted content interpretable.","core_discovery":"On its own terms, the paper's central claim is that political TikTok content is being removed from the platform at rates far exceeding both TikTok's official 1% deletion figure and its claim that 94% of removals happen within 24 hours, and that this deleted content is substantively important rather than noise. Using keyword queries via TikTok's research API plus web scraping to capture full media, the authors assembled a corpus of German election posts and then re-scraped the platform once, ten days after the collection window, to mark which posts had vanished. They report that 17.3% of the Saxony corpus (18,842 of 195,373 posts) was deleted; in the extended federal-election collection, the","pith_inferences":["If the deletion-rate finding generalizes beyond Germany, single-snapshot studies of TikTok politics may be biased in a predictable direction: toward mainstream content, since extremist or intolerant posts are removed first.","The 96-hour lag the authors used suggests TikTok's research API may miss posts that are posted and deleted within that window; the true deletion rate could be even higher than reported.","A natural next step is to track the same post IDs over a longer horizon (months) to distinguish quick platform removals from slow creator withdrawals, and to check whether deletion rates spike around elections.","The low inter-annotator agreement on 'intolerance' and 'eudaimonic entertainment' hints that multimodal implicit intolerance may resist reliable human labeling, which is a challenge for training automated detectors."],"forward_implications":["Researchers can study what kinds of political speech TikTok removes (or creators withdraw) by comparing available and deleted posts on content, user, and engagement attributes.","The multimodal format (video, audio, images, text) allows studying intolerance and humor as they interact across modalities, going beyond text-only or meme-image analyses.","The high deletion rate of election-related posts suggests that analyses relying only on currently available TikTok content may systematically underrepresent far-right and otherwise problematic speech.","The public release of post IDs with hydration code enables replication and extension to other elections or platforms.","The case study's annotated subset provides a benchmark for multimodal intolerance detection, though inter-annotator agreement is below conventional thresholds."],"fun_headline_variants":["German election TikToks deleted at 14–19x official rate","Deleted political TikToks: 1 in 5 has intolerance, majority humor","TikTok deletion of German election posts exceeds official rate 14–19x","PoliTok-DE: 330k+ political TikToks vanished, many with hate and humor","Election TikTok deletion rate 14–19x higher than TikTok admits"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The paper's deletion statistics rest on a single availability check ten days after the collection window and on keyword-API queries that may miss posts deleted within the 96-hour indexing delay or outside the keyword set; if those samples are unrepresentative, the computed deletion rates do not describe the true platform behavior.","fun_headline_variants_meta":{"raw":{"variants":["German election TikToks deleted at 14–19x official rate","Deleted political TikToks: 1 in 5 has intolerance, majority humor","TikTok deletion of German election posts exceeds official rate 14–19x","PoliTok-DE: 330k+ political TikToks vanished, many with hate and humor","Election TikTok deletion rate 14–19x higher than TikTok admits"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000153,"raw_usage":{"total_tokens":1049,"prompt_tokens":754,"completion_tokens":295,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":498,"completion_tokens_details":{"reasoning_tokens":188}},"tokens_in":498,"tokens_out":295,"duration_ms":3279,"temperature":1.0,"reasoning_tokens":188,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T16:03:49.737637+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Replicate the collection with daily re-scrapes over at least 30 days instead of a single check at day 10, and cross-check the keyword-API query against a broader crawl of German election TikTok content. If the resulting platform-deletion rate drops below, say, 5%—or close to TikTok's reported 1%—the paper's headline rate of 13% is an artifact of the single-snapshot measurement; if it stays above 10%, the central claim holds.","supporting_citations":[],"review_version":1}