{"id":"dcb1c3c6-9622-45b2-81c9-17e3c7ab4979","arxiv_id":"2412.05039","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A community-sourced rating dataset shows dark patterns are widespread across mobile games, with free-to-play and in-app-purchase titles significantly more likely to be flagged.","lead":"The paper analyzes 1,496 mobile games using player-submitted ratings from a community website, finding that manipulative 'dark patterns' appear in most games, including many labeled healthy. It then shows free-to-play games and those with in-app purchases are especially likely to be rated as dark, raising questions about the true cost of free gaming.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline '85,388 rated dark pattern instances' aggregates binary user votes, not distinct dark patterns: the taxonomy caps each category at 7-12 patterns per game, yet reported SDs (e.g., 97.66 for monetary) far exceed that cap, so the central prevalence metric conflates rater activity with…","rationale":"The reader's weakest assumption correctly identifies that the crowdsourced ratings may be inaccurate or biased. My concern is more specific and internal: the paper's headline metric conflates the number of positive user votes with the number of distinct dark patterns. This is demonstrable from the paper's own data: Table 1's standard deviations (e.g., SD 97.66 for monetary dark patterns) exceed the maximum possible number of distinct monetary pattern types (11-12) that any single game could contain, so the reported counts must aggregate votes across multiple users. Thus the '85,388 instances' statistic is not a measure of design prevalence but of rater activity, and the central claim's quantitative support is not interpretable as stated. This is a load-bearing flaw because the abstract, results, and conclusion all cite this figure. I located the relevant internal evidence in Section 3 (binary rating procedure), Table 1, and Appendix A (taxonomy list lengths). The paper honestly acknowledges an inability to verify ratings (Section 5.1) and the opaque 'dark'/'healthy' threshold (Section 5.2), but it does not acknowledge the vote-versus-instance conflation. The qualitative conclusion that even 'healthy'-labelled games receive dark pattern ratings may still hold, and the exploratory nature of the study means a revision correcting the unit of analysis could restore validity. I therefore agree with the reader's CONDITIONAL verdict and recommend no change in the verdict category, while adding this specific reanalysis as a required condition for acceptance.","tokens_in":13571,"tokens_out":11671,"duration_ms":118968,"concrete_test":"Re-open the raw JSON crawl from darkpatterns.games and, for each game, compute (a) the number of distinct dark pattern types with at least one positive vote within each category, and (b) the total number of positive votes per category. If (a) is bounded by the taxonomy list lengths (7 temporal, 11-12 monetary, 7 social, 7 psychological) while (b) is substantially larger, the 85,388 figure is a vote count. Then recompute the paper's key statistics using distinct-type counts rather than vote sums, including the 10.76% of games with no dark patterns and the category prevalence comparisons. If the share of games with at least one distinct dark pattern or the relative prevalence of categories changes materially, the paper's quantitative claims require revision.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central quantitative evidence is the 85,388 'instances' of dark patterns, but the underlying data are sums of binary user votes, not counts of distinct dark patterns. Section 3 states that users are asked 'in a binary way whether a dark pattern is present or not,' so each positive response is a vote, not a unique occurrence. Table 2 (Appendix A) lists at most 7 temporal, 11 monetary, 7 social, and 7 psychological pattern types per game. Yet Table 1 reports standard deviations per game of up to 69.67 (temporal, dark) and 97.66 (monetary, dark). A count of distinct pattern types cannot exceed the taxonomy's per-category bound of 7-12, so SDs of 23-98 are mathematically impossible unless the values are aggregates of multiple user ratings. Therefore, the abstract's '85,388 rated dark pattern instances' and the conclusion's 'over 85,000 instances' are vote counts, not instance counts. This conflates the number of engaged raters with the number of dark patterns deployed: a game that many users all flag for 'Grinding' contributes many votes without adding a new dark pattern. The paper's RQ asks 'how prevalent are types of dark patterns,' but the analysis never separates the number of distinct patterns from the number of votes, and no rater-level data (raters per game, inter-rater reliability) are provided. This is not an external validity threat but an internal unit-of-analysis problem, and it is not addressed in the limitations section (Section 6).","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents an exploratory quantitative analysis of user-generated dark pattern ratings from the website darkpattern.games, covering 1,496 mobile games after filtering from an initial crawl of 52,111 games. The authors categorize games as 'dark' or 'healthy' using the website's labels, compare the volume of ratings across four dark pattern categories, and test associations between these labels and game pricing, advertising, and in-app purchase metadata. The central claim is that dark patterns are widespread in mobile games, including games perceived as benign, supported by a headline figure of 85,388 rated dark pattern instances and by significant chi-square associations between dark game labels and free-to-play / freemium revenue models.","tokens_in":13917,"tokens_out":5912,"duration_ms":53484,"significance":"If the prevalence findings were valid, the paper would provide a useful quantitative complement to prior qualitative work on dark patterns in games, and it would strengthen the case for community-based monitoring of manipulative design. The authors are transparent about several limitations, including the inability to verify the provenance of the ratings and the possibility of underreporting. The use of non-parametric tests is appropriate for the reportedly non-normal data. However, the central metric conflates individual binary user votes with distinct dark pattern instances, which undermines the headline prevalence claim and requires either re-analysis or a substantial reframing of the paper's conclusions.","major_comments":[{"comment":"The headline '85,388 rated dark pattern instances' is not a count of distinct dark patterns. Users are asked 'in a binary way whether a dark pattern is present or not' (Section 3), so each positive response is a vote, and the same pattern can be voted for by many users. The taxonomy in Appendix A caps each category at 7 temporal, 11 monetary, 7 social, and 7 psychological pattern types, yet Table 1 reports per-game averages of 22.09 temporal patterns for dark games and standard deviations up to 97.66. Such values are mathematically impossible for counts of distinct pattern types, confirming that the aggregated numbers are sums of user votes. The abstract's '85,388 rated dark pattern instances' and the conclusion's 'over 85,000 instances' are therefore vote counts, not instance counts. The paper's RQ asks about the prevalence of dark pattern types, but the analysis never separates distinct patterns per game from rater activity. Please either re-analyze the data to count distinct patterns per game per category, or explicitly redefine all prevalence claims as 'user ratings of dark patterns' and adjust the abstract, results, and conclusion accordingly.","section":"§3, Table 1, Appendix A"},{"comment":"The Kruskal-Wallis tests in Section 4.1 compare dark pattern ratings between the 'dark' and 'healthy' groups, but those group labels are themselves derived from the website's processing of the same user rating data. The finding that 'dark' games contain significantly more dark patterns in each category is therefore partly built into the definition of the grouping variable. The text acknowledges this in passing ('These differences ... may not be surprising as the data is structured to distinguish between games containing more or less dark patterns'), but the inferential statistics are still presented as evidence. This comparison should be either removed or reframed as a manipulation check of the website's classification, not as an independent empirical result about the prevalence of dark patterns.","section":"§4.1"},{"comment":"The chi-square tests of independence in the revenue analysis report only chi-square statistics, degrees of freedom, and p-values. With N = 1,496 and highly unequal group sizes (e.g., 96.8% of 'dark' games vs. 53.0% of 'healthy' games are free-to-play), p < .001 alone does not convey the magnitude of the associations. For each test, please report an effect size such as Cramér's V or an odds ratio with a 95% confidence interval, and interpret the practical significance of the magnitude in the text.","section":"§4.2"},{"comment":"The data set is reduced from 52,111 listed games to 1,496 games by removing all games without any user ratings. This selection is unlikely to be neutral: games that receive ratings on a dark-pattern reporting website are probably a non-random subset, likely overrepresenting popular games or games that users suspect of manipulative designs. The abstract's claim that dark patterns are 'widespread in mobile games' should be qualified to 'widespread in the games for which users submitted at least one rating on darkpattern.games.' The discussion should explicitly address how this selection may inflate prevalence estimates and limit generalizability.","section":"§3.1, Abstract"}],"minor_comments":[{"comment":"There is a typo in the phrase 'damaging patters' near the beginning of Section 2.1; it should read 'damaging patterns'.","section":"§2.1"},{"comment":"The website name is written inconsistently as both 'darkpattern.games' and 'darkpatterns.games' (e.g., Section 1 vs. Section 3.1, and reference [11]). Please unify the spelling.","section":"Throughout"},{"comment":"The number '1.496' in Section 1 should be '1,496' for consistency with the rest of the paper.","section":"§1"},{"comment":"The text says the study 'revealed over 50k dark pattern instances' while the data reported in Table 1 and elsewhere is 85,388 ratings; please reconcile this figure.","section":"§5.2"},{"comment":"The bar chart labels for 'Price to Download' are ambiguous ('Free', 'Not Free', 'Yes', 'No', 'Unknown' are mixed); please add a legend or clear axis labels for each subplot.","section":"Figure 3"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a timely topic and the authors are appropriately cautious in some limitations, but the central measure is currently mislabeled and the group comparison in §4.1 is circular. If the raw data can be re-analyzed to count distinct patterns per game, that would strengthen the contribution; otherwise, the manuscript must be revised to speak about user ratings rather than instances. There is no data availability statement; given the reliance on a third-party website, the authors should provide the crawler code and aggregated data (with permission) to support reproducibility."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper's headline result — 85,388 dark pattern instances across 1,496 games — is not what it appears to be. The figure is a sum of binary user votes, not a count of distinct dark patterns. The taxonomy in Appendix A allows at most 12 patterns per category, yet the standard deviations in Table 1 go up to 97.66 for monetary patterns in dark games. A per-game count of pattern types can't have that spread. So the analysis conflates how many users bothered to rate a game with how many dark patterns the game actually contains.\n\nTo the authors' credit, this is the first large-scale crawl of darkpatterns.games, and it brings a quantitative angle to a literature that has mostly been qualitative. The revenue comparisons (free-to-play, ads, in-app purchases) are interesting and the authors are honest about the limits of their data, noting they can't verify ratings or the label threshold.\n\nBut the problems run deeper. The dark/healthy comparison in Section 4.1 is partly circular: the labels are derived from the same user ratings being compared. The authors acknowledge they don't know how the website assigns labels, yet they still interpret the differences as evidence about dark pattern prevalence. And since prior work shows users often fail to recognize dark patterns, an absence of ratings in 'healthy' games doesn't mean an absence of manipulative design.\n\nThe paper deserves a serious referee because the research question matters and the dataset is novel, but it needs major revision. The analysis should separate votes from distinct patterns, report rater counts per game, treat the dark/healthy labels as community judgments rather than ground truth, and include effect sizes. Without that, the central claim doesn't hold.","headline":"The paper's headline prevalence number counts user votes, not dark patterns, which undercuts the main claim; the dataset is new and the revenue analysis is suggestive, but the interpretation needs a major rework.","tokens_in":14415,"tokens_out":3811,"would_cite":false,"duration_ms":37836,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Dark patterns appear in 89% of rated mobile games, even ones labeled healthy.","keywords":["dark patterns","deceptive design","mobile games","video games","ethical design","user-generated data","in-app purchases"],"falsifier":"Take a random sample of, say, 50 games that the site labels 'healthy' or 'no dark patterns' and have trained coders audit the actual game interfaces for the same four categories. If the audits find dark patterns in a large majority of those games, the paper's prevalence figures and its 'healthy' group characterization would be unsupported, since the absence of a rating would not mean absence of a pattern.","tokens_in":13385,"feed_emoji":"🎮","tokens_out":7498,"duration_ms":68801,"temperature":0.7,"pith_summary":"This paper sets out to establish, with numbers, how common manipulative design is in mobile games. Working from 85,388 user-generated ratings covering 1,496 games on the community site darkpatterns.games, it finds that only about one game in ten has no reported dark pattern, and that games the community labels 'healthy' still contain thousands of rated instances. The paper also shows that free-to-play games, games with ads, and games with in-app purchases are significantly more likely to be rated 'dark' than paid games. If the findings hold, manipulative patterns are not an edge case in mobile gaming but a structural feature of the freemium revenue model, with real consequences for players who cannot afford to pay up front.","feed_headline":"Dark patterns found in 89% of rated mobile games","feed_subtitle":"Even 'healthy' games carry manipulative designs, led by free-to-play and in-app purchases.","key_machinery":"The analysis runs on the community rating corpus behind darkpatterns.games, which lets players mark each game for the presence or absence of specific dark patterns in four categories: temporal (grinding, playing by appointment), monetary (pay-to-skip, loot boxes, pay walls), social (pyramid schemes, friend spam), and psychological (endowed progress, variable rewards). The authors extract the 1,496 games that have at least one rating, then use non-parametric tests (Kruskal-Wallis for category scores, chi-square for revenue features) to compare 'dark' versus 'healthy' games. The 'dark'/'healthy' split is the website's own threshold over the binary ratings, so the mechanism is entirely user-generated labels, not expert coding.","core_discovery":"The central discovery is that dark patterns are pervasive across mobile games, including games routinely presented as benign. Users reported 85,388 instances of temporal, monetary, social, and psychological dark patterns in the 1,496 games examined, and only 161 of those games (10.76%) had no reported instances. Games labeled 'dark' contained statistically significantly more instances of every category than games labeled 'healthy,' and the revenue analysis shows that free-to-play distribution, advertising, and in-app purchases are all strongly associated with darker ratings. The paper takes this as quantitative support for existing dark-pattern frameworks and as evidence that 'healthy' is a graded label, not a guarantee.","pith_inferences":["If users systematically fail to recognize dark patterns, as the cited user studies suggest, the 89% figure is a lower bound; an expert audit of games with zero reports would likely find additional instances, especially in 'healthy' games.","The stronger link between revenue features (free-to-play, ads, in-app purchases) and 'dark' ratings than between any single pattern category suggests that monetization mechanics, rather than individual manipulative tricks, are the most recoverable risk signal for automated screening.","A testable extension would track the same corpus over time: if regulation or public pressure reduces, say, loot boxes, one would expect the mix of monetary and psychological ratings to shift rather than simply decline."],"forward_implications":["If the reported prevalence is even roughly right, most mobile game players will repeatedly encounter dark patterns: roughly nine in ten rated games have at least one, and 'dark' games average far more.","Free-to-play, ad-supported, and in-app-purchase models are the settings where dark patterns concentrate, which points at the revenue model itself as the main driver rather than a few bad-apple developers.","Since games labeled 'healthy' also contain substantial dark patterns, app-store labels and community ratings should be treated as relative indicators, not certifications of safety.","The quantitative pattern supports the existing qualitative taxonomy of temporal, monetary, social, and psychological dark patterns, giving designers and regulators a measurable target."],"supporting_citations":[{"why":"Supplies the game-specific dark pattern taxonomy (temporal, monetary, social) on which the rating site's categories and this analysis are built.","marker":"[48]"},{"why":"The community website darkpatterns.games: the entire data set of 85,388 ratings across 1,496 games comes from this source.","marker":"[11]"},{"why":"User study of 240 mobile apps showing 95% contain dark patterns and users struggle to detect them; the paper uses this to interpret its own prevalence and underreporting.","marker":"[13]"},{"why":"End-user study with 406 participants on dark pattern recognition; cited to support the assumption that community raters will also miss instances.","marker":"[4]"},{"why":"Study of dark pattern recognition in social media; cited for the finding that even users who can statistically distinguish dark patterns are not equipped to avoid them, supporting the lower-bound interpretation.","marker":"[36]"},{"why":"Large-scale crawl of 11K shopping websites; the model for quantitative dark pattern analysis that this paper extends to mobile games.","marker":"[33]"}],"fun_headline_variants":["Even 'healthy' mobile games are full of dark patterns","Dark patterns found in 89% of games, including 'healthy' ones","Study: Mobile games are drenched in dark patterns, even benign ones","Dark patterns in mobile games: 'healthy' is just a label","89% of mobile games use dark patterns, and 'healthy' isn't safe"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument rests on the assumption that the community ratings on darkpatterns.games faithfully record which dark patterns are actually present in each game, even though the paper cannot verify the ratings' origin and prior work indicates users miss many dark patterns.","fun_headline_variants_meta":{"raw":{"variants":["Even 'healthy' mobile games are full of dark patterns","Dark patterns found in 89% of games, including 'healthy' ones","Study: Mobile games are drenched in dark patterns, even benign ones","Dark patterns in mobile games: 'healthy' is just a label","89% of mobile games use dark patterns, and 'healthy' isn't safe"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000235,"raw_usage":{"total_tokens":1434,"prompt_tokens":813,"completion_tokens":621,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":429,"completion_tokens_details":{"reasoning_tokens":526}},"tokens_in":429,"tokens_out":621,"duration_ms":7067,"temperature":1.0,"reasoning_tokens":526,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T20:57:25.192266+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of, say, 50 games that the site labels 'healthy' or 'no dark patterns' and have trained coders audit the actual game interfaces for the same four categories. If the audits find dark patterns in a large majority of those games, the paper's prevalence figures and its 'healthy' group characterization would be unsupported, since the absence of a rating would not mean absence of a pattern.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the game-specific dark pattern taxonomy (temporal, monetary, social) on which the rating site's categories and this analysis are built."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The community website darkpatterns.games: the entire data set of 85,388 ratings across 1,496 games comes from this source."},{"cited_title":"Dark Patterns at Scale: Findings from a Crawl of 11K Shopping Websites","cited_arxiv_id":"1907.07032","evidence_quote":"Large-scale crawl of 11K shopping websites; the model for quantitative dark pattern analysis that this paper extends to mobile games."}],"review_version":1}