REVIEW 3 major objections 5 minor
Filtering Offensive Content Changes Its Visibility but Not User Behavior: Two Randomized Controlled Trials with 200,000 Users on Nextdoor
T0 review · 3 major / 5 minor · reviewed 2026-08-01 · deepseek-v4-flash
Pith's one-line read This paper reports two randomized controlled trials on Nextdoor showing that filtering offensive content reduced its visibility by up to 95% while leaving platform visitation, content consumption, and content production unchanged.
desk verdict Two huge, clean field RCTs show opaque content filtering changes visibility but not behavior—worth engaging, though the abstract oversells one null and the Study 2 manipulation check is tighter than the headline claims. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The argument is carried by a paired design: each participant is enrolled upon an eligible platform visit, randomized to filtered or visible conditions, and followed 30 or 60 days before and after enrollment; a two-way mixed ANOVA with condition×time interaction isolates the intervention effect while absorbing the enrollment-day activity spike that both groups share. Study 2's manipulation is powered by proactive classification, with content scored by the Perspective API's 'Toxicity' attribute at a threshold of 0.70 and filtered from the newsfeed before accruing views. The escalation from a 12% report-triggered reduction to a 95% proactive reduction is the load-bearing axis that turns two nul
What would settle it
A falsifying observation would come from re-running Study 2's design with the outcome scored by an independent toxicity detector instead of the one used for filtering, or with an added condition that notifies authors when their content is filtered: if offensive-content creation drops in either variant, the nulls are specific to this implementation rather than to visibility reduction generally. Concretely, compare new-offensive-post counts in (a) opaque filtering, (b) filtering plus author notice, and (c) no filtering; a significant drop in (b) but not (a) would settle the paper's claim.
Extended reading notes
Core claim
The central claim, on the paper's own terms, is that filtering offensive content reliably achieves its direct target—reduced visibility—and has no detectable effect on any measured downstream behavior. Study 1 filtered comments in post threads after member reports, reducing views of those comments by 12%. Study 2 scored posts and comments at creation with Google Jigsaw's Perspective API and filtered content with Toxicity ≥ 0.70 from the newsfeed, suppressing related notifications, and reduced offensive-post views by 95%. Across the two trials, the time×condition interaction was null for platform sessions, content consumption, content production, and offensive-content creation, with Bayesian
Load-bearing premise
The load-bearing premise is that the operational definitions of 'offensive' (member reports in Study 1, Perspective Toxicity ≥ 0.70 in Study 2) capture what users actually experience as offensive and that the filter—not some unmeasured part of the treatment such as altered notifications—is the only difference between conditions; the paper itself notes in Section 6.1 that its measures capture behavior rather than attitudes and that both interventions were opaque by design.
Editorial extensions
If this is right
- Filtering can reduce the visibility of offensive content by an order of magnitude without measurably reducing platform engagement, so platforms can treat it as a low-risk visibility tool.
- The behavioral nulls replicate when manipulation strength rises from 12% to 95%, ruling out weak treatment as the explanation for the first study's nulls.
- Opaque visibility reduction does not reduce future creation of offensive content, in contrast to removal-with-notice effects reported in prior moderation research.
- Combining filtering with author-facing feedback—such as a notification that content was filtered—is the paper's suggested next step for turning visibility reduction into behavior change.
- Regulatory regimes requiring notice of demotion (e.g., the EU Digital Services Act) may push platforms toward the transparent variants this paper identifies as untested.
Reading between the lines
- Editorial inference: if mere exposure to offensive content were a major driver of users' own offensive posting, a 95% reduction in such exposure should have moved production; its absence hints that authors' behavior is shaped more by awareness of moderation or by habits formed elsewhere than by what appears in the feed.
- Editorial inference: the 95% manipulation check is partly mechanical because eligibility and outcome share the same Perspective threshold; a replication that scores outcomes with an independent detector would test whether the behavioral nulls survive classifier error.
- Editorial inference: Nextdoor's hyperlocal, real-name setting may dampen effects that would appear in anonymous, interest-based communities, so the generalizability of the nulls is an open empirical question.
- Editorial inference: the one significant secondary effect—lower average toxicity of comments viewed—suggests feed quality moved even though volume did not; future trials measuring exposure quality rather than volume may be more sensitive to filtering's effects.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports two large randomized controlled trials on Nextdoor, each with 100,000 users, testing interventions that reduce the visibility of offensive content without removing it or notifying authors. Study 1 (2022) used a report-triggered filter on comments in post threads, achieving a 12% reduction in views of offensive comments and finding no significant effects on eleven other platform-behavior measures. Study 2 (2023–2024) proactively scored posts and comments at creation with the Perspective API and filtered threshold-scored content from the newsfeed, achieving a 95% reduction in views of offensive posts, with no significant effects on most of thirteen further measures; the one exception was a small significant reduction in the average toxicity of comments viewed. The authors argue that the convergent null results across different content types, classifiers, and countries show that opaque visibility-reduction filters do not change downstream user behavior, while also demonstrating that such filters do not suppress engagement.
Significance. If the results hold, this is one of the first large-scale field experimental evaluations of a ubiquitous but understudied moderation practice. The design has important strengths: preregistration-like clarity in design comparison, broad outcome batteries, manipulation checks showing the filters operated as intended, and complementary frequentist and Bayesian analyses. The two trials are complementary in mechanism and manipulation strength, and the paper's framing as an integrated pair is a useful contribution. The central null result—that even a near-total reduction in views of classifier-defined offensive content did not move most engagement, consumption, or production outcomes—is a valuable addition to the evidence base on content moderation. The main weaknesses are internal inconsistencies in the abstract, the partly mechanical manipulation check in Study 2, and missing tables that prevent full verification of the reported analyses.
major comments (3)
- [Abstract and §5.2] The abstract claims 'across thirteen further measures we again found no significant effects' for Study 2, but §5.2 reports a significant time×condition interaction for average toxicity of comments viewed (F(1,99998)=7.779, p=.005, η²=4.07e−6). While the effect size is tiny, the statement as written is factually incorrect and the paper's cross-study claim of 'same behavioral nulls' is overstated. This inconsistency must be fixed, and the abstract/conclusions should acknowledge the one exception or justify why it is treated as negligible.
- [§5.1–5.2, Table 6] The Study 2 manipulation check is partly mechanical: offensive content is defined as Perspective Toxicity≥0.70 in §5.1, the filter removes exactly that content, and the manipulation check in §5.2 counts views of the same threshold-defined content. The 95% reduction therefore verifies that the filter ran, but it does not by itself establish that users experienced 95% less subjectively offensive content. If the classifier misses or mislabels content that users actually find offensive—e.g., content just below 0.70—the behavioral nulls may reflect the specific classifier and threshold rather than a general property of visibility-only moderation. The paper should either temper the generalization to 'offensive content' or provide evidence (e.g., validation of the threshold against user perceptions, or secondary analyses using alternative thresholds/definitions) that the filtered set correspond
- [Tables 5–7] Tables 5, 6, and 7 are not included in the manuscript; they are placeholders reading '(Carried over from the second manuscript.)' These tables contain the pre/post means, ANOVA F/p values for all fifteen Study 2 measures, and Bayes factors. Without them, the reader cannot verify the central Study 2 results. The tables must be inserted before the manuscript can be properly evaluated.
minor comments (5)
- [§5.2] The sentence reporting the significant interaction on average toxicity of comments viewed should be carried through to the abstract, the general discussion, and the limitations section; currently the abstract and the conclusion present the results as entirely null.
- [§4.1 and §5.1] The eligibility criteria for Study 1 (reported by at least one member for a 'hurtful or harmful' reason and scored above a Nextdoor model threshold) and Study 2 (Perspective Toxicity≥0.70) are presented as cutoffs without documenting how these thresholds were chosen. A sentence on the rationale or sensitivity of results to the specific threshold would strengthen the interpretation.
- [§3.2] The sensitivity analysis is described as indicating that effects of d=.02 can be detected, but no formal justification or simulation details are given. A brief note on the assumed variance structure would help the reader judge the claim.
- [Various tables] Some table headers use inconsistent capitalization and spacing (e.g., 'Manip. Check', 'Consump.', 'Prod.') while the text spells out the full names. Also, the effect sizes (η²) are omitted from Table 3 and only partially reported in text; the authors say they are in the supplement, but including them in the main tables would be more transparent.
- [§6.1] The limitations section does not discuss the validity of the Perspective API threshold as an operationalization of 'offensive' for the study population. Given that the manipulation check is defined by the same classifier, a few sentences on this limitation would be appropriate.
Circularity Check
No significant circularity; the RCT design measures behavioral outcomes independently of the filter definition.
full rationale
This paper is an empirical randomized controlled trial, not a derivation, so the circularity burden is low. The only near-tautological element is the manipulation check: Study 2 defines filter eligibility by Perspective API Toxicity ≥ 0.70 (Section 5.1) and then measures views of threshold-eligible offensive posts, so the 95% drop partly verifies that the filter removed the content it was designed to remove. That is a fidelity check on the intervention, not a prediction from a fitted parameter, and it is not used as evidence for the behavioral nulls. All thirteen downstream outcomes — sessions, consumption, production, and offensive-content creation — are measured from platform records independently of the filtering rule. The creation-outcome threshold (≥0.60) also differs from the filter threshold, further separating the main outcomes from the intervention definition. The paper contains self-citations (e.g., the note that it supersedes two earlier preprints by the same authors), but these are disclosure statements or contextual related-work citations, not load-bearing justifications of the experimental results. There is no fitted parameter renamed as a prediction, no imported uniqueness theorem, and no ansatz smuggled in by citation. Potential construct-validity concerns — such as whether the Perspective threshold captures user-perceived offensiveness or whether the opaque treatment includes additional components — are threats to generalizability, not circularity.
Assumptions & free parameters
free parameters (3)
- Perspective Toxicity filter threshold (Study 2) =
0.70
- Offensive-created-content threshold (Study 2) =
0.60
- Study 1 removal-prediction model threshold =
not specified
assumptions (4)
- domain assumption Perspective API Toxicity score is a valid operationalization of offensive content for both the intervention and the outcomes.
- domain assumption No interference between treatment and control users (SUTVA) within neighborhoods.
- domain assumption Two-way mixed ANOVA inference is valid for highly skewed count outcomes at N=100,000.
- domain assumption Pre/post observation windows (30 and 60 days) capture the behavioral effects of interest.
Cite this review
Pith. "Pith review of Filtering Offensive Content Changes Its Visibility but Not User Behavior: Two Randomized Controlled Trials with 200,000 Users on Nextdoor." pith.science (2026). https://pith.science/paper/AQCDB7GD
@misc{pith2026260721853,
author = {Pith},
title = {Pith review of: Filtering Offensive Content Changes Its Visibility but Not User Behavior: Two Randomized Controlled Trials with 200,000 Users on Nextdoor},
year = {2026},
howpublished = {\url{https://pith.science/paper/AQCDB7GD}},
note = {Machine review of arXiv:2607.21853}
}
read the original abstract
We investigate the effectiveness of interventions that reduce the visibility of offensive content on the local social platform Nextdoor. Content filtering -- hiding or downranking offensive content that brushes against a platform's rules without clearly breaking them -- is deployed across virtually every major platform, yet almost no field evidence exists on whether it changes user behavior. We report two large-scale randomized controlled trials, each involving 100,000 users. Study 1 (2022) tested a report-triggered filter applied to comments in post threads and produced a modest 12% reduction in views of offensive comments; across eleven further measures of platform behavior we found no significant effects. Study 2 (2023-2024) remedied Study 1's central limitation -- a weak manipulation driven by slow, report-based eligibility -- by proactively scoring posts and comments at creation with Google Jigsaw's Perspective API and filtering them from the newsfeed. This produced a near-complete (95%) reduction in views of offensive posts, yet across thirteen further measures we again found no significant effects. Across two independent trials spanning different content types, filtering mechanisms, classifiers, and countries -- and despite manipulation strength rising from 12% to 95% -- filtering reliably reduced the visibility of offensive content without altering platform visitation, content consumption, or content production. These convergent null results provide rare field evidence on a ubiquitous intervention and underscore the complexity of effectively moderating online platforms.
Figures
Reviewed August 1, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.