Pith. sign in

REVIEW 3 major objections 5 minor 16 references

Ingroup bias is prevalent in user reports of hate and abuse online

T0 review · 3 major / 5 minor · reviewed 2026-08-04 · deepseek-v4-flash

Pith's one-line read People flag abuse aimed at their own group 17–63% more often than the same abuse aimed at the other side.

desk verdict The central finding is real and worth publishing, but the abstract's 17–63% range doesn't match the reported numbers, and the missing manipulation check leaves a small but real gap. read the letter →

arxiv 2510.04748 v3 pith:OMDUSJ6Z submitted 2025-10-06 cs.CY cs.HC

classification cs.CYcs.HC
keywords ingroupbiasflagginghatespeechonlineabusecontentmoderationuserreportingpoliticalpolarizationsafety
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper asks whether online 'flag' buttons produce an unbiased report of harmful content or whether group identity distorts what users choose to report. Across five pre-registered online experiments in the US, participants saw mock social-media comments aimed at either their own group or an opposing group and flagged roughly half of abusive comments overall. In every tested context—political affiliation, vaccination stance, climate-change belief, and abortion rights—abuse aimed at the ingroup was flagged more often than identical abuse aimed at the outgroup, by 17% to 63%. This matters because platforms and regulators increasingly rely on user flags as a crowdsourced moderation signal; if flags systematically under-report outgroup-directed abuse, they cannot be treated as a simple severity measure.

What carries the argument

The core instrument is a mock flagging task: each participant reads a neutral news excerpt and then 48 randomized comments (24 abusive, 24 mildly critical) aimed at a named group, and clicks a flag icon to report any they would normally report. The design is pre-registered; equal numbers of participants come from both sides of each issue, and the same comment wording is directed at the ingroup for half of participants and the outgroup for the other half, isolating the effect of target identity. Study 4b presents both targets in a single feed, adding a within-subjects test of whether seeing both sides' abuse reduces bias.

What would settle it

Re-run Study 1 with a manipulation check asking each participant to identify the target group of each comment immediately after flagging. If the ingroup bias vanishes in those who fail the check, the effect is tied to attention to target; a real-world test could use platform logs of who flags which content and compare against the group identity of the comment's target.

Watch

Extended reading notes

Core claim

On the paper's own terms: user flagging is a real but biased signal. In a mock social media feed, participants flagged about half of abusive comments and less than 5% of mild criticism, but they consistently flagged more when the comment attacked their own group. The effect held across all four intergroup contexts and ranged from a 17% to a 63% increase in flagging for ingroup-directed abuse. Seeing both groups abused in the same feed reduced but did not eliminate the bias. The authors conclude that flags cannot be treated as a neutral severity measure; they are shaped by the reporter's group membership.

Load-bearing premise

Participants' subjective group membership is self-reported and the target group is not checked after the task, so the whole ingroup-vs-outgroup comparison rests on respondents both encoding who the comment attacks and feeling the declared affiliation as their own.

Editorial extensions

If this is right

  • When users are explicitly invited to flag, reporting rates are high: roughly half of abusive comments are flagged, so low engagement may be partly a design problem rather than user indifference.
  • Outgroup-directed abuse is systematically under-flagged, meaning moderation queues built from user reports will under-represent attacks on outgroups relative to ingroups.
  • The bias is not limited to hate speech: mild criticism of one's own group is more likely to be misclassified as reportable than the same criticism of an outgroup.
  • Simultaneously showing abuse against both groups weakens the bias, suggesting that design choices about what users see can partially correct it.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • On a real platform, this bias would compound the vulnerability of marginalized groups: outgroup-directed abuse is exactly the category users are least motivated to report, so the people most targeted also get the least crowdsourced protection.
  • The mock-feed setting with explicit instructions and attention checks likely yields higher overall flagging than natural browsing; in a noisy real feed, the relative weight of ingroup bias could be even larger or could dilute as users skim.
  • A direct design implication the authors do not draw: platforms could experimentally test whether showing users abuse aimed at both sides in a single view—as in Study 4b—reduces bias in live flagging data.
  • Because no manipulation check confirms that participants perceived the target group, this bias estimate may be an upper bound for attentive users; skimmers who do not encode the target might show no such bias, which itself would change what the finding means for moderation.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper reports five pre-registered, US-based online experiments (total N ≈ 1,581 after exclusions) testing whether users' flagging of online hate/abuse is biased by group identity. Using mock social media feeds, participants saw abusive and mildly critical comments directed at Republicans/Democrats, pro-/anti-vaccination groups, climate-change believers/non-believers, and pro-/anti-abortion-rights groups. In all studies, participants flagged a majority of abusive comments (roughly 50–65%) and very few critical comments. The key claim is a main effect of target: comments directed at the ingroup are flagged significantly more than comments directed at the outgroup, with the effect replicated in all four social contexts and in both between-subjects (Studies 1–4a) and within-subjects (Study 4b) designs. The authors argue that user flags are therefore not a pure severity signal but are systematically skewed by social identity, with implications for platform moderation and intervention design.

Significance. If the central claim holds, the paper makes an important empirical contribution to the emerging literature on user flagging and online content moderation. Its strengths are substantial: all five studies are pre-registered, adequately powered, use balanced samples recruited on both sides of each issue, and combine a pre-registered ANOVA with a preregistered or complementary GLMM. The paradigm—using a mock feed with matched abusive/critical items—is a useful tool for measuring reporting bias experimentally. The paper also reports open data and materials at an OSF link. However, the quantitative headline in the abstract (a 17%–63% increase in flagging ingroup-directed abuse) is not supported by the reported proportions or odds ratios. This is a load-bearing error because it materially overstates the effect size and robustness that the data actually demonstrate. The qualitative pattern of ingroup bias is nonetheless consistent across all studies, which is the core message of the paper.

major comments (3)
  1. [Abstract] The abstract states that 'participants were between 17% and 63% more likely to flag abuse directed at the ingroup than at the outgroup.' This range is not derivable from any statistics reported in the manuscript. For abusive comments, the relative increases are: Study 1, (50.5−43.8)/43.8 = 15.3%; Study 2, (55.3−49.4)/49.4 = 11.9%; Study 3, (56.5−48.6)/48.6 = 16.3%; Study 4a, (65.4−53.7)/53.7 = 21.8%; Study 4b, (60.4−56.6)/56.6 = 6.7%. None of these reach 63%, and three fall below 17%. If the intended metric is the GLMM odds ratios (1.9, 2.73, 3.25, 2.66, 1.51), those correspond to 90%–225% increases in odds, not 17–63%. Please either correct the abstract to the actual range (≈7%–22% relative increase) or specify exactly how the 17–63% figure was computed. This is a central issue because the abstract is the most visible statement of the paper's findings and the quoted range overstates the
  2. [Study 1 Methods, pp. 9–14] No manipulation check is reported for the target-group manipulation. The design relies on participants perceiving whether each comment is directed at the ingroup or outgroup, but the only checks mentioned are two attention checks. If participants did not encode the target, the observed target main effects could be attenuated or driven by item-level differences despite the matching of items. Given that the central claim is that flagging is biased by group identity, the authors should report whether participants correctly identified the target group (e.g., from a recall check in the OSF materials) or acknowledge this as a limitation. Without such evidence, the internal validity of the key independent variable is not fully established.
  3. [General Discussion, pp. 34–36] The General Discussion states that 'in all four social contexts that we examined, we found a robust ingroup bias effect in flagging.' However, the effect is asymmetrical in the political context: Study 1 shows a significant ingroup bias only among Republicans; Democrats flagged ingroup- and outgroup-directed abuse at similar rates (M_IG = 46.1%, M_OG = 45.6%, p = .901). The same asymmetry appears in Study 4b for abortion stance (pro-abortion participants show bias for abuse but not criticism; anti-abortion participants show bias for criticism but not abuse). The phrase 'robust ingroup bias effect' as a blanket summary is therefore too strong. Please qualify the claim to reflect the group-dependent nature of the effect, or provide a meta-analytic justification for treating the overall effect as robust.
minor comments (5)
  1. [Title / Metadata] The submission title as provided is 'Ingroup bias is prevalent in user reports of hate and abuse online,' but the full-text manuscript title is 'Social bias is prevalent in user reports of hate and abuse online.' Please ensure the metadata and full text are consistent.
  2. [Study 4b, p. 29] The phrase 'much online conversation may polarised in this way' contains a grammatical error; should be 'may be polarized.' Check for similar typos throughout (e.g., Table 1: 'supprters' and 'thieving').
  3. [General Discussion, p. 36] The claim that exposing participants to both ingroup- and outgroup-directed comments 'narrowed the social bias effect' is a post hoc comparison between Studies 4a and 4b. This cross-experiment comparison is not formally tested; the effect sizes (ηp² = .06 vs .054) are not directly comparable without an inferential statistic. Please present this as a tentative observation, not a finding.
  4. [Abstract] The abstract says 'approximately half of the abusive comments in each study reported.' In Studies 4a and 4b, overall abuse flagging was 59.6% and 58.5%, respectively. 'Approximately half' is acceptable, but 'approximately 50–60%' would be more precise.
  5. [Methods/Results] For Study 2, the footnote clarifies that the GLMM was the preregistered primary analysis but the ANOVA is reported first. For consistency, consider making the preregistration status of the GLMM explicit in all study Methods sections, not just as a footnote, so readers know which analysis corresponds to the pre-registration.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the central claim is an empirical behavioral finding, not a derivation from its own inputs.

full rationale

This paper reports five pre-registered experiments measuring flagging behaviour. There is no mathematical derivation chain in which a predicted quantity is definitionally identical to a fitted input: the ingroup/outgroup manipulation is operationalized from participants' own self-reported affiliation (e.g., 'the Target was an ingroup or an outgroup depending on participants’ own political affiliation'), and the outcome (flagging rates) is independently measured behaviour. No parameter is fitted to the target effect and then called a prediction; the target main effects are simply observed differences between conditions. Self-citations (e.g., Enock et al., 2020, 2025; Bright et al., 2024) appear only as background context or as prior survey evidence, not as load-bearing arguments that force the results. The abstract's '17% and 63%' range appears not to follow from the reported proportions, but that is a numerical reporting/consistency issue, not a circular reduction of the conclusion to its premises. Accordingly, no specific circular step can be identified by quoting the paper's own equations or construction.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

No free parameters or invented entities are introduced; the claim rests on the validity of the mock-feed operationalization and the ingroup/outgroup manipulation.

assumptions (4)
  • domain assumption Self-reported group membership (e.g., Democrat/Republican, pro/anti-vaccination, climate believer/non-believer, pro/anti-abortion) corresponds to a psychologically meaningful ingroup in the experimental context.
    The entire ingroup/outgroup coding depends on this; no identity-strength or manipulation check is reported (Methods, p.14).
  • domain assumption Participants perceived the target group of each mock comment as intended.
    No manipulation check is reported; if participants misread target labels, target effects could reflect item content rather than group bias.
  • domain assumption Flagging behavior in the mock feed with explicit instructions reflects real-world online flagging.
    The authors themselves note this limitation (p.37): 'Results are obtained from a mock social media feed... real-world online experiences and behaviours are likely to be different.'
  • domain assumption The researcher classification of comments as abusive vs critical is valid; abusive items are genuinely hateful and critical items are not.
    Abusive items were adapted from Burke-Moore et al. (2025), but no norming or validation data are provided; the low flag rate for criticism (<5%) supports it indirectly.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Ingroup bias is prevalent in user reports of hate and abuse online." pith.science (2026). https://pith.science/paper/OMDUSJ6Z

@misc{pith2026251004748,
  author       = {Pith},
  title        = {Pith review of: Ingroup bias is prevalent in user reports of hate and abuse online},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/OMDUSJ6Z}},
  note         = {Machine review of arXiv:2510.04748}
}
read the original abstract

The prevalence of online hate and abuse is a pressing global problem. While tackling such societal harms is a priority for research across the social sciences, it is a difficult task, in part because of the magnitude of the problem. People's engagement with reporting mechanisms ('flagging') online is an increasingly important part of monitoring and addressing harmful content at scale. However, users may not flag content routinely enough, and when users do engage, they may be biased by group identity and political beliefs. Across five well-powered and pre-registered online experiments, we examine the extent of ingroup bias in people's flagging of hate and abuse in four different intergroup contexts: political affiliation, vaccination opinions, beliefs about climate change, and stance on abortion rights. Overall, participants reported abuse reliably, with approximately half of the abusive comments in each study reported. However, a pervasive ingroup bias was present whereby across studies, participants were between 17% and 63% more likely to flag abuse directed at the ingroup than at the outgroup. Our findings offer new insights into the nature of user flagging online, an understanding of which is crucial for enhancing user intervention against online hate and thus ensuring a safer online environment.

Figures

Figures reproduced from arXiv: 2510.04748 by the authors.

Figure 1
Figure 1. Proportion of flags made for abuse and criticism towards ingroup and outgroup targets by political affiliation. Republican participants flagged abusive and critical comments directed at the ingroup to a greater extent than comments directed at the outgroup, while this social bias effect was not present in Democrats. Study 2: Social bias in flagging abuse about vaccination beliefs In Study 1, we found that participan… view at source ↗
Figure 3
Figure 3. Proportion of flags made for abuse and criticism towards ingroup and outgroup targets by [PITH_FULL_IMAGE:figures/full_fig_p024_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

16 extracted references · 2 canonical work pages

  1. [1]

    Enock1, Helen Z

    1 Social bias is prevalent in user reports of hate and abuse online Authors: Florence E. Enock1, Helen Z. Margetts1,2,3 and Jonathan Bright1 1 Public Policy Programme, The Alan Turing Institute, The British Library, 96 Euston Road, London. NW1 2DB. 2 Oxford Internet Institute, University of Oxford, Stephen A. Schwarzman Centre for the Humanities, Radcliff...

  2. [4]

    and online hate can also provoke and justify violent attacks offline (Enock & Over, 2023; Leader Maynard & Benesch, 2016; Ofcom & Kick It Out, 2025; Siegel, 2020). As such, working to tackle online hate and abuse is a priority area for research across the social sciences, but it is a difficult task, in part because of the magnitude of the problem. Many in...

  3. [5]

    The target of the comments was either the Democrats or the Republicans 10 depending on which between-subjects condition participants were in

    and they were designed to include a various types of harmful language, including incitements of violence, dehumanizing slurs, and general abuse (Leader Maynard & Benesch, 2016). The target of the comments was either the Democrats or the Republicans 10 depending on which between-subjects condition participants were in. Examples of abusive comments were, ‘H...

  4. [7]

    None of the interactions were significant (all ps > .05)

    = 1.69, p =.194, ηp²= .005, showing pro-vaccination and anti-vaccination participants flagged to a similar extent overall. None of the interactions were significant (all ps > .05). Overall, comments directed at the ingroup were flagged to a greater extent than equivalent comments directed at the outgroup, both for abusive and critical speech, and this was no...

  5. [12]

    https://doi.org/10.1038/s41558-022-01527-x Gillespie, T. (2018). Custodians of the Internet: Platforms, content moderation, and the hidden decisions that shape social media. Yale University Press. https://books.google.com/books?hl=en&lr=&id=cOJgDwAAQBAJ&oi=fnd&pg=PA1&dq=Custodians+of+the+Internet:+Platforms,+content+moderation,+and+the+hidden+decisions+th...

  6. [20]

    41 Iyengar, S., & Westwood, S. J. (2015). Fear and Loathing across Party Lines: New Evidence on Group Polarization. American Journal of Political Science, 59(3), 690–707. https://doi.org/10.1111/ajps.12152 Johansson, P., Enock, F., Hale, S., Vidgen, B., Bereskin, C., Margetts, H., & Bright, J. (2022). How can we combat online misinformation? A systematic ...

  7. [41]

    https://doi.org/10.1140/epjds/s13688-025-00556-8 Chang, R.-C., Rao, A., Zhong, Q., Wojcieszak, M., & Lerman, K. (2023). # RoeOverturned: Twitter Dataset on the Abortion Rights Controversy. Proceedings of the International AAAI Conference on Web and Social Media, 17, 997–1005. https://ojs.aaai.org/index.php/ICWSM/article/view/22207 Cikara, M., & Van Bavel,...

  8. [308]

    = 16.12, p < .001, ηp²= <.05. Pairwise comparisons showed that pro-abortion rights participants flagged ingroup-directed abuse to a greater extent than outgroup-directed abuse, (MIG = 63.1%, MOG = 56.8%, p < .001), but flagged critical comments to a similar extent for both target groups (MIG = 2.1%, MOG = 1.6%, p = .325). Anti-abortion rights participants fl...

Show all 16 references
  1. [312]

    = 6.45, p =.012, ηp²=.020, suggesting that the effect of target group was dependent on political affiliation. Pairwise comparisons showed that Republicans flagged abuse directed at the ingroup to a greater extent than abuse directed at the outgroup (MIG = 54.9%, MOG = 41.8%, p =...

  2. [315]

    None of the other interactions were significant (ps > .05)

    = 6.73, p =.010, ηp² = .021, though these effects were not relevant for our 27 key research questions2. None of the other interactions were significant (ps > .05). Overall, comments directed at the ingroup were flagged to a greater extent than equivalent comments directed at the...

  3. [316]

    None of the other interactions were significant (all ps > .05)

    = 17.74, p < .001, ηp²= .053, which was not relevant to our key research questions and showed that while criticism was flagged to a similar extent by both groups, climate change believers flagged abuse to a greater extent than non-believers. None of the other interactions were s...

  4. [851]

    Keipi, T., Näsi, M., Oksanen, A., & Räsänen, P. (2016). Online hate and harmful content: Cross-national perspectives. Taylor & Francis. https://library.oapen.org/handle/20.500.12657/22350 Leader Maynard, J., & Benesch, S. (2016). Dangerous speech and dangerous ideology: An int...

  5. [2021]

    (Meta, 2025). High levels of online hate and abuse are problematic for many reasons – exposure can cause severe harm to the psychological wellbeing of targets, with experiences linked to depression, anxiety, fear, low self-esteem and escalation of self-harm (Keipi et al., 2016...

  6. [2023]

    efforts should focus not only on increasing overall engagement with these tools, but also on enhancing the accuracy and fairness of reports. Across all studies, while participants consistently flagged abusive comments to a greater extent than critical ones and comments directed...

  7. [2025]

    and in the US, survey research found that 41% of adults in America had directly experienced online harassment, with a quarter reporting personal experience with stalking, harassment, and physical threats (Vogels, 2021). Platform reports also highlight the scale of the problem ...

  8. [7811]

    https://doi.org/10.1038/s41586-020-2281-1 Kahan, D. M. (2012). Ideology, motivated reasoning, and cognitive reflection: An experimental study. Judgment and Decision Making, 8, 407–424. Kahan, D. M., Hoffman, D. A., Braman, D., & Evans, D. (2012). They saw a protest: Cognitive i...

Pith tools

Reviewed August 4, 2026 · model on record in the stance chip above.