Pith. sign in

REVIEW 14 cited by

Generative Language Models and Automated Influence Operations: Emerging Threats and Potential Mitigations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2301.04246 v1 pith:275U7PLB submitted 2023-01-10 cs.CY

classification cs.CY
keywords operationsinfluencelanguagemodelscontentmitigationsactorsgenerative
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Generative language models have improved drastically, and can now produce realistic text outputs that are difficult to distinguish from human-written content. For malicious actors, these language models bring the promise of automating the creation of convincing and misleading text for use in influence operations. This report assesses how language models might change influence operations in the future, and what steps can be taken to mitigate this threat. We lay out possible changes to the actors, behaviors, and content of online influence operations, and provide a framework for stages of the language model-to-influence operations pipeline that mitigations could target (model construction, model access, content dissemination, and belief formation). While no reasonable mitigation can be expected to fully prevent the threat of AI-enabled influence operations, a combination of multiple mitigations may make an important difference.

Discussion (0). Sign in to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 146 citations worldwide. Full citation record

  1. InfoOps Bench: A live information operations safety benchmark

    cs.AI 2026-07 conditional novelty 7.0 of 10

    Most of 17 frontier models can be co-opted to promote live state-backed information-operation claims, with integrity from 8.8% to 94.5% and large China-specific filtering in several Chinese-developed models.

  2. On Capturing the Narrative: Social Media Manipulation Wargaming for Cyberliteracy

    cs.CY 2026-07 conditional novelty 7.0 of 10

    A 4,000-NPC, 108-team LLM-bot wargaming competition yielded no gain in participants' bot-detection confidence and an apparent decrease in their emotional discomfort with misinformation.

  3. The AI Cognitive Trojan Horse: How Large Language Models May Bypass Human Epistemic Vigilance

    cs.HC 2026-01 accept novelty 7.0 of 10

    Conversational AI may evade human epistemic vigilance through 'honest non-signals'—genuine traits like fluency and helpfulness that no longer carry the trust information they do in humans.

  4. Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm

    cs.CR 2025-09 conditional novelty 7.0 of 10

    A trigger-tag watermark embedded by fine-tuning lets modified LLMs mark their own phishing outputs for cheap detection.

  5. Large language models can effectively convince people to believe conspiracies

    cs.AI 2026-01 conditional novelty 6.0 of 10

    In three experiments, GPT-4o instructed to argue for a conspiracy raised believers' confidence about as much as it lowered it when arguing against; a truth-constraining prompt and a corrective debrief largely undid the harm.

  6. Towards Trustworthy AI: Characterizing User-Reported Risks across LLMs "In the Wild"

    cs.CY 2025-09 conditional novelty 6.0 of 10

    Across seven AI chatbots, Reddit users report mostly reliability failures, with each chatbot showing a distinct pattern of safety, privacy, and security complaints.

  7. EPT Benchmark: Evaluation of Persian Trustworthiness in Large Language Models

    cs.CL 2025-09 conditional novelty 6.0 of 10

    A new Persian-Islamic trustworthiness benchmark ranks Claude highest and Qwen lowest across eight LLMs and finds safety is the weakest dimension.

  8. When Autonomy Goes Rogue: Preparing for Risks of Multi-Agent Collusion in Social Systems

    cs.AI 2025-07 conditional novelty 6.0 of 10

    In a 1,000-agent social simulation, decentralized groups of malicious AI agents spread more misinformation and commit more fraud than centralized groups, and they adapt to evade content moderation.

  9. Sword and Shield: Uses and Strategies of LLMs in Navigating Disinformation

    cs.HC 2025-06 conditional novelty 6.0 of 10

    In a 25-participant Werewolf-style game, all roles used an LLM chatbot strategically, as a sword for disinformation and a shield against it.

  10. AI Propaganda factories with language models

    cs.CR 2025-08 conditional novelty 5.0 of 10

    Small language models sustain political personas and become more ideologically extreme when replying to counter-arguments, according to a language-model judge.

  11. Pruning General Large Language Models into Customized Expert Models

    cs.CL 2025-06 conditional novelty 5.0 of 10

    Cus-Prun identifies and removes neurons that are irrelevant to a user's target language, domain, and task, producing specialized expert models without post-training.

  12. How does Misinformation Affect Large Language Model Behaviors and Preferences?

    cs.CL 2025-05 conditional novelty 5.0 of 10

    MisBench provides a 10.3M-example benchmark of styled, conflict-based misinformation and shows LLMs' detection accuracy depends strongly on conflict type and textual style.

  13. Towards a Humanized Social-Media Ecosystem: AI-Augmented HCI Design Patterns for Safety, Agency & Well-Being

    cs.HC 2025-11 conditional novelty 4.0 of 10

    Five browser-side design patterns (rewriter, integrity meter, feed curator, micro-withdrawal, recovery mode) aim to give users control over social feeds through an explainable AI intermediary, with evaluation still pending.

  14. Social Media Information Operations

    cs.SI 2025-08 accept novelty 2.0 of 10

    A tutorial that frames social media information operations as an optimization problem and surveys the analytics, threat models, and countermeasures that support it.

Pith tools