REVIEW 9 cited by
Can LLM-Generated Misinformation Be Detected?
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The advent of Large Language Models (LLMs) has made a transformative impact. However, the potential that LLMs such as ChatGPT can be exploited to generate misinformation has posed a serious concern to online safety and public trust. A fundamental research question is: will LLM-generated misinformation cause more harm than human-written misinformation? We propose to tackle this question from the perspective of detection difficulty. We first build a taxonomy of LLM-generated misinformation. Then we categorize and validate the potential real-world methods for generating misinformation with LLMs. Then, through extensive empirical investigation, we discover that LLM-generated misinformation can be harder to detect for humans and detectors compared to human-written misinformation with the same semantics, which suggests it can have more deceptive styles and potentially cause more harm. We also discuss the implications of our discovery on combating misinformation in the age of LLMs and the countermeasures.
Forward citations
Cited by 9 Pith papers
-
Tailored untruths: How personalisation challenges LLM safeguards
A 1.6-million-text study of eight LLMs in four languages finds that adding demographic personae to disinformation prompts raises jailbreak rates from 78% to 82%.
-
Sword and Shield: Uses and Strategies of LLMs in Navigating Disinformation
In a 25-participant Werewolf-style game, all roles used an LLM chatbot strategically, as a sword for disinformation and a shield against it.
-
Do people rely on ChatGPT more than their peers to detect deepfake news?
In a lab deepfake-detection task, students shifted more toward ChatGPT's advice than toward peers' advice (weight-of-advice 0.59 vs 0.33), though in 2025 sessions they trusted linguistic experts slightly more than ChatGPT.
-
An Audit and Analysis of LLM-Assisted Health Misinformation Jailbreaks Against LLMs
LLM-generated jailbreak prompts elicited health misinformation from GPT-3.5, Llama 3.1-8B, and Gemini 2.0 Flash at high rates, and both LLM judges and simple classifiers detected the resulting texts with high accuracy.
-
Are Today's LLMs Ready to Explain Well-Being Concepts?
AI judges can score explanations of well-being concepts, and small models fine-tuned with preference data score better than larger models, although judges and explainers are all AIs.
-
The Compositional Architecture of Regret in Large Language Models
The paper claims that regret in LLMs is encoded by interacting neuron groups detectable in the final hidden layer, using new S-CDI, RDS, and GIC metrics.
-
Embodied AI: Emerging Risks and Opportunities for Policy Action
A policy analysis arguing that embodied AI risks are real, under-covered by current US/EU/UK frameworks, and best handled through certification, benchmarks, clarified liability, and economic adaptation.
-
Invariant-based Robust Weights Watermark for Large Language Models
An invariant-based weights watermark embeds per-user keys into the null space of transformer invariants and uses noise to repel collusion.
-
AI-Generated Content in Cross-Domain Applications: Research Trends, Challenges and Propositions
A cross-domain vision paper that surveys AI-generated content and proposes research directions, without introducing new empirical results.
Discussion (0). Continue with ORCID to comment.