Pith. sign in

Paper Citation Record · LEDGER

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM

As of 18 August 2026, this Paper Citation Record lists 100 of 251 outbound references and 6 inbound Pith citation observations for arXiv:2508.05775.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.05775 v3

Coverage vector

measured 100 of 251 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T23:13:04.217122Z

measured 106 of 106 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T07:06:20.077337Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-25T07:05:26.724164Z

Reference resolution

100 of 251 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved90
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch8

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 261d18ad-fab4-4f43-9dd3-ecb088b7fbac · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T23:12:59.700739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:12:59.700739Z digest=sha256:b9b372f0ccd89395036c4bf09ad3cd5ccb1122d281947f2f30901e7e3e0e0a30

Observation df7a80e7-73ef-4c35-b707-6f4833e9d6b6 · outbound

This paper cites Certifying LLM Safety against Adversarial Prompting.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Certifying LLM Safety against Adversarial Prompting

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T23:12:59.802317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:12:59.802317Z digest=sha256:bce28913b9eaad0997a947b8ab29bf60ee0c75edff369e141b326232a9483f97

Observation 47a90b26-9d41-46ec-b784-b65d7cb3fdac · outbound

This paper cites and Arora, Simran and Mazeika, Manias and Hendrycks, Dan and Lin, Zinan and Cheng, Yu and Koyejo, Sanmi and Song, Dawn and Li, Bo , title =.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM and Arora, Simran and Mazeika, Manias and Hendrycks, Dan and Lin, Zinan and Cheng, Yu and Koyejo, Sanmi and Song, Dawn and Li, Bo , title =

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T23:12:59.919407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:12:59.919407Z digest=sha256:f46d038a9f0df06ece269565b8ae2c6c0c071f3896284aeb9339e18d91e22b13

Observation 8b639068-14e2-43aa-bcb9-7271a1414879 · outbound

This paper cites TrustGPT: A Benchmark for Trustworthy and Responsible Large Language Models.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM TrustGPT: A Benchmark for Trustworthy and Responsible Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:00.005776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:00.005776Z digest=sha256:c16331cda046cb50f0edbcfe5ea8fc0b19ad7f8fd464c1dbb75f14bf40f2bf56

Observation caaf4c2e-3351-4090-8fc0-c064cb0b8f5f · outbound

This paper cites Discourse & Society , volume=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Discourse & Society , volume=

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:00.104978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:00.104978Z digest=sha256:bf1a35f70de4ef971ca8c78b2d48f4c96c5d31480528e39791e52118db6e2a7b

Observation 56cbbbe5-c55e-4f2c-a79d-55d70c7ac078 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Advances in Neural Information Processing Systems , volume=

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:00.181127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:00.181127Z digest=sha256:5a39affbe4042e88c396b1a1dda21f4f4eb99482db843f5261f2a4e22dfd0b4c

Observation 31ebaa58-fc08-479b-8ccd-d8f0970e2863 · outbound

This paper cites Proceedings of the 26th International Symposium on Research in Attacks, Intrusions and Defenses , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the 26th International Symposium on Research in Attacks, Intrusions and Defenses , pages=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:00.321368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:00.321368Z digest=sha256:4033651ecc2e186028fac53d16555282daad8e7d35ef47275f2429d30f74480b

Observation 3024722d-7cfa-4993-86b6-52c8b1367780 · outbound

This paper cites Findings of the association for computational linguistics: EMNLP 2023 , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Findings of the association for computational linguistics: EMNLP 2023 , pages=

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:00.411654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:00.411654Z digest=sha256:20781b3b8d53075ba51e73b164b652ca316b210c0fca3291b5accd95b074cabc

Observation b62dcf92-388f-4ed0-812f-e6cd008e0d32 · outbound

This paper cites IEEE Transactions on Cognitive and Developmental Systems , year=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM IEEE Transactions on Cognitive and Developmental Systems , year=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:00.524651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:00.524651Z digest=sha256:fb5e03d5666b5c3e54f2dff0d4c909678f86119cd2f14cb975ddcdc19bb05366

Observation 1852a291-a39b-4b13-9bd1-39bd7f229e87 · outbound

This paper cites Proceedings of the 2022 ACM SIGSAC Conference on Computer and Communications Security , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the 2022 ACM SIGSAC Conference on Computer and Communications Security , pages=

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:00.699894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:00.699894Z digest=sha256:6a72d78986db13947ba1cb70e3fe6083c87d6cb4ebeb9b4196d5a55270fdfe9b

Observation 7bce24a2-e0a2-49ac-ba41-f6ce07a77dcd · outbound

This paper cites Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:00.793432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:00.793432Z digest=sha256:3eade4033b272e766969a0a23f02a18e572799939de2e92d3aaab96e08f4c666

Observation 588c10bd-6657-46d9-8169-db93e98a0fa8 · outbound

This paper cites Large Language Models are Vulnerable to Bait-and-Switch Attacks for Generating Harmful Content.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Large Language Models are Vulnerable to Bait-and-Switch Attacks for Generating Harmful Content

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:00.922194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:00.922194Z digest=sha256:06bc43f045a18fe0cb51d18f7ee61216810fa14ec8b6efe9c9d9e7fc00e2cdab

Observation 8cdb97cf-27b1-4e68-9b34-49645946c50e · outbound

This paper cites Jigsaw Puzzles: Splitting Harmful Questions to Jailbreak Large Language Models.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Jigsaw Puzzles: Splitting Harmful Questions to Jailbreak Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:01.081890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:01.081890Z digest=sha256:386bdfd8e7148baa72e97968fd60894c7320a083defb2d47c7537efdf12905a8

Observation 0f7c2101-7c04-434e-bff4-bf7a40f3b1d4 · outbound

This paper cites Security and Communication Networks , volume=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Security and Communication Networks , volume=

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:01.164060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:01.164060Z digest=sha256:c4238b19f2022a81f9723e90f1124b996d74eedc6ba0852e0fef12145f5954bf

Observation af481b02-ce75-483f-923f-36f1c96e498c · outbound

This paper cites Proceedings of the First Workshop on Social Influence in Conversations (SICon 2023) , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the First Workshop on Social Influence in Conversations (SICon 2023) , pages=

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:01.264593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:01.264593Z digest=sha256:71453fb89deacc238ddc20f87150f0788879c30ff48a2fea5810e4123fcab613

Observation bc88f121-4999-4919-9bfb-05db88f4f74a · outbound

This paper cites Alignment is not sufficient to prevent large language models from generating harmful information: A psychoanalytic perspective.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Alignment is not sufficient to prevent large language models from generating harmful information: A psychoanalytic perspective

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:01.410828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:01.410828Z digest=sha256:847cc7d427adb0c029c937b69d60fedf9e448613d2e6b4b536a9afac4a9b615e

Observation 3be65711-c14d-4c93-828e-947d429cc5fc · outbound

This paper cites Systematic Rectification of Language Models via Dead-end Analysis.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Systematic Rectification of Language Models via Dead-end Analysis

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:01.483650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:01.483650Z digest=sha256:df32eba860c7d23bc23590653d75f5df559d8216d754450a514fd202fcfb6ebf

Observation f5f12ff1-141f-460e-8f75-c76ff8b69b78 · outbound

This paper cites Successor Features for Efficient Multisubject Controlled Text Generation.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Successor Features for Efficient Multisubject Controlled Text Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:01.514336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:01.514336Z digest=sha256:5a35e09c0152d843b48983a2101c55f99f4092ff2bac87b3076a4d736a0bd39f

Observation 699178a1-2ce7-42cb-8f5a-055746c3e91c · outbound

This paper cites Expert-Guided Extinction of Toxic Tokens for Debiased Generation.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Expert-Guided Extinction of Toxic Tokens for Debiased Generation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:01.555889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:01.555889Z digest=sha256:c291415564ccfab20f9d5799507aec97950dbeb7963c5995f514285fd163ca2c

Observation cfde5f27-fd0a-4007-bba6-ae62fc3df428 · outbound

This paper cites Adversarial DPO: Harnessing Harmful Data for Reducing Toxicity with Minimal Impact on Coherence and Evasiveness in Dialogue Agents.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Adversarial DPO: Harnessing Harmful Data for Reducing Toxicity with Minimal Impact on Coherence and Evasiveness in Dialogue Agents

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:01.607883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:01.607883Z digest=sha256:d036821bb8540324dd5c7a436a03d28f9c9c78dddd111206734ee960c57a5c5f

Observation 78d438c3-a207-46c7-ae91-31d8267607d5 · outbound

This paper cites Plug and Play with Prompts: A Prompt Tuning Approach for Controlling Text Generation.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Plug and Play with Prompts: A Prompt Tuning Approach for Controlling Text Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:01.681830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:01.681830Z digest=sha256:604ddf183b8c80c2d21a4decff433426c06a810ed642ba6b03f629fd6b7bd373

Observation 1fdd39d6-cfdd-4101-a7fb-caaac4a64647 · outbound

This paper cites 2024 IEEE Security and Privacy Workshops (SPW) , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM 2024 IEEE Security and Privacy Workshops (SPW) , pages=

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:01.769771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:01.769771Z digest=sha256:3ab708bec9adceefa1cf55c13d5707c9b772de3771b35913811d8e93ec3ab5f0

Observation 32390486-8e5d-4ff5-916e-cbf37a9521c9 · outbound

This paper cites I’m fully who I am.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM I’m fully who I am

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:01.803231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:01.803231Z digest=sha256:55a231ec374210b5841ed2e147cb2b02f835748b0378a781504c0edd772f19a8

Observation fed7984b-8f98-4228-b648-4a15ac91245a · outbound

This paper cites From Text to MITRE Techniques: Exploring the Malicious Use of Large Language Models for Generating Cyber Attack Payloads.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM From Text to MITRE Techniques: Exploring the Malicious Use of Large Language Models for Generating Cyber Attack Payloads

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:01.893879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:01.893879Z digest=sha256:305d868d5d0b87f83e8e2eae2f3eb31165206bc7d4fe05cc0e0fb7119322088c

Observation 988ccafc-7a5e-4aaa-8ef2-6504b7d2a28e · outbound

This paper cites Scientific Reports , volume=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Scientific Reports , volume=

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:01.986502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:01.986502Z digest=sha256:a1d2cb41811199e990c2f8cfd998a23df610618489a169aaebbe691af9858fc6

Observation ae41fc50-c3db-4742-887b-f06e3f739fa2 · outbound

This paper cites Otolaryngology--Head and Neck Surgery , year=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Otolaryngology--Head and Neck Surgery , year=

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:02.013049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:02.013049Z digest=sha256:d732cbbd726456acd5dd8822091b55f6e554bd9ebf1de9d65ab2499821e0258b

Observation b8bb058e-c906-4d8a-880f-626b48f65a86 · outbound

This paper cites The 2024 ACM Conference on Fairness, Accountability, and Transparency , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM The 2024 ACM Conference on Fairness, Accountability, and Transparency , pages=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:02.039730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:02.039730Z digest=sha256:0c3f2696b1d002b952deefcc8b46545f2eb51abc539a901bcaf663d6118031ca

Observation d56f67aa-6057-4433-8d0b-4a4e33158069 · outbound

This paper cites The 2024 ACM Conference on Fairness, Accountability, and Transparency , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM The 2024 ACM Conference on Fairness, Accountability, and Transparency , pages=

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:02.115290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:02.115290Z digest=sha256:8fbfcb7aa1ac20aa3731f7d40e29a242ec5459295c078226947ec2aa25ec8d39

Observation 4c0ad005-a921-4062-870c-d34e4c8d7ed4 · outbound

This paper cites Inclusivity in Large Language Models: Personality Traits and Gender Bias in Scientific Abstracts.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Inclusivity in Large Language Models: Personality Traits and Gender Bias in Scientific Abstracts

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:02.198098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:02.198098Z digest=sha256:bde413f96e8e41c8c90d00785dbd63811fb6d9945024d790dfcdee7014157a86

Observation fe70d7f0-6775-42a6-bfbb-7b348e29d36b · outbound

This paper cites The African Woman is Rhythmic and Soulful: An Investigation of Implicit Biases in LLM Open-ended Text Generation.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM The African Woman is Rhythmic and Soulful: An Investigation of Implicit Biases in LLM Open-ended Text Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:02.239622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:02.239622Z digest=sha256:73a6d3513895864a530263c452be4745bc35f5e06b03712fd07c2499961cdbd5

Observation 3af60680-828a-4fa1-978f-01de151e4a68 · outbound

This paper cites TATuP-Zeitschrift f.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM TATuP-Zeitschrift f

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:02.263456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:02.263456Z digest=sha256:6cbb7476a4337fb12b9e86d43cd85895001e4c08b8020d203fa6ecfa7d58cd1a

Observation 1d570c05-8e04-4d25-a096-1c8596171a84 · outbound

This paper cites Gender Bias in Decision-Making with Large Language Models: A Study of Relationship Conflicts.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Gender Bias in Decision-Making with Large Language Models: A Study of Relationship Conflicts

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:02.311367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:02.311367Z digest=sha256:459ddd88ac9c030e8fb949e877ad585d5a90c2f80de93c78753c6a408c81e457

Observation 66a13d2b-461c-4acf-bf25-6351a60fde69 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Advances in Neural Information Processing Systems , volume=

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:02.418496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:02.418496Z digest=sha256:46a0ebbfcb114e4609d262d760eee89c6212d91a1734b0d49aff1b36fac59917

Observation cbef038d-dd7b-4ab2-abae-e276cb072e90 · outbound

This paper cites Attack Prompt Generation for Red Teaming and Defending Large Language Models.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Attack Prompt Generation for Red Teaming and Defending Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:02.478257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:02.478257Z digest=sha256:b0dda518d84f22a3bc9ba1c6f755bf448ec7b711b154f335b3ee260827401227

Observation b58ed619-db02-496a-ac0b-a68b93782ad9 · outbound

This paper cites Exploring the Adversarial Capabilities of Large Language Models.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Exploring the Adversarial Capabilities of Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:02.527686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:02.527686Z digest=sha256:0d5a2eb142bfda1d493368af2de56e8daca4ebb2fc604d82e7762e3ec8261eaf

Observation f4f56dd4-36fb-4e56-8077-3db47afe5ee1 · outbound

This paper cites Safety Alignment in NLP Tasks: Weakly Aligned Summarization as an In-Context Attack.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Safety Alignment in NLP Tasks: Weakly Aligned Summarization as an In-Context Attack

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:02.596616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:02.596616Z digest=sha256:9517d02a9da16e0d4e449a428d1420daae7cdd8002a7d0e82360f24c6b47ee7c

Observation 485469a0-bdc5-4ccf-80e5-84675f2a08ea · outbound

This paper cites The Dark Side of Human Feedback: Poisoning Large Language Models via User Inputs.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM The Dark Side of Human Feedback: Poisoning Large Language Models via User Inputs

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:02.685636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:02.685636Z digest=sha256:49ae38a75ef72dc819509fb043363749337e641c07a1066a379e7439502927a9

Observation 042ae570-0dfa-4358-bcfb-f09ecfaf014c · outbound

This paper cites F2A: An Innovative Approach for Prompt Injection by Utilizing Feign Security Detection Agents.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM F2A: An Innovative Approach for Prompt Injection by Utilizing Feign Security Detection Agents

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:02.828938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:02.828938Z digest=sha256:bf30805b85c152e03975af320cf16f88e89d03ac8fe0868622af10bc6bf5a84a

Observation 79103127-ed0f-494b-82bf-a8c035d439d7 · outbound

This paper cites an unresolved cited work.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:02.907233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:02.907233Z digest=sha256:c258a6b860b18f6099978b6e5ace8bcd48f974603b011c6ee0b4155b86547739

Observation ce3e6df4-cdd3-46e3-8ef4-7ed728a81656 · outbound

This paper cites NeurIPS 2023 Workshop on Instruction Tuning and Instruction Following , year=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM NeurIPS 2023 Workshop on Instruction Tuning and Instruction Following , year=

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:02.998152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:02.998152Z digest=sha256:ffe8503aee46c6715a788a6605a4dfcd7b9a362e06a75361c3f78d2975c842c8

Observation 40bdc7da-2a16-4037-9803-63721fd03371 · outbound

This paper cites Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:03.063808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:03.063808Z digest=sha256:21d0e495c0ceb4eff49c7b862b951d605d907a8b6d40457284b64485736a879d

Observation 173e19ea-556e-41c7-b0b1-fbe63074ad43 · outbound

This paper cites Fifty Shades of Bias.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Fifty Shades of Bias

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:03.146442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:03.146442Z digest=sha256:3cb24f6a7e94fddd62ea7b3742dc1834faa69dae0bfe8e6e51727bdbea5ce92b

Observation 78ea9e60-7d1d-44ab-baff-8594fc267619 · outbound

This paper cites K o C o S a: K orean Context-aware Sarcasm Detection Dataset.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM K o C o S a: K orean Context-aware Sarcasm Detection Dataset

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:03.198820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:03.198820Z digest=sha256:74e9982791cdef633fd000d33062d8c70b5beb8745cacdfec9341f3bf64dba8e

Observation 0203ae01-2eae-47bf-a360-bd268aa7305e · outbound

This paper cites ParaFusion: A Large-Scale LLM-Driven English Paraphrase Dataset Infused with High-Quality Lexical and Syntactic Diversity.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM ParaFusion: A Large-Scale LLM-Driven English Paraphrase Dataset Infused with High-Quality Lexical and Syntactic Diversity

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:03.303646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:03.303646Z digest=sha256:e6ecb4cbd07b69ac745e1bf903f72ce3afa3013cf8771c4083f3bb09dcf1c8bd

Observation 2eb00353-2d98-4488-94c3-cc9b9f61f1b9 · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:03.368332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:03.368332Z digest=sha256:e7c3dce6c017326f07c89acf36d6a45b6b6721e32fd7e32fe17a6787fd119281

Observation 7673df10-b083-4358-b07d-637e107f30ab · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) , pages=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:03.407488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:03.407488Z digest=sha256:b30b3636da222b9d71c1f1e525353347067c75c571c2abf617567674e99afb65

Observation 14afb384-feac-4544-b26d-75ad2f90ccbd · outbound

This paper cites Text generation for dataset augmentation in security classification tasks.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Text generation for dataset augmentation in security classification tasks

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:03.471388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:03.471388Z digest=sha256:c7011ab3a9b72ae920b63aabf1c48193433eaea0d857feaadaee91e0f0ca30a7

Observation 8fa9f4d3-d632-40d9-a50d-d4ed6826d947 · outbound

This paper cites Explore, Establish, Exploit: Red Teaming Language Models from Scratch.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:03.550174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:03.550174Z digest=sha256:85a424b04c6ef60aa5fb63669ac914e4f3d79569fbfa5d8b25637e587398be1d

Observation 4b0c8281-d9ad-4be3-88b5-307123e3bc52 · outbound

This paper cites International Conference on Machine Learning , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM International Conference on Machine Learning , pages=

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:03.580538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:03.580538Z digest=sha256:bc42a2638d4a0c6d63ab21d3b437289a499c75033fe5cccfe58dcff267d14f9b

Observation 8705ad8d-9cd4-466b-b034-18c9e58deb4f · outbound

This paper cites Adversarial Fine-Tuning of Language Models: An Iterative Optimisation Approach for the Generation and Detection of Problematic Content.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Adversarial Fine-Tuning of Language Models: An Iterative Optimisation Approach for the Generation and Detection of Problematic Content

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:03.657036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:03.657036Z digest=sha256:a31477c0f49b33ea21ab183c087ca84700616c815940479ee969e2abba640729

Observation 79165f25-f3b1-4ee8-a196-890c4414871e · outbound

This paper cites TroubleLLM: Align to Red Team Expert.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM TroubleLLM: Align to Red Team Expert

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:03.721737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:03.721737Z digest=sha256:63535ab37970852e4c610063607dd8e802e619c31850eda05a8248a6fea81252

Observation e3620943-fba8-4cdf-8c73-9b0b3674bd14 · outbound

This paper cites Learning diverse attacks on large language models for robust red-teaming and safety tuning.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Learning diverse attacks on large language models for robust red-teaming and safety tuning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:03.820155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:03.820155Z digest=sha256:b472cbca13b4a43c48350f12a028a6a557145c466792bbeff38f6f9578817b42

Observation 79871108-0d5c-49d3-bc79-dd7c6aee161c · outbound

This paper cites Outcome-Constrained Large Language Models for Countering Hate Speech.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Outcome-Constrained Large Language Models for Countering Hate Speech

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:03.900039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:03.900039Z digest=sha256:552a200064fa95f8475c167db7ae7bc45bade8d7788a6b6a2c3bf51ed9e3b00c

Observation 7f166ca7-d232-427c-9d3c-aa44f147d535 · outbound

This paper cites Proceedings of the CHI Conference on Human Factors in Computing Systems , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the CHI Conference on Human Factors in Computing Systems , pages=

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:03.984396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:03.984396Z digest=sha256:4b79cfaa47d1eb47c2559cf5cf56241b25d17735263de86056f0cec79d7d77a9

Observation 8545ce46-1cb4-42ac-8900-8c4ca76e4704 · outbound

This paper cites an unresolved cited work.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.024459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.024459Z digest=sha256:c138f6dd7567b94f800e6d41d570f9b18d7ba18b2872b81beae7bf71dfdff217

Observation 2e9cd479-162b-4609-8a59-504cfcf47563 · outbound

This paper cites Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) , pages=

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.028618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.028618Z digest=sha256:d103a21ea4c8dff77bc730729729aa643912a70689f4cf28bfbd770890985aed

Observation 9765cd3b-ca29-4b0b-8c55-9a5c99e9e018 · outbound

This paper cites Seeing Seeds Beyond Weeds: Green Teaming Generative AI for Beneficial Uses.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Seeing Seeds Beyond Weeds: Green Teaming Generative AI for Beneficial Uses

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.032739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.032739Z digest=sha256:64dcc7671352d21ac649a5c40dd56d29e3eb835f8498a4a4dc91c3c1ec512c04

Observation ffb8d0e4-52c2-4072-94a7-aaa4be383cf5 · outbound

This paper cites Intrinsic Self-correction for Enhanced Morality: An Analysis of Internal Mechanisms and the Superficial Hypothesis.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Intrinsic Self-correction for Enhanced Morality: An Analysis of Internal Mechanisms and the Superficial Hypothesis

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.037075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.037075Z digest=sha256:dbd3e52dfcc9014e8932ddc3ac132ea959202cc35cd5c671c7a123b39a8dc452

Observation a3aa813f-3e0d-4da8-a522-c6652f44ff70 · outbound

This paper cites Probing the Safety Response Boundary of Large Language Models via Unsafe Decoding Path Generation.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Probing the Safety Response Boundary of Large Language Models via Unsafe Decoding Path Generation

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.041425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.041425Z digest=sha256:0e2c62f8bc10a5b40a04895966ce5f7e6ec2f3a241886f37699dce840e2351f0

Observation eb41db75-9280-4d7b-b079-445d4b108a1e · outbound

This paper cites Creativity Has Left the Chat: The Price of Debiasing Language Models.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Creativity Has Left the Chat: The Price of Debiasing Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.046627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.046627Z digest=sha256:0360114cd656718db43608b0eaedec6f4cdf641d8f7f7d2b448b8dac2389a2be

Observation 4fb36029-f869-4632-84e2-4798f5ba304c · outbound

This paper cites Controlled Text Generation for Large Language Model with Dynamic Attribute Graphs.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Controlled Text Generation for Large Language Model with Dynamic Attribute Graphs

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.050958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.050958Z digest=sha256:064258b6056f55166764e9f7f2b2ee20b5be26086a65b274306809f47dc8098a

Observation 7e68a8cf-756a-4b8e-8909-781950102777 · outbound

This paper cites Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) , pages=

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.055130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.055130Z digest=sha256:87ac5eeb1f654b167cd664d62cc64dfb8253545f89b8bf2f29a54cb4a7fdb609

Observation fadb5382-8c1b-4e3f-88e9-d283ed392b10 · outbound

This paper cites Yale JL & Tech.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Yale JL & Tech

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.059034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.059034Z digest=sha256:f1469bc28fba9cf01cc6c25ffa08743220a09965d8432a826a3af0ee7b8e764c

Observation 7f71c83e-dfd6-4dcf-b82e-b45d171c540b · outbound

This paper cites Engineering, Technology & Applied Science Research , volume=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Engineering, Technology & Applied Science Research , volume=

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.062924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.062924Z digest=sha256:910e1ea482a338c8412a013627cc29c34feefe8c93424a5b061140c055d1836a

Observation 49673c17-e9d5-4249-9414-3d001d1ff6f3 · outbound

This paper cites Machine-Generated Tweets , author=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Machine-Generated Tweets , author=

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.066922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.066922Z digest=sha256:ed90b0c8cc995850e5197c68cf43ba0562f9bb3dd5e8f5497a972e50b4fa5e31

Observation a8c6e79f-2a1f-468c-8481-f059a24e2c7b · outbound

This paper cites Natural Language Processing , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Natural Language Processing , pages=

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.070937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.070937Z digest=sha256:9cb752c958e8b117dcdc39a089a78c92d152d9b2894c6e2572a540a81261fd7f

Observation fa620c31-7410-4c82-ad9c-610d6e605801 · outbound

This paper cites International Conference on Computational Science and Its Applications , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM International Conference on Computational Science and Its Applications , pages=

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.075456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.075456Z digest=sha256:ab76330796eca096c6dcc1bd92793233ce22014d67e44fb10788df55b653b969

Observation ec361d9f-ec3f-4332-88fd-83b657ee20e4 · outbound

This paper cites Modeling subjectivity (by Mimicking Annotator Annotation) in toxic comment identification across diverse communities.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Modeling subjectivity (by Mimicking Annotator Annotation) in toxic comment identification across diverse communities

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.079715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.079715Z digest=sha256:8683182c56f501570d8676476607c4892c23b5d3029af0d8a3eaa2e111344c68

Observation 4613330a-9ecb-48f5-bb4f-e743afd3f148 · outbound

This paper cites an unresolved cited work.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Unresolved cited work

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.084319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.084319Z digest=sha256:7dd83c90c9ca92352c7e6915fdf071ab5b28f00e45ae85ec7ae9f3fab65c6ed8

Observation baf565a0-4225-4baa-9631-8a37500fc0ce · outbound

This paper cites Regulating Hate Speech Created by Generative AI , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Regulating Hate Speech Created by Generative AI , pages=

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.088949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.088949Z digest=sha256:7754a06d9db55a095f0e5c502fa4d980f7af343d5fc0f6a757111dc2486877d1

Observation 005b5e48-1803-4b5c-a922-4f6a68568577 · outbound

This paper cites A Study of Slang Representation Methods.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM A Study of Slang Representation Methods

Reference 71

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T23:13:08.069771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T23:13:04.092860Z digest=sha256:c892fa1d17d15a46c4b0e120a46bcdcdcffa7d2f480651ea1412f7653049df3e

Observation 6f761ee3-f867-44ab-ad42-3032a44c39ae · outbound

This paper cites Perplexed by Quality: A Perplexity-based Method for Adult and Harmful Content Detection in Multilingual Heterogeneous Web Data.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Perplexed by Quality: A Perplexity-based Method for Adult and Harmful Content Detection in Multilingual Heterogeneous Web Data

Reference 72

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T23:13:08.047629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T23:13:04.096997Z digest=sha256:d3900b2c00da2c312e6f74b859a2439b12a1c8805854867bb374737a9113a42f

Observation 52f71db4-b053-4998-a6a1-ac5e0bbd3db4 · outbound

This paper cites Human-Guided Fair Classification for Natural Language Processing.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Human-Guided Fair Classification for Natural Language Processing

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.101234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.101234Z digest=sha256:e39c86a96a6faa811306de89ddc7a22b5f5f46ebaf11b487d0d1883657d19be7

Observation 65b692fc-8d19-4aab-9192-b7358227d056 · outbound

This paper cites arXiv preprint arXiv:2301.12534 , year=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM arXiv preprint arXiv:2301.12534 , year=

Reference 74

Resolution
verified exact
raw_fallback, observed 2026-08-05T23:13:08.009527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T23:13:04.105543Z digest=sha256:18a2f164354b5602a9d6d577ae59bb295297f5d73fc7ff74fc4a7a382553d83a

Observation 3042f772-ca5c-4b91-8fcb-c770d6a45a37 · outbound

This paper cites International Conference on Advances in Social Networks Analysis and Mining , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM International Conference on Advances in Social Networks Analysis and Mining , pages=

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.109738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.109738Z digest=sha256:83b4289f249e2dc98e8b18e30b48cfaea6793824ccc475a525b887b2f1081530

Observation 5f7ed78f-2433-4640-90f3-a95f2603ed0a · outbound

This paper cites Explicit Toxicity Detection Models with Interactive Visualization , year=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Explicit Toxicity Detection Models with Interactive Visualization , year=

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.114159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.114159Z digest=sha256:e5bcf36be4462a20ab192a5b8815840309df2ee26b643c9adf2e8dadf5cc82e9

Observation 164a9473-9137-4419-a9a1-adad34d60acc · outbound

This paper cites Which Argumentative Aspects of Hate Speech in Social Media can be reliably identified?.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Which Argumentative Aspects of Hate Speech in Social Media can be reliably identified?

Reference 77

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T23:13:07.905145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T23:13:04.118824Z digest=sha256:82c1651473f437c4a717de2767493fbbc8b16cc97232e4134a46f01d0b3da666

Observation 25099d55-a1e2-4d1b-a80b-f0077fbc4f34 · outbound

This paper cites Proceedings of the Second Workshop on NLP for Positive Impact (NLP4PI) , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the Second Workshop on NLP for Positive Impact (NLP4PI) , pages=

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.123304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.123304Z digest=sha256:922ea53b3247b392a6f1cf9b290700e53237d0cb5675c2929a7b3663f91fbe58

Observation a458cb7d-4282-49cc-bc73-972d9425e3d6 · outbound

This paper cites an unresolved cited work.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Unresolved cited work

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.127327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.127327Z digest=sha256:7cc4590ca8a2618aa8aa15a12fa42edb85dd0ea5db76aadc8d6e67fd280a4626

Observation 803845f0-f56b-479d-9657-43c8d43a3d27 · outbound

This paper cites Topological Data Mapping of Online Hate Speech, Misinformation, and General Mental Health: A Large Language Model Based Study.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Topological Data Mapping of Online Hate Speech, Misinformation, and General Mental Health: A Large Language Model Based Study

Reference 80

Resolution
verified exact
local_arxiv, observed 2026-08-05T23:13:07.883128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T23:13:04.131563Z digest=sha256:08c2bb8b4e95c0e897e4561bb4e9170210e7981ac7c1ae658c1d4fb6a3a7f6d7

Observation 53d99bd2-5e9f-4264-a72f-8f076dcc6532 · outbound

This paper cites Demonstrations Are All You Need: Advancing Offensive Content Paraphrasing using In-Context Learning.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Demonstrations Are All You Need: Advancing Offensive Content Paraphrasing using In-Context Learning

Reference 81

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T23:13:07.861922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T23:13:04.135809Z digest=sha256:f01e542dc913b10c00cb9f6cd67d1aea6315a1b45873f006a9dfbef2f33ef468

Observation d6661e22-6d2c-4b57-b6d2-934321f5527a · outbound

This paper cites Beyond plain toxic: building datasets for detection of flammable topics and inappropriate statements.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Beyond plain toxic: building datasets for detection of flammable topics and inappropriate statements

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.139949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.139949Z digest=sha256:83dc258dd6e764d5e4915e722e30a080658ba21857cdcfb7e66aa1fc605265c3

Observation d9798d37-cc31-433b-a0b5-a14eb29082c5 · outbound

This paper cites Evaluation of ChatGPT and BERT-based models for Turkish hate speech detection.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Evaluation of ChatGPT and BERT-based models for Turkish hate speech detection

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.144250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.144250Z digest=sha256:eae3f34db39c1916db4f5d254aefbb85a7875540f2b85a7e29bce994285cffba

Observation 5ec65427-89d8-4191-b8ae-ac3903d98754 · outbound

This paper cites HateRephrase: Zero- and Few-Shot Reduction of Hate Intensity in Online Posts using Large Language Models.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM HateRephrase: Zero- and Few-Shot Reduction of Hate Intensity in Online Posts using Large Language Models

Reference 84

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T23:13:07.840022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T23:13:04.148362Z digest=sha256:b93782943aab502766c5c00a6f04bc7687dfd2ef1613231db80430939805aec8

Observation 10435c6c-a17d-4409-9d63-6408f817ac91 · outbound

This paper cites FFT: Towards Harmlessness Evaluation and Analysis for LLMs with Factuality, Fairness, Toxicity.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM FFT: Towards Harmlessness Evaluation and Analysis for LLMs with Factuality, Fairness, Toxicity

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.152649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.152649Z digest=sha256:4f95f9ab192ca47f49a90718526b89005cda6bc1fc038abcf09c438a0cf173e9

Observation 302948ef-6193-4c33-911b-47b1b120795a · outbound

This paper cites Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.157130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.157130Z digest=sha256:ec3edae925f304374a13a36af2b8ab801e9df08f02af869a7266a4c4e35a69a6

Observation a5365e4f-0f2d-47a2-8d76-5b266eafd26e · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.161107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.161107Z digest=sha256:04ef10083248e11dfc018c0c54bfd3ee492c3fc6cc8c726d2d7706d57272da18

Observation 78766c66-76e8-4a0a-9996-8cb3380fc3b7 · outbound

This paper cites 2024 , isbn =.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM 2024 , isbn =

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.165543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.165543Z digest=sha256:8b6ec4d62500eaa7bea39fba944257e833b4b1c8c66ba1eac1f2da352df49d60

Observation dd78b58d-881b-4fd8-9d82-aad9612a5cf3 · outbound

This paper cites Eagle: Ethical Dataset Given from Real Interactions.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Eagle: Ethical Dataset Given from Real Interactions

Reference 89

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T23:13:07.712366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T23:13:04.169918Z digest=sha256:68baf8369feb72fb662c61283ab0819e560ce40ab9d6a8549550bc4de9d88747

Observation 54e798c2-c413-47ea-bb80-6a08a6c8fc0c · outbound

This paper cites Proceedings of the first workshop on language technology for equality, diversity and inclusion , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the first workshop on language technology for equality, diversity and inclusion , pages=

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.174303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.174303Z digest=sha256:12b8e28ac7c5e4919187a86166c5d5f3645898c4d69a5b47431301d316d2933f

Observation 54f73beb-ca2e-4f98-b9b6-d76066cfccfe · outbound

This paper cites Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.178745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.178745Z digest=sha256:8b9c4c53bf8d16cc53fbfa2d3747b3f8cfcbafdba713fcab0813128c5cc2818d

Observation 34bdbfab-2e6a-4ead-bfb0-5ad5db868f85 · outbound

This paper cites M isgender M ender: A Community-Informed Approach to Interventions for Misgendering.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM M isgender M ender: A Community-Informed Approach to Interventions for Misgendering

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.183378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.183378Z digest=sha256:75c6ce54807b7dda3bec7eafe89601a818a896d0ea31575a2a131fd2cd0ef2f2

Observation 41a2da99-3a2f-442a-9f47-98af30d9b87c · outbound

This paper cites HateTinyLLM : Hate Speech Detection Using Tiny Large Language Models.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM HateTinyLLM : Hate Speech Detection Using Tiny Large Language Models

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.187341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.187341Z digest=sha256:4d5cc612f7eef70f72ebb05772339cc8cf9c1a9db3bea731185fce49519c861c

Observation d422327e-7c3c-459e-9c4c-774677d84c23 · outbound

This paper cites Toxicity Classification in Ukrainian.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Toxicity Classification in Ukrainian

Reference 94

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T23:13:07.675095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T23:13:04.191642Z digest=sha256:6c1f7012d555abbf90b21cd99356c3afbcc017fa0c58f4c45b73e2e30379262d

Observation b7910204-72e1-4448-a7f3-d09091fa36b5 · outbound

This paper cites Proceedings of the International AAAI Conference on Web and Social Media , volume=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the International AAAI Conference on Web and Social Media , volume=

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.196099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.196099Z digest=sha256:6abe0393e2b3e63e7528804d7c842bb492ae2b3098854d61e618dd7c3c11e44e

Observation 2fae51d7-23e7-4aef-aae7-40ffa4c11406 · outbound

This paper cites Proceedings of the International AAAI Conference on Web and Social Media , volume=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the International AAAI Conference on Web and Social Media , volume=

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.200044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.200044Z digest=sha256:26a6b8fedbbcfa87a2accae2eac4cfcd9e3b6fbc8643c6685472ee2dc4c9ce2c

Observation ea2dfb9d-e2ee-4ff1-b6a5-f35c656b1f1a · outbound

This paper cites Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: Student Research Workshop , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: Student Research Workshop , pages=

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.204022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.204022Z digest=sha256:de17746b1e6691721ee60ac19e62cbf7091a80e1d40375dedcdb69494ed02d4f

Observation 0e70d85c-1aa2-4384-84c5-2ffbfd12e9b8 · outbound

This paper cites A Community-Centric Perspective for Characterizing and Detecting Anti-Asian Violence-Provoking Speech.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM A Community-Centric Perspective for Characterizing and Detecting Anti-Asian Violence-Provoking Speech

Reference 98

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T23:13:07.653077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T23:13:04.208197Z digest=sha256:1fe91773f6d42814f78666fb70738972a21375bbe53eb267d37f4f2e879f691c

Observation fc7888dc-7bf9-4d8b-8bba-3b06a2bb1d91 · outbound

This paper cites and Saha, Sriparna and Pasupa, Kitsuchart , title =.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM and Saha, Sriparna and Pasupa, Kitsuchart , title =

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.213203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.213203Z digest=sha256:34988864c043fea11606bde16885c41268604f0b350c51110ccfabb4a8b1f320

Observation 35c4bcef-1500-4a11-a17b-48205080ce62 · outbound

This paper cites On Calibration of LLM-based Guard Models for Reliable Content Moderation.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM On Calibration of LLM-based Guard Models for Reliable Content Moderation

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.217122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.217122Z digest=sha256:c0a05bc36efded95b13e7eba7048abea3909985948605ff5421695469870cf93

Pith citing papers

Observation 56e7e935-1af9-4a4d-aa2f-86ed48c0b5b6 · inbound

BarrierSteer: LLM Safety via Learning Barrier Steering cites this paper.

BarrierSteer: LLM Safety via Learning Barrier Steering Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-29T00:24:32.402891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-25T07:02:03.058731Z digest=sha256:f6023cab2b94d1135afaf5ab63765ab97d64bcc01732f8bb880b7fe998303741

Observation 9b2ca48d-322a-4c1d-9ec2-44a2779397db · inbound

Why Do Large Language Models Generate Harmful Content? cites this paper.

Why Do Large Language Models Generate Harmful Content? Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-29T00:24:32.402891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T15:31:13.545599Z digest=sha256:10fdd82422bea8f0ffb4a68dacadb8b02bed5949978d31ab9c1abcdeac58bb9d

Observation ee3c235e-d736-4b70-b9b0-507b8df37bea · inbound

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue cites this paper.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-29T00:24:32.402891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T11:17:19.079380Z digest=sha256:c69372d8e4f31c1a614128873ddb18332e1530a99c172b4bc9776e70663e517f

Observation 9c3d6022-748a-4c66-b3d8-d7f6c01daa7a · inbound

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue cites this paper.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-29T00:24:32.402891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T07:53:06.928500Z digest=sha256:8e9a69c21efd0263df458cb7dfb8fb10f1ac3777bfda192e5edb589c52f9bf39

Observation a6083394-f2d9-440b-aa9c-11737631a9ca · inbound

Do Coding Agents Understand Least-Privilege Authorization? cites this paper.

Do Coding Agents Understand Least-Privilege Authorization? Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-07-29T00:24:32.402891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T16:34:14.379419Z digest=sha256:0584aaa53063b53a37d76ea618039c06ddc8ea32cc5bd4c5a4433c9708041e3a

Observation 62ac8f37-4f52-40ad-b1cc-44d7042493d1 · inbound

Operational Evidence Gaps for LLMs in Fraud Detection and Trust-and-Safety Workflows cites this paper.

Operational Evidence Gaps for LLMs in Fraud Detection and Trust-and-Safety Workflows Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T07:06:20.077337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:06:20.077337Z digest=sha256:90ed50d432c9b0bb3d94daf2b61b65262a83f6dde76b35e37e99ee54664de631