Pith. sign in

Paper Citation Record · LEDGER

Learning diverse attacks on large language models for robust red-teaming and safety tuning

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2405.18540.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.18540 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T18:14:25.844619Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T15:41:17.036171Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f015b579-a826-4ced-b02f-389e2d8872fc · inbound

Neural Genetic Search in Discrete Spaces cites this paper.

Neural Genetic Search in Discrete Spaces Learning diverse attacks on large language models for robust red-teaming and safety tuning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T18:14:25.844619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:14:25.844619Z digest=sha256:fa5a28d6594629b87ff56a50ec01de4808cd1b217fb3f4acf8b289b5c6c51459

Observation 9691966f-846a-41ba-8fad-cb9078025d7d · inbound

Adversarial Preference Learning for Robust LLM Alignment cites this paper.

Adversarial Preference Learning for Robust LLM Alignment Learning diverse attacks on large language models for robust red-teaming and safety tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:21.669202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:21.669202Z digest=sha256:857178e0b9f2a3a99de0017e5c5ee35b36d5a2333b4afe89a28e01ee2dbc6b79

Observation 8b7839bf-1144-4e59-9590-e49c9f2bbb38 · inbound

Jailbreak-R1: Exploring the Jailbreak Capabilities of LLMs via Reinforcement Learning cites this paper.

Jailbreak-R1: Exploring the Jailbreak Capabilities of LLMs via Reinforcement Learning Learning diverse attacks on large language models for robust red-teaming and safety tuning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:08.855205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:08.855205Z digest=sha256:bc2931ff3be69fd095ce6d1b34780e817699ac7e53c0418b0449bb7afba652bd

Observation 26fcdc41-bd95-4bb1-b0d9-aef4fda1bea2 · inbound

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation cites this paper.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Learning diverse attacks on large language models for robust red-teaming and safety tuning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:22.585003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:22.585003Z digest=sha256:15491500a60d5973f9f4b64095b938547d156bb900087b7425c9dbe8eab406be

Observation a1881534-8996-4c2b-9edf-fed53cb73eb8 · inbound

From Seed to Harvest: Augmenting Human Creativity with AI for Red-teaming Text-to-Image Models cites this paper.

From Seed to Harvest: Augmenting Human Creativity with AI for Red-teaming Text-to-Image Models Learning diverse attacks on large language models for robust red-teaming and safety tuning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:20.466196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:20.466196Z digest=sha256:14d5db4b3289d1bafef36ea87d3560245d3ff5cbb40e23234996202b78ad34ec

Observation e3620943-fba8-4cdf-8c73-9b0b3674bd14 · inbound

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM cites this paper.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Learning diverse attacks on large language models for robust red-teaming and safety tuning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:03.820155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:03.820155Z digest=sha256:6d90622fab415e7ddb2017d1127dfd89aff602ddd8a01a4609139d419e9db0d4

Observation eec2fb5a-a4f3-44fc-835b-60ebb1957c71 · inbound

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation cites this paper.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Learning diverse attacks on large language models for robust red-teaming and safety tuning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:36.553261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:36.553261Z digest=sha256:3a333eecc2c82a9f8d12c1c189279384441741d3b141ea2eea4702864051f54a

Observation 7c358abf-7bf5-4559-8649-380cf8d6e359 · inbound

LLM in the Middle: A Systematic Review of Threats and Mitigations to Real-World LLM-based Systems cites this paper.

LLM in the Middle: A Systematic Review of Threats and Mitigations to Real-World LLM-based Systems Learning diverse attacks on large language models for robust red-teaming and safety tuning

Reference 169

Resolution
unresolved
no resolver link, observed 2026-08-04T17:46:19.425971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:46:19.425971Z digest=sha256:8c3a0b8d4c5847e1fe651ed9e357ede978370be037d4a047ddeaaf6e4cd037b9

Observation 305359b0-5c56-4033-a8ad-0c73cf638e2f · inbound

Stable-GFlowNet: Toward Diverse and Robust LLM Red-Teaming via Contrastive Trajectory Balance cites this paper.

Stable-GFlowNet: Toward Diverse and Robust LLM Red-Teaming via Contrastive Trajectory Balance Learning diverse attacks on large language models for robust red-teaming and safety tuning

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:41:17.115469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-09T19:29:36.348680Z digest=sha256:f3cbde52e05652b7d256ec4b832371f85f0397c563c5c7a0d16ff10db80c2791