Pith. sign in

Paper Citation Record · LEDGER

Attack Prompt Generation for Red Teaming and Defending Large Language Models

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2310.12505.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.12505 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:35:20.826210Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T05:25:54.592662Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 15d46676-6f82-46d9-8fe3-7b01c8c96701 · inbound

Adversarial Preference Learning for Robust LLM Alignment cites this paper.

Adversarial Preference Learning for Robust LLM Alignment Attack Prompt Generation for Red Teaming and Defending Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:20.826210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:20.826210Z digest=sha256:3c83e258c8e08dacb363c51ab9be8b8008d5241560c333a062eaa833ade869f1

Observation b06c2a86-7fd2-4537-92e4-511cf268330d · inbound

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming cites this paper.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Attack Prompt Generation for Red Teaming and Defending Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.050296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.050296Z digest=sha256:3272bf833cdf3dc547788b274914a3113c612f9d3caff749bc6b93f815b7ea2d

Observation 2e56ca2d-ef73-4583-ba60-800825df4643 · inbound

GuardVal: Dynamic Large Language Model Jailbreak Evaluation for Comprehensive Safety Testing cites this paper.

GuardVal: Dynamic Large Language Model Jailbreak Evaluation for Comprehensive Safety Testing Attack Prompt Generation for Red Teaming and Defending Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:40:52.210500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:40:52.210500Z digest=sha256:69fd4a84ec549facd4efd8f109eb8c496d3e8eca092899f8c80153437ee72a94

Observation cbef038d-dd7b-4ab2-abae-e276cb072e90 · inbound

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM cites this paper.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Attack Prompt Generation for Red Teaming and Defending Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:02.478257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:02.478257Z digest=sha256:54c6441f70b683faca3dcdc83f1f2012d43471dc2d2b63aa1c9a568513417dde

Observation 3409c0bc-093a-48b3-8b4c-02487b8338b1 · inbound

Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint cites this paper.

Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint Attack Prompt Generation for Red Teaming and Defending Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T23:10:21.359755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:10:21.359755Z digest=sha256:643fca53d086c1846ab237a348dfda4f59a068a71d1faad85a5779ca3f26143a

Observation bdcdeac0-827d-4bbc-9058-f8fe59b29b4a · inbound

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models cites this paper.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models Attack Prompt Generation for Red Teaming and Defending Large Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:25:54.594563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:1a39821fc7bcdde1db51c50b46b12f115ebec1a5734343032d56567c2490a23d

Observation cf3b0c90-f06b-44cb-ba23-b8b51fe25423 · inbound

Beyond Context: Large Language Models' Failure to Grasp Users' Intent cites this paper.

Beyond Context: Large Language Models' Failure to Grasp Users' Intent Attack Prompt Generation for Red Teaming and Defending Large Language Models

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:11:13.577168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T20:09:25.827452Z digest=sha256:945598be1858465721c53029c0f72089c867b7615f82854a5375948f32e96ff0

Observation 72a24dfa-255a-4af3-803c-8c8f48093676 · inbound

Adaptive Instruction Composition for Automated LLM Red-Teaming cites this paper.

Adaptive Instruction Composition for Automated LLM Red-Teaming Attack Prompt Generation for Red Teaming and Defending Large Language Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:06:03.793513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-09T23:34:36.280092Z digest=sha256:a7aaf89833537a2bb493f5463e9548bf39b420fcd7524dfd2f518d78f17e9f46