Pith. sign in

Paper Citation Record · LEDGER

Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2309.14348.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.14348 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T14:37:39.700462Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T13:14:40.788988Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation aeaa40a0-7599-463a-b09e-f0b523f7175e · inbound

Jailbreak Attacks and Defenses Against Large Language Models: A Survey cites this paper.

Jailbreak Attacks and Defenses Against Large Language Models: A Survey Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:20:44.561823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:20:44.368219Z digest=sha256:2ce735b7a6db8079b8cb59f303562facfb4639e9f4d2d1c148743233e7b3fca8

Observation 89d4f606-6489-4d8a-9803-2cf5ee4b2273 · inbound

Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey cites this paper.

Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:58:25.804371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T20:58:16.237327Z digest=sha256:962e4fc727a23ca06bcde4e4fd2c18df1113761105b18416d91de00a05c20d31

Observation 1987a853-bb8e-40d6-a5bc-1edff36c0992 · inbound

Training Users Against Human and GPT-4 Generated Social Engineering Attacks cites this paper.

Training Users Against Human and GPT-4 Generated Social Engineering Attacks Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T14:37:39.700462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:37:39.700462Z digest=sha256:d35aaeb701f076209961632eb64034f7b9d5539c10efc2a7aef9dd9442309e4a

Observation 64256897-cf8a-4449-8d82-8bfdd6f80501 · inbound

Large Language Model Adversarial Landscape Through the Lens of Attack Objectives cites this paper.

Large Language Model Adversarial Landscape Through the Lens of Attack Objectives Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-09T10:29:50.240137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T10:29:50.240137Z digest=sha256:609ed1f774ff687247f983c622d2bcc7dc491024ce0f416ae2d2f66244cf076d

Observation 6b0d11a4-0216-4169-b477-edaa041104e2 · inbound

JBShield: Defending Large Language Models from Jailbreak Attacks through Activated Concept Analysis and Manipulation cites this paper.

JBShield: Defending Large Language Models from Jailbreak Attacks through Activated Concept Analysis and Manipulation Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T12:25:30.523910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:25:30.523910Z digest=sha256:bcb54329e9681cf5bdfad4df88dd85a00c211ba3130593201964535c5cf2dc67

Observation 0bd44a55-90f6-4db0-8a57-d90f4b2cfaa9 · inbound

SafeAgent: Safeguarding LLM Agents via an Automated Risk Simulator cites this paper.

SafeAgent: Safeguarding LLM Agents via an Automated Risk Simulator Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:44.993336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:44.993336Z digest=sha256:6a4344b8f9decb5bc65d3ec7499c60e52e85074d83b7d621595912828ffaa1aa

Observation 4538763a-b358-4a70-9d37-fc0f6b1d787a · inbound

Breaking the Ceiling: Exploring the Potential of Jailbreak Attacks through Expanding Strategy Space cites this paper.

Breaking the Ceiling: Exploring the Potential of Jailbreak Attacks through Expanding Strategy Space Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:55.748392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:35:55.748392Z digest=sha256:8ac3bac4d0b1ca72194157b41adf813960705cf0bc982f957908a16288b8e097

Observation 4b1863eb-3d29-49fb-a8e3-40e0256476e6 · inbound

Paper Summary Attack: Jailbreaking LLMs through LLM Safety Papers cites this paper.

Paper Summary Attack: Jailbreaking LLMs through LLM Safety Papers Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T16:28:49.811945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:28:49.811945Z digest=sha256:be0c5aed2397ff434ea3bda4ec524bef1fdca7e7adcaa2fe0365ca13ca400939

Observation fd3d5b42-2e5e-4763-ab2e-2b2addca7818 · inbound

Attack the Messages, Not the Agents: A Multi-round Adaptive Stealthy Tampering Framework for LLM-MAS cites this paper.

Attack the Messages, Not the Agents: A Multi-round Adaptive Stealthy Tampering Framework for LLM-MAS Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T04:46:10.420990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:46:10.420990Z digest=sha256:23ff08b9683f4643ae7375c985380e4db8255f9a63c7a43750baa3ea3f5b0fca

Observation ea5ea22f-bccc-43c0-ae88-8d3c153020ed · inbound

ReasoningGuard: Safeguarding Large Reasoning Models with Inference-time Safety Aha Moments cites this paper.

ReasoningGuard: Safeguarding Large Reasoning Models with Inference-time Safety Aha Moments Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-19T01:02:54.673549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T01:02:07.088724Z digest=sha256:b24973707f5d63d8a6aa69025304e4cd6dad9c7586b8be7464b9b30e70bc3587

Observation b66173f8-dd3b-4e8b-abdb-0f16668dd1ca · inbound

Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm cites this paper.

Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM

Reference 124

Resolution
unresolved
no resolver link, observed 2026-08-04T22:33:31.926040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:33:31.926040Z digest=sha256:41d0a15892dc7b9cd7d1b1646f2a5a96677fb211ceddd193b98bd2cbb62dea88

Observation c72c90da-83e9-4b29-bd0f-a194d0003486 · inbound

AI Trust OS -- A Continuous Governance Framework for Autonomous AI Observability and Zero-Trust Compliance in Enterprise Environments cites this paper.

AI Trust OS -- A Continuous Governance Framework for Autonomous AI Observability and Zero-Trust Compliance in Enterprise Environments Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:55:50.655648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T18:48:26.121708Z digest=sha256:14910eea2a77ab92fd9fbe4a58ee747611550db568f933433afe9d4ea4022047

Observation 751e182f-d888-4e85-bff5-6d7c21639e53 · inbound

ADAM: A Systematic Data Extraction Attack on Agent Memory via Adaptive Querying cites this paper.

ADAM: A Systematic Data Extraction Attack on Agent Memory via Adaptive Querying Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:45:57.624741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T17:57:59.918483Z digest=sha256:30e93e4b1f52370fd023d2a716b21b6610345936dc9e6eb97e404049a6b524a9

Observation c63dfc51-94c3-4b1c-b58a-251dc347d6bf · inbound

Ellipsoid Control: A White-list Jailbreak Defense via Benign Latent Modeling cites this paper.

Ellipsoid Control: A White-list Jailbreak Defense via Benign Latent Modeling Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:14:40.790523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T13:10:44.728498Z digest=sha256:6cce9094402d7651d896243b58047cabe018223b78a9b6f10ceb11a9537d338f

Observation 5e6ea70a-49cb-425d-bb61-b5a8cff0c5f5 · inbound

Learning from Mistakes: Can LLM Self-Recover after Misalignment? cites this paper.

Learning from Mistakes: Can LLM Self-Recover after Misalignment? Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-13T18:51:10.298187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:51:10.298187Z digest=sha256:857daaedef43a01d5f658ec0d980fd76f49383dfb2a12975f612a384ab63be0d

Observation ba8886ca-b654-4e61-8022-fc96553d7efc · inbound

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks cites this paper.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:13.252157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:13.252157Z digest=sha256:35605c9f28015683cc0c944dd85a4b61003902110b0d1d21f6aadb6cfaee15c8