Pith. sign in

Paper Citation Record · LEDGER

Bypassing the Safety Training of Open-Source LLMs with Priming Attacks

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2312.12321.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.12321 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T19:32:24.131114Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T01:15:13.251025Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c2ac0286-28ca-49a7-a4d0-56b6ededdb7f · inbound

Fast Proxies for LLM Robustness Evaluation cites this paper.

Fast Proxies for LLM Robustness Evaluation Bypassing the Safety Training of Open-Source LLMs with Priming Attacks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T19:32:24.131114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T19:32:24.131114Z digest=sha256:6c4326be1e88b7e47ec0b40daeae2f1d0897cc496af8ecc98b489b2a6d665673

Observation 72be71cf-0c7e-4576-a367-8697020cb15d · inbound

Saffron-1: Safety Inference Scaling cites this paper.

Saffron-1: Safety Inference Scaling Bypassing the Safety Training of Open-Source LLMs with Priming Attacks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:25.890123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:25.890123Z digest=sha256:12b12054d9e891fd17fb8a536e614e0969f8b89a19466a53d386b08e6e288ba9

Observation e0609a61-8420-4d81-8112-4db5f00d0667 · inbound

Few-Shot Truly Benign DPO Attack for Jailbreaking LLMs cites this paper.

Few-Shot Truly Benign DPO Attack for Jailbreaking LLMs Bypassing the Safety Training of Open-Source LLMs with Priming Attacks

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:07:27.007452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T07:06:46.387088Z digest=sha256:c693173fac3e07bc1aa16aef610688a7a35cb23108c7427a68d5d1a6ea8afff5

Observation 935d416b-ec31-4a3f-8205-bcbc4af244cc · inbound

Reducing the Safety Tax in LLM Safety Alignment with On-Policy Self-Distillation cites this paper.

Reducing the Safety Tax in LLM Safety Alignment with On-Policy Self-Distillation Bypassing the Safety Training of Open-Source LLMs with Priming Attacks

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:37:39.857254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-19T16:34:47.856606Z digest=sha256:762fa893c6a3c209893f2727d234a717b502d9c32d4bbc601053fb2336946a29

Observation e8332121-2dd1-4729-baf1-0eac60c9ba9f · inbound

Internal-State Probes Read the Situation, Not the Action: Three Negative Results for Pre-Action Misalignment Monitoring cites this paper.

Internal-State Probes Read the Situation, Not the Action: Three Negative Results for Pre-Action Misalignment Monitoring Bypassing the Safety Training of Open-Source LLMs with Priming Attacks

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:44:22.034796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-30T07:37:47.420986Z digest=sha256:be8bd90bb3a7e3c66754039b3a6d521a11ebae96f2851fcadc950e347de36668

Observation cf048a05-d422-4224-963d-5725ba18bb6e · inbound

Wait, am I Being Fair? Characterizing Deductive Stereotyping and Mitigating It with Fair-GCG cites this paper.

Wait, am I Being Fair? Characterizing Deductive Stereotyping and Mitigating It with Fair-GCG Bypassing the Safety Training of Open-Source LLMs with Priming Attacks

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-07-01T01:15:13.253118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T01:13:13.329619Z digest=sha256:c842bef660c7ddc2a12068e6c6666974583da58b1bfb27ada8ef1845ce2c85c7