Pith. sign in

Paper Citation Record · LEDGER

Bypassing the Safety Training of Open-Source LLMs with Priming Attacks

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2312.12321.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.12321 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T19:32:24.131114Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T01:15:13.251025Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c2ac0286-28ca-49a7-a4d0-56b6ededdb7f · inbound

Fast Proxies for LLM Robustness Evaluation cites this paper.

Fast Proxies for LLM Robustness Evaluation Bypassing the Safety Training of Open-Source LLMs with Priming Attacks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T19:32:24.131114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T19:32:24.131114Z digest=sha256:1d1e715dbfd5ac9d283d4401edc46447bd688e95f5b3f736a397795297e3b824

Observation 72be71cf-0c7e-4576-a367-8697020cb15d · inbound

Saffron-1: Safety Inference Scaling cites this paper.

Saffron-1: Safety Inference Scaling Bypassing the Safety Training of Open-Source LLMs with Priming Attacks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:25.890123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:25.890123Z digest=sha256:7bf76acb9d8b082fa7b62393ff3e5d62e4bee4f6a26fff48218d9d2aa4da263c

Observation e0609a61-8420-4d81-8112-4db5f00d0667 · inbound

Few-Shot Truly Benign DPO Attack for Jailbreaking LLMs cites this paper.

Few-Shot Truly Benign DPO Attack for Jailbreaking LLMs Bypassing the Safety Training of Open-Source LLMs with Priming Attacks

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:07:27.007452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T07:06:46.387088Z digest=sha256:9247e8c3618d0cd3a49f933b3d9a549c412a262bb9365c0ac2037f5b1d822371

Observation 935d416b-ec31-4a3f-8205-bcbc4af244cc · inbound

Reducing the Safety Tax in LLM Safety Alignment with On-Policy Self-Distillation cites this paper.

Reducing the Safety Tax in LLM Safety Alignment with On-Policy Self-Distillation Bypassing the Safety Training of Open-Source LLMs with Priming Attacks

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:37:39.857254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-19T16:34:47.856606Z digest=sha256:fb2134b8b752adcce32879640f923f5bf463f4f2a6dfad393c011f70efd84955

Observation e8332121-2dd1-4729-baf1-0eac60c9ba9f · inbound

Internal-State Probes Read the Situation, Not the Action: Three Negative Results for Pre-Action Misalignment Monitoring cites this paper.

Internal-State Probes Read the Situation, Not the Action: Three Negative Results for Pre-Action Misalignment Monitoring Bypassing the Safety Training of Open-Source LLMs with Priming Attacks

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:44:22.034796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-30T07:37:47.420986Z digest=sha256:9c5ef5cd56a1c212e89e7f3a2b464c507dc6e9910cfb3f16d1ec205732f3ae65

Observation cf048a05-d422-4224-963d-5725ba18bb6e · inbound

Wait, am I Being Fair? Characterizing Deductive Stereotyping and Mitigating It with Fair-GCG cites this paper.

Wait, am I Being Fair? Characterizing Deductive Stereotyping and Mitigating It with Fair-GCG Bypassing the Safety Training of Open-Source LLMs with Priming Attacks

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-07-01T01:15:13.253118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T01:13:13.329619Z digest=sha256:b3b2099eabeafa5e58e85e6af9647a67a9c6879667160738f1e955546b87b287