Pith. sign in

Paper Citation Record · LEDGER

Weak-to-Strong Jailbreaking on Large Language Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2401.17256.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.17256 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T22:06:39.024687Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T09:29:44.257667Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c586368a-150e-4f55-84a3-29d1350e9d56 · inbound

Jailbreak Attacks and Defenses Against Large Language Models: A Survey cites this paper.

Jailbreak Attacks and Defenses Against Large Language Models: A Survey Weak-to-Strong Jailbreaking on Large Language Models

Reference 117

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:20:44.504149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T02:20:44.368219Z digest=sha256:ae921a39eece7188a5ea67ba03864148280306403ed0875b5fef4067db3b08f7

Observation 22481f73-652b-4e7a-abb2-ec51598fee79 · inbound

Peering Behind the Shield: Guardrail Identification in Large Language Models cites this paper.

Peering Behind the Shield: Guardrail Identification in Large Language Models Weak-to-Strong Jailbreaking on Large Language Models

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-23T03:45:21.416651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T03:45:14.234545Z digest=sha256:148d1599713433357ef1ac34b9498e362c8f708d2d03dcbeef4f51bf0e0e3791

Observation 2ed9ab0e-3aec-4809-979b-34f74362d5ad · inbound

Confidence Elicitation: A New Attack Vector for Large Language Models cites this paper.

Confidence Elicitation: A New Attack Vector for Large Language Models Weak-to-Strong Jailbreaking on Large Language Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-08T22:06:39.024687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:06:39.024687Z digest=sha256:09513144e4c78531d744e9da17a4333679d219ddd67f92991f51ec36e90913e4

Observation e17858c5-6305-4ec4-a4cc-c9d78526c34f · inbound

Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety cites this paper.

Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety Weak-to-Strong Jailbreaking on Large Language Models

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:42:34.219547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T04:39:04.591722Z digest=sha256:d3b33207387841fe94c4ffb1af87ade505ebed62c427639c02448ba720c74afa

Observation 65d59232-1a7e-4c6b-8b98-3e67cee7e60b · inbound

Bridging the Safety Gap: A Guardrail Pipeline for Trustworthy LLM Inferences cites this paper.

Bridging the Safety Gap: A Guardrail Pipeline for Trustworthy LLM Inferences Weak-to-Strong Jailbreaking on Large Language Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-08T10:23:06.305919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:23:06.305919Z digest=sha256:b2eea452a699d3f4f0ee8247f2a3502789e410117d8957f364d30b56c456baa5

Observation 36ca2ae1-9574-45a8-a4bf-cc5df38bc296 · inbound

Exploring Jailbreak Attacks on LLMs through Intent Concealment and Diversion cites this paper.

Exploring Jailbreak Attacks on LLMs through Intent Concealment and Diversion Weak-to-Strong Jailbreaking on Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:15.103581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:41:15.103581Z digest=sha256:4794964c067fbc5f966dcc90df5e4ec2a28d337ecdcef8ebe00482f01ec9ce3f

Observation eb482b29-547d-4846-bd6b-42e9cf4b000f · inbound

Security Concerns for Large Language Models: A Survey cites this paper.

Security Concerns for Large Language Models: A Survey Weak-to-Strong Jailbreaking on Large Language Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:03.759736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:03.759736Z digest=sha256:92cdc773e19e5462ea595a4a95ce31d1a7b52db4f59c7f0f0607f13e21996f3f

Observation c35c1c48-c59d-4924-9014-51bde1ac83e3 · inbound

SafeTy Reasoning Elicitation Alignment for Multi-Turn Dialogues cites this paper.

SafeTy Reasoning Elicitation Alignment for Multi-Turn Dialogues Weak-to-Strong Jailbreaking on Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:04:47.630473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:04:47.630473Z digest=sha256:51465df90182c632f33eb4fb2e20974ff1d6ec273539895d04716cf9058cc3aa

Observation 595452c4-6357-4760-8f52-30d432d8a57b · inbound

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures cites this paper.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Weak-to-Strong Jailbreaking on Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.606600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.606600Z digest=sha256:af24c79450522727c6969d2387b2c31ca7f9ada9bced66ff5a70ea3189d30567

Observation 014a108e-00d5-492c-8dad-fb42b13f5e67 · inbound

Toward Principled LLM Safety Testing: Solving the Jailbreak Oracle Problem cites this paper.

Toward Principled LLM Safety Testing: Solving the Jailbreak Oracle Problem Weak-to-Strong Jailbreaking on Large Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:42:12.897176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T08:40:56.186349Z digest=sha256:5f19a45b0ae7b5d494f521d4459d4e0a64dcbd1c66e7c6a9a9ba2e832dd03ad8

Observation 31368c05-277e-4bbf-b0e3-6e2c92dff6f8 · inbound

SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models cites this paper.

SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models Weak-to-Strong Jailbreaking on Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:57.504508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:57.504508Z digest=sha256:458c856162d6e62f3c2850216d7e5e25b24b0de91cf1bdc4e14b0b974aa973f6

Observation 2687ee65-1b68-46c1-91c9-905b049d5d40 · inbound

Logit Arithmetic Elicits Long Reasoning Capabilities Without Training cites this paper.

Logit Arithmetic Elicits Long Reasoning Capabilities Without Training Weak-to-Strong Jailbreaking on Large Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:01.155732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:01.155732Z digest=sha256:cd50e6616585bf6119272c9fcff8ad7c464f28541eed2e07628942aa64c04262

Observation f6b8c244-c238-478b-9765-920e0081a3a9 · inbound

Beyond Surface-Level Detection: Towards Cognitive-Driven Defense Against Jailbreak Attacks via Meta-Operations Reasoning cites this paper.

Beyond Surface-Level Detection: Towards Cognitive-Driven Defense Against Jailbreak Attacks via Meta-Operations Reasoning Weak-to-Strong Jailbreaking on Large Language Models

Reference 196

Resolution
unresolved
no resolver link, observed 2026-08-06T04:47:25.675116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T04:47:25.675116Z digest=sha256:21b33225c0aa4b764e2e45a022fbf54f9a04ced05b92a8ad2fb5d1c88b645334

Observation cd0431ae-c85a-46f5-878f-6dcce83602e6 · inbound

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation cites this paper.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Weak-to-Strong Jailbreaking on Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:33.631918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:33.631918Z digest=sha256:6bc3dd713f71c2faede727bff978e3d4df781577132f39c5c329d729d338f6cc

Observation 4bfedeb4-c79b-4671-9818-36d308a6b432 · inbound

SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses cites this paper.

SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses Weak-to-Strong Jailbreaking on Large Language Models

Reference 241

Resolution
unresolved
no resolver link, observed 2026-08-04T09:25:58.581856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:25:58.581856Z digest=sha256:12cd5452697bc671e90c4e7f2e306fc4e95c04bdd2ac478e766c29c05c93790a

Observation b2e2cef7-bad0-463a-8847-12988eecd8cc · inbound

Representation-Aware Unlearning via Activation Signatures: From Suppression to Entity-Signature Erasure cites this paper.

Representation-Aware Unlearning via Activation Signatures: From Suppression to Entity-Signature Erasure Weak-to-Strong Jailbreaking on Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T10:20:04.101125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:20:04.101125Z digest=sha256:22272d38bbae8dcaba5dc43bb4fb1a3c5f1c9f639c62be0ba47caae486d99b16

Observation a4cca6e4-e5df-47f5-9d7f-d15684c99b5a · inbound

CausalDetox: Causal Head Selection and Intervention for Language Model Detoxification cites this paper.

CausalDetox: Causal Head Selection and Intervention for Language Model Detoxification Weak-to-Strong Jailbreaking on Large Language Models

Reference 8

Resolution
malformed identifier
arxiv_id, observed 2026-05-10T12:05:21.962460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T12:04:13.987667Z digest=sha256:683fcaa69e47d677eae25ef88cb52b96d1957c498deb61272d4405ff56a0dd0a

Observation fe3c3664-c480-4d42-b6c2-57b112b72c48 · inbound

Skills as Verifiable Artifacts: A Trust Schema and a Biconditional Correctness Criterion for Human-in-the-Loop Agent Runtimes cites this paper.

Skills as Verifiable Artifacts: A Trust Schema and a Biconditional Correctness Criterion for Human-in-the-Loop Agent Runtimes Weak-to-Strong Jailbreaking on Large Language Models

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-19T18:27:43.007360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T18:25:12.407203Z digest=sha256:a02c8ff0f96c33a4d74cb3151b3ac40e330daca3ad4dfc969edefbf742a7b8a6

Observation e0b49d19-bc3f-44ad-a27c-e72e6c28fd64 · inbound

SlotGCG: Exploiting the Positional Vulnerability in LLMs for Jailbreak Attacks cites this paper.

SlotGCG: Exploiting the Positional Vulnerability in LLMs for Jailbreak Attacks Weak-to-Strong Jailbreaking on Large Language Models

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T13:26:59.307397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T01:16:07.252429Z digest=sha256:544ac320387aa10b0ad2200cd94c3fc755631a42cbc4bfddc078b4d482df1ef7

Observation 030fe290-72b6-4030-bff2-7b26a25262c6 · inbound

The Geometry of Refusal: Linear Instability in Safety-Aligned LLMs cites this paper.

The Geometry of Refusal: Linear Instability in Safety-Aligned LLMs Weak-to-Strong Jailbreaking on Large Language Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:29:44.259446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T09:52:52.603708Z digest=sha256:2d06a7ee84c925e85e447161aaac75ec814448acdcc10dfbbc6bd4022debda09

Observation 760ed90a-bad4-42fb-85e2-c538bd131da1 · inbound

The Geometry of Refusal: Linear Instability in Safety-Aligned LLMs cites this paper.

The Geometry of Refusal: Linear Instability in Safety-Aligned LLMs Weak-to-Strong Jailbreaking on Large Language Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:45:35.637196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T07:07:45.177242Z digest=sha256:1ec30dd94822ac1b9217a8682611523f0d789131eb3e6973a8b1ddc1a79c5dd3

Observation 52fadcb1-083a-450b-ba00-b19f4d4df369 · inbound

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions cites this paper.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Weak-to-Strong Jailbreaking on Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:55.053927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:55.053927Z digest=sha256:3f00603f52b66f03b3d47d18da15cb5e447894bc7867347f71353ef8cb4d872a

Observation 47529e3a-b68b-475c-a2c1-45808eaf61f4 · inbound

Weak-to-Strong On-Policy Distillation cites this paper.

Weak-to-Strong On-Policy Distillation Weak-to-Strong Jailbreaking on Large Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:23.761661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:23.761661Z digest=sha256:54ccc982c117d3ad79da6da7bfe366a19a20975e199ffaecdbc1bd81b7ff9bce