Pith. sign in

Paper Citation Record · LEDGER

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models

As of 10 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2510.20129.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.20129 v2

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-18T05:22:30.509147Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact20
  • verified fuzzy6
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 60d19697-c414-403d-8613-47f1d2a45592 · outbound

This paper cites GPT-4 Technical Report.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models GPT-4 Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:25:54.588103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:cb8f2754b374cfdc48343f047e89136c531f4c914b5962b53cda815a8a11cfc4

Observation 62a58878-b48d-4562-bf88-5231492a62b3 · outbound

This paper cites Detecting Language Model Attacks with Perplexity.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models Detecting Language Model Attacks with Perplexity

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:25:54.601980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:795a9e2ccedcd4477cbe07fed2571e2db505fe382ad1b5f7735cb527062fab65

Observation d0fb6236-87cb-456b-af2e-ef2b9a3d95de · outbound

This paper cites Efficient Training of Language Models to Fill in the Middle.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models Efficient Training of Language Models to Fill in the Middle

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:25:54.581354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:dc1c555b0aac4fa3b77290f7e63c80843bae4accfa981b6e9a00bdae843267e4

Observation b0aa928f-7f2a-492b-87c5-85c78b430438 · outbound

This paper cites On the dangers of stochastic parrots: Can language models be too big? InProceedings of the 2021 ACM conference on fairness, accountability, and transparency, pp.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models On the dangers of stochastic parrots: Can language models be too big? InProceedings of the 2021 ACM conference on fairness, accountability, and transparency, pp

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:25:54.975897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:b1a91bdef5ae8fa8294bdb2c88cf5ff945b9f11e18b1643f9adf3bb938768f8e

Observation 80d9fd52-404a-445d-8d57-deb9b2c81485 · outbound

This paper cites Single-pass Detection of Jailbreaking Input in Large Language Models.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models Single-pass Detection of Jailbreaking Input in Large Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:25:54.598504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:0f9a495475287106e4e0e84afeb6d17de4bc614d7566f8851fe1b647b528adb5

Observation 8ac98264-c6e6-4e9f-bfb7-50c9b0c50203 · outbound

This paper cites A survey on evaluation of large language models.ACM transactions on intelligent systems and technology, 15(3):1–45, 2024a.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models A survey on evaluation of large language models.ACM transactions on intelligent systems and technology, 15(3):1–45, 2024a

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:25:54.969547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:93dc7e56ea1911f7bac9357d98e4ea772d588eba5e3623d6e7e12381c1fd2ab6

Observation a33a43cc-c162-4524-a220-dcbf48b5f7a0 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.See https://vicuna.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.See https://vicuna

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:25:54.972940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:5cd57220b072b0bb0962ec738357be3ef99f5b9e6182bce57a62c3190106cd1d

Observation bdcdeac0-827d-4bbc-9058-f8fe59b29b4a · outbound

This paper cites Attack Prompt Generation for Red Teaming and Defending Large Language Models.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models Attack Prompt Generation for Red Teaming and Defending Large Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:25:54.594563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:1a39821fc7bcdde1db51c50b46b12f115ebec1a5734343032d56567c2490a23d

Observation 46a53fd7-b283-434a-8f45-3a266a9a2acb · outbound

This paper cites Shaping the Safety Boundaries: Understanding and Defending Against Jailbreaks in Large Language Models.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models Shaping the Safety Boundaries: Understanding and Defending Against Jailbreaks in Large Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:25:54.605699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:cf2575b40fde2e84b2b8eaeacd4a37ab3234252a3dbcbd9abbd816f9c636fb9c

Observation 5653751b-b50d-4246-aec2-17d7a6739f09 · outbound

This paper cites ChatGLM-RLHF: Practices of Aligning Large Language Models with Human Feedback.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models ChatGLM-RLHF: Practices of Aligning Large Language Models with Human Feedback

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:25:54.591261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:da56a34a4b29d339d62478d56f0d2753fa36f966485ab0779a3f06d7b4252f09

Observation 3334fb1b-1f0b-4af7-83ed-f8955a7ebb35 · outbound

This paper cites Baseline Defenses for Adversarial Attacks Against Aligned Language Models.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models Baseline Defenses for Adversarial Attacks Against Aligned Language Models

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:25:54.584578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:126ad0d341da3255b4e3ace3b328f06272f76ebbe1c85a0497b4cc76710a511b

Observation c39e4aac-3abc-478b-9ea1-51ed769a62ea · outbound

This paper cites Scaling Laws for Neural Language Models.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models Scaling Laws for Neural Language Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:25:54.608709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:5d6bba8a0fb64b28a8e3dc11c901e4277528a03a2848fa3090e694fa96291424

Observation 692ae75b-197e-4f6e-8e29-e1d2b887ca24 · outbound

This paper cites DeepInception: Hypnotize Large Language Model to Be Jailbreaker.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models DeepInception: Hypnotize Large Language Model to Be Jailbreaker

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-19T10:37:07.279605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:fb37b9a157d79c4ad12d9282c8670a5d490cbc0bfdbd91951d465ebb19cdba5d

Observation 3630742f-6a01-4f88-983f-14b53afa37ef · outbound

This paper cites Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T05:25:54.619328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:7fba049ffb9e3e62949e1524ad5c37b1cc9e6000e008345c4fdae7662d2f8b29

Observation a4f16710-d197-41d0-9f06-c86237e9b564 · outbound

This paper cites The Unlocking Spell on Base LLMs: Rethinking Alignment via In-Context Learning.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models The Unlocking Spell on Base LLMs: Rethinking Alignment via In-Context Learning

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:25:54.623202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:4821e37cb20848c1a2148bc2431260cbb365951ef4c14879cadb7c9fdf119df2

Observation 7604e27f-5245-40be-ac1c-8aba2983e0c0 · outbound

This paper cites AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:25:54.615726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:cf4a409961221ee0fce26a56df86c3dce78d396ebf5a1ddca5284f9a2cfd9f11

Observation eddb566c-e329-4d43-988d-fec1b861e745 · outbound

This paper cites FlipAttack: Jailbreak LLMs via Flipping.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models FlipAttack: Jailbreak LLMs via Flipping

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-20T00:00:16.351675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:0e79068c33bae82a7c4425d10260b5a355913fa3247124daca1a861b6fdddaf3

Observation f005551b-3535-4fbc-8b50-44a2164ce102 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:25:54.612403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:1e8acb9f730b218bbc260e91a93bbd7d58ffc019b4296828b0ff692930e8eed6

Observation 8b9af65f-eef9-4e39-b148-2c7c4b2f4287 · outbound

This paper cites Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:25:54.630924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:4f745949140f52e932492eae6721908c4ecd1d1a76d0c6ae64d2bb56744ab4dd

Observation f4a23f62-8ce5-41e7-8783-73cd59afbd12 · outbound

This paper cites Beyond Surface-Level Patterns: An Essence-Driven Defense Framework Against Jailbreak Attacks in LLMs.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models Beyond Surface-Level Patterns: An Essence-Driven Defense Framework Against Jailbreak Attacks in LLMs

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:25:54.651445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:0e6fc10aedd9f025332cb1bbe3ed0866bbba4f76611d9fcb0dafc9a382c56d32

Observation fcdc1f0f-0310-4bd3-85ac-c878661eb97d · outbound

This paper cites Aligning Large Language Models via Fine-grained Supervision.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models Aligning Large Language Models via Fine-grained Supervision

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:25:54.634495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:6d222a77734433e235146b66e57eeb37187e863c9be75bcaf287e047c8e1aa44

Observation e1dc0f50-33f0-44c8-a642-1a38297e42ae · outbound

This paper cites SQL Injection Jailbreak: A Structural Disaster of Large Language Models.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models SQL Injection Jailbreak: A Structural Disaster of Large Language Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:25:54.641114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:d5ff292e1a6fcd4e1d234d186770c5d662f4ae88422ff91fde3973786fdaea55

Observation 47c0bf16-1594-4405-a165-94c5d2ae83f1 · outbound

This paper cites A Survey of Large Language Models.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models A Survey of Large Language Models

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:25:54.637716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:c86c4fc33e45cdd5d990833cca40fa87ace83a536abb22b5655840c06fcd83b9

Observation bde3143e-20e0-43e6-a32c-d846de78f1bd · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:25:54.644521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:19823c1ec521714690be9602c84a3bb44d6997a547db629676dd1ee5f13273ed

Observation 26b76b3f-b214-4dfe-876f-9241edfc35ca · outbound

This paper cites comment- ing out.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models comment- ing out

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:25:54.981075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:0333de7fb1bebaf9a2ba36c8bb1222790390c5353e1b30623c802f05fcc31c49

Observation 0ff281eb-4bc6-480f-8262-4ba0b439132c · outbound

This paper cites Can I”, significantly outperform prefixes that carry a strong personal or instructional intent, like “Can you teach me to do this to others.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models Can I”, significantly outperform prefixes that carry a strong personal or instructional intent, like “Can you teach me to do this to others

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:25:54.983873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:79263aef162d624564cfc496e720d65cf56a3d5468b5fb3371c6de4855be837b

Observation e22bf799-5ac6-4635-9ffa-62842cc5e0f9 · outbound

This paper cites role-playing.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models role-playing

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:25:54.978587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:19c35375c8a81bad902f431fc0057742a413b4e294f2e4e5065fc2e4b23625b9

Pith citing papers

No inbound Pith citation observations are available.